gilly.space
Field guide · Claude lineup Specs verified 2026-10-01 Selected: Claude Opus 5.5

The Model
Sequence.

Four models, plotted the way an astronomer would: capability against latency, priced by point size. It's not quite a main sequence (more of a giant branch), but the physics of the tradeoff is real. Pick a star and the whole guide follows.

◀ FASTER · RESPONSE LATENCY · SLOWER ▶ CAPABILITY ▶ FASTEST FAST MODERATE SLOWER CLAUDE MYTHOS 5.1 · SAME SPECS · INVITATION-ONLY (PROJECT GLASSWING) HAIKU 4.5 $1 / $5 SONNET 5.5 $2 / $10 OPUS 5.5 $4 / $20 FABLE 5.1 $10 / $50 POINT SIZE ∝ PRICE PER MTOK DASHED RING = INVITATION-ONLY TWIN
CLICK OR TAB TO A STAR · THE CARD, CONSOLE & SNIPPET BELOW FOLLOW YOUR SELECTION
01 · Object of interest Spectral card Pricing & specs from the Claude Platform docs
02 · Spectrograph Feed in a task, read out a model Answer all three; the verdict updates live
Q1 · How hard is the task, honestly?
Q2 · What are you optimizing?
Q3 · Any scale factors?
03 · Throttle The effort parameter One dial, five stops; it shapes every token, thinking and tool calls included
high

Typical use
Token spend
Tool behavior
Available on

# ● on the rail marks the selected model's default
response = client.messages.create(
    model="claude-opus-5",
    max_tokens=4096,
    messages=[{"role": "user", "content": prompt}],
    output_config={"effort": "high"},  # the whole trick
)

Gotchas worth memorizing

  • The default is per model now: high on Fable 5.1 and Sonnet 5.5, medium on Opus 5.5. Omitting the parameter runs Opus 5.5 one notch lower than Opus 5 did.
  • Change effort mid-conversation with a per-message output_config (beta header mid-conversation-output-config-2026-07-01); it keeps the prompt cache. Changing the top-level value between requests still restarts it.
  • It's a behavioral signal, not a hard budget: at low effort, Claude still thinks on genuinely hard problems, just less.
  • Lower effort means fewer, terser tool calls: operations get combined and preambles vanish.
  • Thinking is always on for Opus 5.5 and Fable 5.1; thinking: disabled returns a 400. Sonnet 5.5's floor is between_tools, and that one is only legal at low, medium, and high.
  • Sonnet 5.5's levels are recalibrated: a level does not buy the same thinking it did on Sonnet 5. Re-run the effort sweep instead of carrying settings over.
  • Set max_tokens generously at high and above: it caps thinking plus reply combined. For agentic coding on Sonnet 5.5 the docs say 128k and stream.
  • Where you set it: the API (output_config) and Claude Code. In claude.ai it mostly stays behind the curtain.
04 · Observing proposals Thirty tasks from your actual week Recommendations are this guide's judgment, applying the official guidance
Nothing matches; clear the search or switch category.
05 · The catalog Side by side Click a column header to sort
Model ↕ $ In/MTok ↕ $ Out/MTok ↕ $ Cache hit ↕ Context ↕ Max output ↕ Latency ↕ Knowledge cutoff ↕ Default effort Thinking Effort levels

Cache hit = price per MTok read back from the prompt cache (0.025× input on Fable 5.1, 0.05× on Opus 5.5, 0.1× elsewhere). Batch API discounts stack on top. Claude 4.7+ models use a tokenizer that yields roughly 30% more tokens for the same text.

Still callable, previous generation (Claude API, tentative retirement not sooner than): Fable 5 $10/$50 (2027-06-09) · Opus 5 $5/$25 (2027-07-24) · Sonnet 5 $2/$10 (2027-06-30) · Opus 4.8 / 4.7 / 4.6 $5/$25 · Sonnet 4.6 $3/$15. Haiku 4.5 itself is the oldest current model: not sooner than 2026-10-15.

06 · Photometry What will it cost? Per request, list prices; Batch API would halve it