Field guide · Claude lineupSpecs verified 2026-10-01Selected: Claude Opus 5.5
The Model Sequence.
Four models, plotted the way an astronomer would: capability against latency, priced by point size. It's not quite a main sequence (more of a giant branch), but the physics of the tradeoff is real. Pick a star and the whole guide follows.
CLICK OR TAB TO A STAR · THE CARD, CONSOLE & SNIPPET BELOW FOLLOW YOUR SELECTION
01 · Object of interestSpectral cardPricing & specs from the Claude Platform docs
02 · SpectrographFeed in a task, read out a modelAnswer all three; the verdict updates live
Q1 · How hard is the task, honestly?
Q2 · What are you optimizing?
Q3 · Any scale factors?
03 · ThrottleThe effort parameterOne dial, five stops; it shapes every token, thinking and tool calls included
high
Typical use
Token spend
Tool behavior
Available on
Solar analogue: active region
# ● on the rail marks the selected model's default
response = client.messages.create(
model="claude-opus-5",
max_tokens=4096,
messages=[{"role": "user", "content": prompt}],
output_config={"effort": "high"}, # the whole trick
)
Gotchas worth memorizing
The default is per model now: high on Fable 5.1 and Sonnet 5.5, medium on Opus 5.5. Omitting the parameter runs Opus 5.5 one notch lower than Opus 5 did.
Change effort mid-conversation with a per-message output_config (beta header mid-conversation-output-config-2026-07-01); it keeps the prompt cache. Changing the top-level value between requests still restarts it.
It's a behavioral signal, not a hard budget: at low effort, Claude still thinks on genuinely hard problems, just less.
Lower effort means fewer, terser tool calls: operations get combined and preambles vanish.
Thinking is always on for Opus 5.5 and Fable 5.1; thinking: disabled returns a 400. Sonnet 5.5's floor is between_tools, and that one is only legal at low, medium, and high.
Sonnet 5.5's levels are recalibrated: a level does not buy the same thinking it did on Sonnet 5. Re-run the effort sweep instead of carrying settings over.
Set max_tokens generously at high and above: it caps thinking plus reply combined. For agentic coding on Sonnet 5.5 the docs say 128k and stream.
Where you set it: the API (output_config) and Claude Code. In claude.ai it mostly stays behind the curtain.
04 · Observing proposalsThirty tasks from your actual weekRecommendations are this guide's judgment, applying the official guidance
Nothing matches; clear the search or switch category.
05 · The catalogSide by sideClick a column header to sort
Model ↕
$ In/MTok ↕
$ Out/MTok ↕
$ Cache hit ↕
Context ↕
Max output ↕
Latency ↕
Knowledge cutoff ↕
Default effort
Thinking
Effort levels
Cache hit = price per MTok read back from the prompt cache (0.025× input on Fable 5.1, 0.05× on Opus 5.5, 0.1× elsewhere). Batch API discounts stack on top. Claude 4.7+ models use a tokenizer that yields roughly 30% more tokens for the same text.
Still callable, previous generation (Claude API, tentative retirement not sooner than): Fable 5 $10/$50 (2027-06-09) · Opus 5 $5/$25 (2027-07-24) · Sonnet 5 $2/$10 (2027-06-30) · Opus 4.8 / 4.7 / 4.6 $5/$25 · Sonnet 4.6 $3/$15. Haiku 4.5 itself is the oldest current model: not sooner than 2026-10-15.
06 · PhotometryWhat will it cost?Per request, list prices; Batch API would halve it