Claude Desktop's model picker lists around twenty entries, and the names don't explain themselves. This article sorts them into the four families that matter, says what each is for and what it costs your key, and explains why most models appear twice.
Cornell bills gateway usage at Anthropic's published rates with no markup, so every price below is what actually lands on your KFS account.
ℹ️ Prices and version numbers move. Everything here is accurate as of September 2026. AI at Cornell's Available Models & Pricing page (NetID sign-in) is the source of truth for current rates; where the two differ, trust that page and tell us.
The four families
Every Claude model is the same assistant with a different balance of how hard it thinks, how fast it answers, and what it costs. The family name is the part that tells you what to expect; the number after it is just the generation.
| Family |
Thinking power |
Speed |
Input |
Output |
Reach for it when… |
| Fable |
●●●●● Highest |
●○○○○ Slowest |
$10 |
$50 |
You've already tried Opus and it wasn't enough |
| Opus |
●●●●○ Very high |
●●○○○ Deliberate |
$5 |
$25 |
The problem is hard, long, or high-stakes |
| Sonnet |
●●●○○ High |
●●●●○ Quick |
$2 |
$10 |
Almost anything — the everyday choice |
| Haiku |
●●○○○ Solid |
●●●●● Instant |
$1 |
$5 |
The job is short, simple, or repeated a lot |
Prices are US dollars per million tokens. Input is what you send (your prompt, attached documents, and the whole conversation so far); output is what Claude writes back.
💰 What a token costs you in practice. A token is roughly three-quarters of a word, so a 50-page PDF is about 25,000 tokens. Asking for a two-page summary of it costs roughly 3¢ on Haiku, 7¢ on Sonnet, 16¢ on Opus, and 33¢ on Fable. The ratios are what to remember: Sonnet is about twice Haiku, Opus about five times, Fable about ten times.
One thing that surprises people: the whole conversation is re-sent with every message, so a long chat gets more expensive per turn as it goes. Starting a fresh conversation for a new topic is the cheapest habit you can form.
Sonnet — start here
The right default for most of the day: drafting and editing, summarising papers and meetings, building a deck or a spreadsheet, thinking a problem through with you. Strong enough for nearly everything, fast enough not to interrupt you, and cheap enough that a busy month stays in the single digits of dollars.
The trade-off: on the very hardest reasoning it may need a second pass where Opus would have got there first time.
Opus — when it has to be right
Use it when the task has real depth: reworking a grant narrative against a solicitation, reading a full manuscript for gaps in the argument, planning a multi-stage analysis, or anything where being wrong is expensive.
The trade-off: it takes visibly longer to answer and costs two and a half times Sonnet for the same work. Asking Opus for a synonym is like booking a lecture hall for a phone call.
Haiku — the quick one
Best for short, bounded, repetitive jobs: reformatting a list, tidying references, pulling the dates out of an email thread, a first-pass sort across many items.
The trade-off: it will take a confident swing at something complicated rather than reasoning it through. Give it work you could check at a glance.
Fable — the top tier
Fable is Anthropic's most capable tier, above Opus. It is the right tool for a small number of tasks — deep research synthesis, long autonomous work in the Code tab, problems where Opus visibly falls short — and the wrong tool for everything else, because it costs twice Opus and ten times Haiku for the same words.
One difference worth knowing. Fable is hosted for Cornell by Amazon Web Services under a slightly different arrangement from the other models: AWS retains prompts and responses sent to it for up to 30 days for automated abuse detection (flagged requests may be reviewed by a person at AWS; nothing is shared with Anthropic). AI at Cornell has confirmed this sits within Cornell's AWS agreement and remains compatible with Moderate Risk data, so it needs nothing from you — it is just the one model where "we don't store your prompts" has a footnote.
A rule of thumb. Start on Sonnet. Move up to Opus when you catch yourself re-explaining or correcting it. Drop to Haiku when you're doing the same small thing over and over. Reach for Fable only when Opus has already let you down on something that matters.
Why most models appear twice — the "1M" entries
In the picker, most models show up as a pair: claude-sonnet-5 and claude-sonnet-5 1M, claude-opus-5 and claude-opus-5 1M, and so on.
They are the same model at the same price. The only difference is how much the conversation can hold at once — the model's context window:
|
Context window |
Roughly |
| Standard |
200,000 tokens |
150,000 words — a 500-page book |
| 1M |
1,000,000 tokens |
750,000 words — five of them |
For almost all work the standard entry is more than enough; a conversation would have to contain several full books before it hit the ceiling. Pick the 1M entry only when you need Claude to hold a very large body of material in a single conversation — a whole dissertation with its sources, a semester's worth of transcripts, a large codebase — and would otherwise see it say the conversation is too long.
⚠️ "Same price" means the same price per token — not the same bill. A 1M conversation can hold five times as much, and you pay for everything it holds on every turn. A chat carrying 800,000 tokens of documents costs about $1.60 per message on Sonnet, $4 on Opus, and $8 on Fable in input alone, before Claude writes a word. Use the room when you need it, and start a fresh conversation when you don't.
Not every model has a 1M entry. Haiku and the pre-4.6 generations (Opus 4.5 and 4.1, Sonnet 4.5 and 4) only come in the standard size, which is why they appear once.
Older version numbers
The picker also lists earlier generations — Opus 4.8, 4.7, 4.6, 4.5, 4.1; Sonnet 4.6, 4.5, 4. They stay available so that work built on them keeps running, but there is no reason to choose one for new work: the newest number in each family is at least as capable and never more expensive.
Two of them are worth actively avoiding:
- Opus 4.1 costs
$15 / $75 — three times the current Opus for a less capable model. Anthropic is retiring it on 8 October 2026.
- Sonnet 4 is past its retirement date and may disappear from the list at any time.
Everything else in the picker is a current-generation model at the current-generation price.
ℹ️ What the number means. Opus 5 is newer than Opus 4.8, which is newer than Opus 4.7. A higher number within a family is always the more recent model. Between families the number tells you nothing — Haiku 4.5 is not "older" than Sonnet 5 in any way that matters; it is just a smaller model.
Where to watch what you're spending
Every request also carries a flat gateway fee of $0.0002 (two cents per hundred requests) on top of the token cost. It is too small to change any of the advice above, but it is why very short exchanges never quite cost zero.
Your key's Spend History in the key portal shows exactly what you have spent and on which models — see Managing, Rotating, and Renewing Your Key. If a model that was working suddenly stops with a message about the server "temporarily limiting requests", you have most likely hit your monthly spend cap; that article explains how to tell and what to do.
Further reading
Questions about a model, its cost, or whether it is right for your work? Email itrequests@business.cornell.edu.