← All insights

Fable 5. Which Claude Model Should You Actually Use?

Fable 5 vs Sonnet 5 vs Opus 4.8 (and GPT‑5.6)

Newsletter artwork for “Fable 5. Which Claude Model Should You Actually Use?”
Real benchmarks, real token prices, and a role-by-role cheat sheet — plus the Cowork 2× trick heavy users are sleeping on. No spec-sheet dumping, I promise.

Hey! Welcome to the latest Creators’ AI edition.

Let me tell you about the dumbest money I spent this spring.

Back in March I did what every overexcited AI person does when a new frontier model drops: I set it as my default for everything. Every Claude Code session, every content draft, every “hey summarize this email” — all routed to the biggest, smartest, most expensive model I had access to. Because why would I use the cheap one? I’m a professional. I want the good stuff.

Then the API bill came in. My always-on agent — the one that runs my agency and this newsletter out of a single Telegram chat — went from a comfy ~$25/month to $500+ for doing the exact same work it did the month before. Same tasks. Same output. Just a far dumber way of paying for them. (I told that whole horror story in I Rebuilt My OpenClaw Setup on Hermes + GPT‑5.5.)

Hero — five circles growing small to large, a spectrum of the models; the middle one ringed in amber as the everyday default

Here’s the thing I learned the expensive way, and the whole point of today’s post:

The question is never “which model is the smartest?” It’s “which model, for THIS task, on THIS surface, at THIS price?”

Capability got cheap in 2026. Routing is the actual skill now. And most people — smart people, people who build for a living — are still reaching for the most expensive model by reflex.

So let’s fix that. By the end of this you’ll have a one-screen cheat sheet, by role and by task, plus a Cowork trick that’s quietly saving heavy users half their usage. Have a seat, it’s going to be practical.

The five models on the table (60-second version)

If you’ve been anywhere near AI Twitter this month you already know the drama: Anthropic launched Fable 5, the government yanked it offline three days later, then it came back — this time with Sonnet 5 riding shotgun. (We covered the whole soap opera in the Fable 5 Goes Public and Fable 5 Shutdown digests — genuinely wild fortnight.)

Here’s the cast, one honest line each:

Claude Sonnet 5 is the new default and the real story of the year — near-Opus quality at Sonnet money. Claude Opus 4.8 is the reliable, level-headed one you call in for long, gnarly, don’t-lose-the-thread work. Claude Fable 5 is the frontier — the biggest brain, the biggest lead on the biggest problems, and a price tag to match. On the OpenAI side, GPT‑5.6 just landed as a tiered family (Sol, Terra, Luna), with Sol setting a genuine terminal-coding record — and GPT‑5.5, the incumbent workhorse that half the agent setups on the internet still quietly run on a $20 subscription.

That’s the whole field. Now let’s talk about who’s actually good at what — with receipts.

Keep your mailbox updated with practical knowledge & key news from the AI industry!

What the benchmarks actually say (and where they lie)

I’m going to show you the numbers, but I want you to hold them loosely, because we’ll get to why in a second.

On raw coding, Fable 5 is the king — 95% on SWE-bench Verified, 80.3% on SWE-bench Pro. Opus 4.8 sits at 69.2% Pro, Sonnet 5 at 63.2%. Fable’s lead is real, and here’s the important part: it grows as the task gets harder. On small stuff everyone’s basically tied; on the big, ugly, multi-system problems, Fable pulls away. That single fact is your escalation signal, so tuck it in your pocket.

But then you look at terminal work — actual autonomous agents living in a shell, running commands — and the leaderboard scrambles. GPT‑5.6 Sol takes the crown at 91.91%. GPT‑5.5 is at 83.4%. And here’s the plot twist that should change how you work: Sonnet 5 (80.4%) beats Opus 4.8 (74.6%) on this one. The cheaper Claude out-agents the pricier one in the terminal.

cai-model-benchmarks-table.png

Now, the “hold them loosely” part. Benchmarks in 2026 are saturating and, frankly, a little compromised. Datacurve audited SWE-bench Pro — the one everybody quotes — and found 8% false positives and 24% false negatives. And GPT‑5.6 Sol? Its own system card admits “instances of the model cheating on tasks and fabricating research results.” An independent evaluator clocked its cheating rate higher than any public model they’d tested.

Three near-identical contestants measured by a warped, bent tape measure — benchmarks are saturating and easy to game

So treat every number here as directional, not gospel. (When DeepSWE first scrambled the leaderboard, we walked through what it exposed in Opus 4.8, $965B, GPT‑5.5 Wins DeepSWE.) Which is exactly why a benchmark table alone is useless, and why the rest of this post exists.

The part that can kill your project: what this stuff costs

Every other “which model” guide stops at benchmarks. That’s like reviewing cars and never mentioning the price. So here’s the table that actually runs your bank account:

cai-model-pricing-table.png

Stare at that for a second and one thing jumps out: output tokens cost about 5× what input does. And guess what agents produce buckets of? Output. Every reasoning step, every tool call, every “let me try that again” — output tokens, ticking like a taxi meter in traffic.

A little robot cranking out pages with a taxi meter on its back ticking up — output tokens are the real cost driver

Here’s the math that fixed my brain. Take one meaty coding task that spits out ~200K output tokens. On Sonnet 5 that’s $2. On Fable 5 it’s $10. Same task. Now multiply by the dozens of runs you do a week, and you understand how I got to a $500+ surprise while feeling like a responsible adult the whole time.

The flip side: there’s now a genuinely cheap tier for grunt work. GPT‑5.6 Luna at $1/$6 is built for high-volume, low-drama jobs — bulk copy, classification, tagging. And if you’re hammering the same context over and over, GPT‑5.5’s cached input at ~90% off is criminally underused.

Share this post with a friend who’s about to default everything to Fable 5!

Share

The cheat sheet, by who you are

Okay. Story time’s over, here’s the payoff. Screenshot this section — it’s the whole post in one place.

A task-line branching at a junction, the right path lit amber to its matching circle — choosing a model is routing to the right-sized option

If you’re a developer

Make Sonnet 5 your daily driver in Claude Code and mean it. It’s fast, it holds a million tokens of context, and — as we just saw — it out-agents Opus in the terminal. Escalate to Opus 4.8 when you hit a long, multi-file refactor and Sonnet starts losing the plot, and pull out Fable 5 only for the genuinely novel, brain-melting problems where its lead actually shows up. Terminal-agent purist with preview access? GPT‑5.6 Sol is your record-holder; if not, GPT‑5.5 is the dependable fallback. The one rule: don’t set Fable 5 as your Code default. That’s the most expensive habit in this whole post, and I already made it so you don’t have to. (If you want to see how other builders wire their stacks, I reverse-engineered three of them in 3 Creator AI Stacks, Fully Reverse-Engineered.)

If you’re a vibe coder

Sonnet 5, full stop. It’s the cheapest path to “wait, it just… worked,” and the giant context window means you can paste your whole messy project in and let it sort you out. Graduate to Opus 4.8 when your app grows past the toy stage and things start breaking in ways you can’t describe. Skip Fable 5 — you’d be paying 5× on output for genius you’re not going to use yet. (New to this? Our vibe coding cases and tips is a good on-ramp.)

If you’re a marketer

Live in Cowork with Sonnet 5 — it’s built for exactly the multi-source, “pull the brief, the data, and last quarter’s campaign together” work you do all day. Need 300 ad variations? Drop down to GPT‑5.6 Luna or cached GPT‑5.5 and save your premium budget for the thinking. When it’s real positioning or strategy, Opus 4.8 earns its keep. Fable 5 for a subject line is like renting a stadium for a dinner party. (We built a full Claude Cowork for Marketing playbook if you want the actual workflows.)

If you’re a content creator

Sonnet 5 for drafting and voice-matching — the long context means it can actually read your back catalog and sound like you, not like LinkedIn. When you’re shipping a flagship piece and want a real editor’s eye on structure, bump up to Opus 4.8. Save Fable 5 for the occasional deeply-researched, genuinely-hard investigative piece.

If you’re a founder juggling everything

This one’s personal, because it’s me. Run an operator harness — I use Hermes — on a flat $20 subscription + GPT‑5.5 for all your always-on ops: briefings, standup summaries, “what’s overdue” pings. Do not meter Fable 5 for that stuff (see: $500+). Then reach for Opus 4.8 or Fable 5 deliberately, per hard task — the fundraising model, the thorny analysis — not as a background default. And for anything cross-tool, live in Cowork and abuse the 2× window (more on that in a sec). If you’re scaling past one agent, How to Hire & Manage AI Agents is the companion piece.

cai-model-by-role-matrix.png

The same thing, by task (for people who think in jobs)

If roles aren’t your mental model, here’s the other lens. Daily coding and most agent work → Sonnet 5. Long-horizon, multi-file, don’t-lose-the-thread coding → Opus 4.8. The hardest, most novel problem you’ve got, the kind whose complexity keeps growing → Fable 5. Terminal-native autonomous agents → GPT‑5.6 Sol if you’re in the preview, GPT‑5.5 if not. High-volume cheap generation → GPT‑5.6 Luna or cached GPT‑5.5. Cross-tool synthesis when you’re not a dev → Sonnet 5 in Cowork (wiring your tools together with MCP makes this dramatically better).

And one gotcha worth its own line: if you send Fable 5 anything touching cybersecurity, biology, chemistry, or model distillation, it silently routes to Opus 4.8 anyway. So you’d be paying Fable rates for Opus work and not even know it. Don’t.

The Cowork 2× trick (do this before Saturday)

Quick tactical one, because the clock is literally ticking as you read this.

Through July 5, 2026, Anthropic doubled the 5-hour usage limit inside Claude Cowork for everyone on Pro, Max, or Team. No coupon, no clicking — if you’re on an eligible plan it’s just on. The catch: it’s Cowork only (not web/desktop/mobile chat), and it’s the 5-hour limit, not the weekly one.

Translation: for the next few days you can run twice as much premium-model work — Opus and Fable-class heavy lifting — inside Cowork before you hit a wall. If you’ve got a big synthesis project, a campaign build, a research crunch? This is the window. Batch it now.

By the time you’re reading this the exact promo may have closed — Anthropic reruns these, so the tactic (batch heavy Cowork work into a 2× window) is evergreen. Check if it’s live and act accordingly. We broke down what to actually run in Cowork in 10 Claude Cowork Prompts That Replace Your Boring Admin Work.

Sharing is caring! Refer someone who’s still paying frontier prices for grunt work.

Refer a friend

How I actually keep my bill sane

Since I opened with my $500+ mistake, let me close the loop on what fixed it.

I stopped defaulting and started routing — cheapest model that clears the task, escalate only when it visibly fails. I lean on cached input for anything with repeated context (that 90% GPT‑5.5 discount is free money). And the big structural fix: I moved my always-on agent off metered API tokens and onto a flat $20 subscription through Hermes, self-hosted on a $5 VPS. Same eight workflows, same Telegram chat, back down to ~$25/month all-in — and it no longer flinches every time a lab changes its pricing. The full migration is in I Rebuilt My OpenClaw Setup on Hermes + GPT‑5.5 if you got stranded the same way I did.

cai-openclaw-cost-chart.png

The honest caveats (there’s always a catch)

Because I’d be a hypocrite to sell you certainty after that benchmark section: the numbers saturate and can be gamed, so treat them as directional. GPT‑5.6 Sol and Terra are preview-gated right now under a US-gov restriction — they may simply not be available to you yet, record or no record. Sol’s own paperwork admits it cheats, so verify its output. Fable 5’s safety fallback means you won’t actually get frontier-class scores in cyber/bio/chem. And nearly every price here is promo-loaded — Sonnet 5’s intro rate, the Cowork 2× — and will drift once the IPO earnings-call era kicks in. Re-check before you hard-wire any of this into a workflow.


So, what should you actually do?

If you remember one sentence, make it this: Sonnet 5 is your new default, Opus 4.8 is your judgment call, Fable 5 is for the genuinely hard, and GPT‑5.6/5.5 cover terminal-native and cheap-bulk. Capability is cheap now. Routing is the skill. Stop paying frontier prices for everyday work.

Screenshot the role matrix, pick your default, and be honest about when you actually need to escalate — it’s a lot less often than your ego wants it to be.

Now I’m curious about yours: what’s your daily-driver model right now, and what do you escalate to? Drop your routing setup in the comments — I collect these, and the good ones end up in a future edition.

Leave a comment

Archive note

This article was first published in the Creators AI newsletter. View the original edition.

Keep exploring

More in Claude & Anthropic.