Skip to main content

Claude Opus 5 puts Fable on trial: when is the top model still worth double?

11 min read

Claude Opus 5 keeps Opus 4.8 pricing while matching or beating Fable 5 on several launch tests. Here is where the cheaper model wins, where Fable still earns a place, and how builders should route real workloads.

Claude Opus 5 puts Fable on trial: when is the top model still worth double?

Imagine opening your model router on Monday and discovering that the “second-best” option now wins several of the tests you cared about, costs half as much as the top tier, and is already the default on your paid plan. That is the awkward decision Anthropic just created.

Anthropic released Claude Opus 5 on July 24. It is available across Claude’s paid products and API, where it keeps Opus 4.8’s price of $5 per million input tokens and $25 per million output tokens. Fable 5 remains above it in Anthropic’s lineup at $10 and $50.

That simple 2x price gap is more interesting than the launch adjectives. Opus 5 beats Fable 5 on several of Anthropic’s headline evaluations, comes within fractions of a point on others, and loses a few that matter. The question is no longer “Which Claude model is smartest?” It is “Where does Fable earn the extra bill?”

My answer: fewer places than it did yesterday, but not zero.

Anthropic squeezed a lot of Fable into the Opus price

Opus 5 is now the default model on Claude Max and the strongest model offered on Claude Pro. Developers can call it with claude-opus-5. It has a 1 million-token context window, can return up to 128,000 tokens in one response, and has a May 2026 reliable knowledge cutoff.

Those specifications are already serious. The launch table is where the product line gets uncomfortable.

On Frontier-Bench v0.1, which Anthropic describes as agentic terminal coding, Opus 5 scored 43.3%. Fable 5 reached 33.7%, GPT-5.6 Sol 34.4%, and Opus 4.8 21.1%. On GDPval-AA v2 knowledge work, Opus 5 led the same group with 1861.

Frontier-Bench v0.1 launch results

Anthropic-reported mean reward over five attempts per task. Higher is better; bars are scaled to Opus 5’s 43.3% result.

Claude Opus 5
43.3%
GPT-5.6 Sol
34.4%
Claude Fable 5
33.7%
Claude Opus 4.8
21.1%

ARC-AGI-3 is the wild-looking result: 30.2% for Opus 5, compared with 7.8% for GPT-5.6 Sol and 1.5% for Opus 4.8. Fable 5 was not reported in that row, so it would be wrong to turn this into an Opus-versus-Fable win.

There are more:

  • Opus 5 scored 90.8% on BrowseComp, just ahead of GPT-5.6 Sol at 90.4%.
  • It reached 70.6% on OSWorld 2.0, ahead of Fable 5 at 66.1%.
  • It posted 26.0% on AutomationBench, while the next result in Anthropic’s table was GPT-5.6 Sol at 18.1%.
  • On Humanity’s Last Exam with tools, Opus 5 scored 64.7% and Fable 5 scored 63.9%.

These are Anthropic-published launch evaluations, not independent MustHave.ai testing. Some also depend on effort settings, tools, harnesses, and fallback behavior. I would use them to choose which models deserve an internal trial, not to skip that trial.

This is not a clean sweep

The same table contains the reasons not to delete Fable from your router.

Fable 5 edged Opus 5 on DeepSWE v1.1, 69.7% to 68.8%. GPT-5.6 Sol led that row at 72.7%. Fable also beat Opus by a tenth of a point on FrontierCode v1.1 Main, 53.5% to 53.4%, and led the held-out Legal Agent Benchmark, 13.3% to 11.7%.

Without tools, Fable scored 56.5% on Humanity’s Last Exam and Opus scored 56.3%. In health, Mythos 5 led HealthBench Professional at 66.0%, while Opus 5 reached 59.8%.

That pattern matters. Opus 5 looks like the better default, but a default is not a universal winner. If your business depends on one narrow task, the average benchmark story is almost irrelevant. Your own pass rate is the benchmark.

This is the same discipline I recommend in our guide to choosing an AI model without losing a weekend to benchmarks: test the work you actually sell or ship, include failure cases, and count the human repair time.

The real control knob is effort

Anthropic is not presenting Opus 5 as one fixed cost-performance point. The model can spend different amounts of effort on a task, and the launch charts plot quality against cost at those settings.

That changes how I would deploy it. I would not send every email draft, classification job, and one-line code edit to maximum effort. Anthropic’s prompting guide says low and medium effort can preserve strong quality with fewer tokens and lower latency. For difficult coding and agentic work, it recommends starting at xhigh, then measuring.

The important word is measuring.

Run the same representative task set at low, medium, high, and xhigh. Record final correctness, tool failures, retries, latency, tokens, and the minutes a human spends repairing the answer. The cheapest setting is the one with the lowest completed-task cost, not the smallest API line item.

Anthropic also says Opus 5 verifies its work without needing the old “check everything again” scaffolding many agent prompts accumulated. Keeping redundant verification instructions can make it over-check and burn tokens. The model may also expand a task’s scope or narrate more than prior Opus releases, so concise scope and response-length rules are worth adding.

That is an unusually practical migration note. A smarter model can cost more because your old prompt tells it to do work it already does by default.

What the 2x price gap means in a real budget

The base API prices are straightforward:

  • Opus 5: $5 per million input tokens, $25 per million output tokens.
  • Fable 5: $10 per million input tokens, $50 per million output tokens.
  • Sonnet 5: $2 and $10 through August 31, 2026, then $3 and $15.
High-volume tier
Claude Sonnet 5
  • Input / MTok$2, then $3
  • Output / MTok$10, then $15
  • Context1M tokens
Hard-work default
Claude Opus 5
  • Input / MTok$5
  • Output / MTok$25
  • Context1M tokens
Frontier premium
Claude Fable 5
  • Input / MTok$10
  • Output / MTok$50
  • Context1M tokens

Sonnet 5’s $2/$10 introductory rates run through August 31, 2026. Its published standard rates are $3/$15.

Take a modest monthly agent workload that consumes 10 million input tokens and 2 million output tokens. Opus 5 would cost $100: $50 in and $50 out. The same volume on Fable 5 would cost $200.

That is a $1,200 annual difference for one workload of that size. Ten similar production workflows turn it into $12,000. This is simple arithmetic, not a promise about anyone’s actual bill, but it gives the benchmark debate a useful unit: Fable has to save more than $100 per month in errors, labor, or missed outcomes on that workload.

Fast mode needs its own line in the budget. Anthropic says Opus 5 Fast runs about 2.5 times the default speed, but its token prices double to $10 and $50. That is exactly Fable 5’s base token rate. You are buying speed, not moving up to the Fable capability tier.

For background on why this kind of routing is becoming normal, our recent AI price-war breakdown explains how quickly model families are spreading across price and performance tiers.

The behavior change may save more than the token discount

The strongest part of Anthropic’s launch case is not a chart. It is the claim that Opus 5 keeps working until the job is genuinely checked.

In one reported evaluation, the model had to reconstruct a machine part from an image it could not inspect through the normal visual interface. It built a computer-vision pipeline to extract the geometry from pixel data, then created the part in FreeCAD. In another case, it found the root cause of a real package-manager bug and caught an edge case missed by an existing community patch.

Anthropic also describes an engineer asking Opus 5 to build a market-data feed for a new exchange. With no live feed available, the model created a test harness to validate its parser.

These are hand-picked vendor examples, so I would not treat them as a reliability guarantee. They do point at the right evaluation target: does the agent find a way to prove its own work, or does it stop when the output merely looks finished?

For builders, I would test four things:

  1. Give Opus 5 the full task, the completion criteria, and the tools it needs.
  2. Remove repeated verification instructions unless your eval shows they still help.
  3. Put a hard boundary around scope, permissions, and what the agent must not change.
  4. Score the evidence it returns alongside the final answer.

That is also why our plain-English guide to agentic AI focuses on the job boundary and proof of completion rather than the “autonomous” label.

Safety here is a routing system, not one big refusal switch

Anthropic reports that Opus 5 scored 2.3 on its automated behavioral audit for overall misaligned behavior, the lowest result among its recent models. The company says it followed Claude’s Constitution more consistently, showed less deceptive behavior, and was harder to trick into misuse than Opus 4.8, Sonnet 5, or Fable 5.

That is an Anthropic assessment, not a neutral certification. Still, the implementation details are worth understanding.

Anthropic says it deliberately avoided cyber-specific training for Opus 5. General capability gains still made it much better at finding software vulnerabilities, where it approaches Mythos 5. It remains well behind Mythos at turning those findings into working exploits.

The cyber classifiers reflect that split. They allow source-code vulnerability discovery, while blocking binary-based scanning, penetration testing, and exploit generation. Anthropic expects those classifiers to intervene about 85% less often than Fable 5’s classifiers.

If a request is flagged in Claude.ai, Claude Code, or Claude Cowork, it falls back to Opus 4.8 by default. API users can opt into automatic fallbacks. Anthropic also launched mid-conversation tool changes that do not invalidate the prompt cache, which should make it easier to tighten or expand an agent’s tool set as the job changes.

There is another operational difference hiding below the benchmarks: Opus 5 has no general-access data-retention requirement. Fable 5 requires 30-day retention and is not available under zero-data-retention arrangements. For some enterprise workloads, that policy decides the route before a benchmark does.

My routing recommendation

A practical three-tier router

Start with the lowest tier that passes your own task evaluation, then escalate only when the result earns the price.

S
Route routine volume to Sonnet 5
Use it for the high-frequency jobs that already meet your quality bar at the lower rate.
O
Make Opus 5 the hard-work default
Start here for difficult coding, long documents, computer use, and agentic jobs that need careful verification.
F
Do not pay for Fable by habit
Keep it only where your eval shows enough gain to cover a 2x token-price premium or a policy requirement does not exclude it.

If you are on Opus 4.8, Opus 5 deserves the first migration test. The price is unchanged, the context window stays at 1 million tokens, and Anthropic says existing Opus 4.8 prompts should work well. Do not flip production traffic blindly; re-run your effort sweep and watch for longer narration, broader scope, and extra self-checking.

If you use Sonnet 5, keep it for routine volume that already passes your eval. Opus 5 should earn the harder cases: multi-file coding, long documents, computer use, difficult tool loops, and work where the cost of a plausible-but-wrong answer is high.

If you pay for Fable 5, make it defend the premium. Keep it on the few workloads where it wins your tests by enough to offset double the token rate, or where its long-horizon autonomy changes the outcome. Route the rest to Opus 5.

This is how I already think about Claude versus ChatGPT in real work: choose by task, not by team jersey. Opus 5 makes that approach more important inside Anthropic’s own lineup.

Opus 5 is the model that makes routing boring

Fable 5 still has a job. Sonnet 5 still has a job. Opus 5 simply takes a much larger share of the middle, including some tasks that no longer look “middle” at all.

That is why this launch matters to me. The useful result is not that Anthropic won another benchmark screenshot. It is that builders can put a near-frontier model on more everyday work without paying the frontier rate every time.

Start with Opus 5 at a sensible effort level. Keep Sonnet underneath it for cheap volume. Let Fable prove where it earns the premium. Then revisit the router when the next model lands, because at this release pace, it will not stay settled for long.

Check the sources before you switch

Which workload would you trust to Opus 5 first, and what would Fable have to do better to stay in your router?

Research note: benchmark results, behavioral-audit findings, safety behavior, and launch examples in this article are reported by Anthropic. Prices and model specifications were checked against Anthropic’s documentation on July 24, 2026. The monthly workload calculation is a MustHave.ai example, not a customer case study.

Leave a comment

Your email address will not be published. Required fields are marked *