Claude Opus 5.5 model release — Anthropic AI pricing and benchmarks

Anthropic Just Made Fable-Level Performance Cost $20 Per Million Tokens. Here's What Actually Changed.

Claude Opus 5.5 launched September 22 at $4/$20 per million tokens — 40% cheaper than Opus 5 on typical workloads, and Fable 5.1-level on most benchmarks. Terminal-Bench ties GPT-6 Astra. Cache reads down 60%. But the 40% savings assume medium effort, there are 5 breaking changes, and one fails silently.

TL;DR

- Anthropic released Claude Opus 5.5 on September 22. Pricing: $4 per million input tokens, $20 per million output tokens — down from $5 and $25 for Opus 5. - Cache reads dropped 60%, from $0.50 to $0.20 per million tokens. Total operating cost is approximately 40% lower than Opus 5 on typical workloads — but that figure assumes your workload runs at medium effort, which is Opus 5.5's default. Opus 5's default was high. If you need high or extended effort, the savings narrow to around 20-30%. - On Terminal-Bench 4.0, Opus 5.5 scores 59.6% — tying GPT-6 Astra and up 11 points from Opus 5. On the private Artificial Analysis Briefcase benchmark, Opus 5.5 hits an Elo of 1,822, up 143 from Fable 5.1. It's the first Anthropic model to beat GPT-5.6 Sol on presentation quality. - Anthropic describes Opus 5.5 as matching Claude Fable 5.1 "on most tasks." Fable 5.1 costs $10 per million input tokens and $75 per million output tokens. Opus 5.5 lists at 40% of Fable's per-token price. - Sonnet 5.5 and Haiku 5.5 are coming "in the coming weeks." Measure Opus 5.5 at low effort against Sonnet 5 before assuming the larger model costs more per completed task. - There are five breaking changes. Four are documented. One fails silently. - This is also Anthropic's first model release since Dario Amodei publicly called for pacing the AI frontier.

The headline for Opus 5.5 is not the benchmark. It is the pricing structure.

Fable 5.1, Anthropic's current frontier model, costs $10 per million input tokens and $75 per million output tokens. Opus 5.5 is priced at $4 input and $20 output, and Anthropic says it matches Fable 5.1 "on most tasks." If that holds under your workload — and the independent benchmarks suggest it largely does — you are getting Fable-level capability at 27% of the output price.

That is the decision this release puts in front of every team currently running Fable 5.1 in production.

What the pricing actually is

$4 per million input tokens. $20 per million output tokens. Cache writes at $5 per million (five-minute) and $8 per million (one-hour). Cache reads at $0.20 per million — down 60% from Opus 5's $0.50. Batch API at $2 input and $10 output.

Anthropic's claim that Opus 5.5 costs 40% less than Opus 5 on typical workloads is real but requires a specific condition: it assumes you are running at medium effort, which is the default for Opus 5.5. Opus 5's default was high. If your production workload was running Opus 5 at high effort, and Opus 5.5's medium produces equivalent output quality on your specific tasks, you get the full 40% reduction — from cheaper per-token prices plus fewer tokens generated.

If your workload requires high or extended effort on Opus 5.5 to match your existing output quality, you are looking at closer to 20-30% savings from the token price cuts alone.

This matters because the 40% figure is being quoted in most coverage without the effort-level caveat. Measure your actual cost per completed task on your own workload before committing to a migration plan based on that number.

For teams on Fable 5.1, the math is more straightforward. At 40% of Fable's per-token price, you need roughly similar output quality to see substantial savings. The benchmarks suggest that threshold is largely met.

What the benchmarks show

Artificial Analysis, the independent benchmarking service, placed Opus 5.5 at the top of its intelligence index immediately after release. The model hits 59.6% on Terminal-Bench 4.0, tying GPT-6 Astra and beating Opus 5 by 11 points. That is the benchmark most relevant to long-horizon agent and terminal work — the same benchmark where, last week, StepFun Step 5 Preview scored 33.3%.

On GDPval-AA v2.1, four of five Opus 5.5 effort levels land on the Pareto frontier — matching or beating every other model above 50 on the intelligence index for cost per task. On the private AA-Briefcase benchmark measuring presentation quality, Opus 5.5 achieves an Elo of 1,822, up 143 from Fable 5.1 — the first Anthropic model to beat GPT-5.6 Sol on that evaluation.

Anthropic's own coding benchmark, FrontierCode 1.1, shows Opus 5.5 approaching Fable-level performance at half the cost, with 22% improvement over Opus 4.7 on hard agentic coding tasks and lower variance run-to-run. Lower variance matters for production agent deployments — an unpredictable p95 on task completion is often the real operational cost, not the average.

On life sciences evaluations covering structural biology, organic chemistry, and bioinformatics, Opus 5.5 improves over Opus 5 across the board. Anthropic notes it carries the same cybersecurity and biology safeguards as Fable 5.1 — meaning both models have the same restrictions on discovering exploits in compiled programs and biological weapon capability research.

The safety guardrails caveat

This is worth dwelling on. Opus 5.5 is Anthropic's first model with Fable 5.1-level capability safeguards — restrictions on the most sensitive use cases in cybersecurity and biology. This means Opus 5.5 carries both Fable-level capability and Fable-level restrictions. If your production workload previously ran on Opus 5 and relied on capabilities that fall into those restricted categories, Opus 5.5 may refuse tasks that Opus 5 completed.

Anthropic's documentation on the biology safeguards specifically notes restrictions around "developing recognizable biological weapons." The cybersecurity restrictions limit autonomous exploitation of compiled programs. For enterprise deployments in security research, vulnerability disclosure programs, or life sciences, review the Fable 5.1 safeguard documentation before migrating — the same capabilities that make Opus 5.5 the strongest model in its tier also come with the same capability restrictions.

The breaking changes

Five. Four documented, one silent.

The four documented breaking changes affect: tool use schema validation (stricter in Opus 5.5, some previously passing schemas now rejected), thinking block handling (Fable 5.1 thinking blocks cannot be carried into Opus 5.5 conversations), default effort level (medium, not high — changing output behavior and token counts for integrations that relied on default effort), and extended thinking timeout defaults (changed, may affect long-running agent tasks).

The fifth breaks silently: certain edge cases in structured output generation that previously returned a best-effort response now return an empty string. No error is raised. Integrations that depend on non-empty structured output without explicit validation may fail in ways that are not immediately visible.

If you are migrating from Opus 5 or Fable 5.1, run your full integration test suite against Opus 5.5 before cutover. The silent failure mode on structured output is the one most likely to produce subtle production bugs rather than obvious errors.

What's coming

Sonnet 5.5 and Haiku 5.5 are coming "in the coming weeks." Before assuming Sonnet 5 is the cheaper option until then, Anthropic's guidance is to measure Opus 5.5 at low effort against Sonnet 5 on your specific workload. At low effort with lower token counts, Opus 5.5's per-task cost may be competitive with Sonnet 5 for some query types.

For Claude Code users on subscription: Anthropic says five-hour usage limits are 20% higher under Opus 5.5, and the model goes approximately 25% further within those limits. If you are hitting context limits regularly, this is a meaningful practical change.

The timing

Anthropic released Opus 5.5 three days after Dario Amodei's public essay calling for a paced approach to frontier AI development drew same-day agreement from Altman, Musk, and Hassabis. This is the first model Anthropic has shipped since that announcement.

The release is not inconsistent with the "pacing" framing — Opus 5.5 is explicitly positioned as a cost reduction and efficiency improvement on existing capabilities, not a capability leap beyond the current frontier. Anthropic's blog post references the infrastructure investment required to sustain this kind of improvement curve and says more details are coming. The model matches Fable 5.1. It doesn't beat it by a wide margin on most benchmarks.

What pacing means in practice, for a lab releasing models, is a question this release begins to answer: improve efficiency and access to existing capability levels rather than racing to push the absolute frontier. Whether the industry follows that framing or treats it as a one-week talking point before the next capability release is a different question. Anthropic's model release calendar will be the evidence.

The decision rule

If you are running Fable 5.1 today: run Opus 5.5 at high effort against your hardest production tasks. It lists at 40% of Fable's price. The Terminal-Bench and GDPval benchmarks suggest you will find it largely equivalent on coding and knowledge work. Check the three likely breaking changes (schema validation, thinking blocks, effort defaults) and measure the output quality delta on your specific use case before full migration.

If you are running Opus 5 today: run Opus 5.5 at medium effort and measure actual cost per completed task against your existing workload, not the 40% headline. If medium-effort Opus 5.5 matches your high-effort Opus 5 output, the savings are real. If it doesn't, set effort to high and expect closer to 20-30%. Check the silent structured output failure mode explicitly.

If you are waiting for Sonnet 5.5: it's coming in weeks. In the meantime, benchmark Opus 5.5 at low effort against Sonnet 5 before assuming the smaller model is always cheaper per task.