Future Product
Issue № 007 · 27/09/2026
Capability

Opus 5.5 Forces a New Agent Cost Test

Anthropic’s new Opus matches Fable 5.1 performance at a lower price, putting long-running coding agents back on every product team’s spreadsheet.

AI-written, human-edited, never fabricated. How this is made

an unnamed product manager comparing two agent workflow cost curves on a large screen

The first product decision is not whether to celebrate Anthropic’s latest model. It’s whether your agent budget is now based on a model that costs too much. Claude Opus 5.5 arrived on September 22 with reported Fable 5.1-level performance, output that is more than 30% faster, and pricing of $4 per million input tokens and $20 per million output tokens. Anthropic also says cache reads are 60% cheaper. For a team running a coding agent across a large repository, those are not cosmetic improvements. They change which workflows can make financial sense.

I’m using “agentic coding” here in the practical sense: a model that can inspect a codebase, call tools, make changes, run checks, and continue through several steps rather than answer one isolated programming question. That distinction matters because the cost of a single prompt is rarely the cost that breaks a product plan. The expensive bit is the loop. A model reads context, proposes a change, calls a tool, reads the result, revises its plan, and repeats. Lower the price of every pass, and suddenly a workflow that was reserved for a small pilot can look plausible at production volume.

The headline comparison is unusually direct. Anthropic says Opus 5.5 matches the performance of Fable 5.1 while costing 40% less, alongside stronger scores on agentic benchmarks. Its reported result on Terminal-Bench 4.0 is 66.4%, a number meant to capture performance on tasks that require a model to work through a terminal rather than simply generate a code snippet. Anthropic’s launch post presents the pricing and benchmark claims. The independent Artificial Analysis report places Opus 5.5 at the top of its index, with a score in the 57.6 to 58 range and improved cost efficiency.

Those figures reset the control sample. If your team last benchmarked an agent against an older premium model, that result is now a historical baseline, not a safe planning assumption. The relevant question is no longer, “Can the agent complete this task?” It is, “Can it complete the task reliably enough, quickly enough, and cheaply enough to justify leaving it in the loop?” A model that is only slightly better at the task but materially cheaper per run can win. A model that is both cheaper and stronger deserves an immediate slot in the test harness, even if you have no intention of switching vendors this week.

Reprice the loop

The spreadsheet needs to stop treating tokens as a flat unit cost. Input and output are priced differently, cache reads have their own economics, and long-running agents generate a shape of usage that a chatbot forecast misses. Opus 5.5’s stated $4 input and $20 output prices make the output side particularly important when the agent is explaining intermediate steps, generating patches, or returning large results. The 60% reduction in cache-read costs matters when the same repository context is carried across repeated turns. It can make persistence less punitive, but it does not make an undisciplined agent free.

Here is the concrete exercise I’d put in front of a PM and an engineering lead this week. Take one existing workflow, not a demo task. Record the repository or document context sent on every turn, the number of tool calls, the output tokens, the number of retries, and the point at which a human takes over. Then run the same workflow with Opus 5.5 and compare total cost per successful completion, not cost per request. If the old model finishes in fewer turns, it may still win. If Opus 5.5 takes more turns but its lower price and faster output produce a cheaper successful run, that is a meaningful product result. Without this accounting, “40% cheaper” is a headline floating above the actual unit economics.

The speed claim deserves similar discipline. More than 30% faster output sounds like a direct improvement to user experience, but an agent’s wall-clock time includes tool execution, queueing, retries, and the time a person spends reviewing changes. Faster generation may shorten the part users notice least. Or it may make an interactive coding workflow feel less like waiting for a remote colleague and more like working alongside one. The only honest way to know is to measure time to accepted result, with the human review step included.

The model you benchmarked last quarter may already be the wrong model for your cost model.

That is why the Terminal-Bench result is useful and insufficient at the same time. A 66.4% score gives product teams a common signal for terminal-based agent work. It says Opus 5.5 belongs in serious comparisons, not merely in a marketing appendix. But it doesn’t tell you how the model handles your repository, your test suite, your deployment permissions, or the strange internal conventions that make a seemingly simple change risky. A benchmark can move a candidate into the queue. It cannot move a candidate through your release process.

Benchmark the handoff

The product consequence is broader than a model swap. When the cost of a strong agent falls, teams can reconsider where the human enters the workflow. A developer might ask for a first-pass implementation, inspect the diff, and send the agent back to address failing checks. A support product might use an agent to trace a reported issue through a codebase before handing a concise diagnosis to an engineer. These workflows were always technically imaginable. The new pricing makes their operating assumptions worth revisiting.

That does not mean handing Opus 5.5 every task. A cheaper premium model can still be the wrong choice for a lightweight request, and the evidence here does not establish that every task is faster or more accurate. Product teams should compare it against the model they use today on the work that creates real cost: multi-step changes, repeated context, failed attempts, and review-heavy output. A small benchmark of polished examples will flatter any capable system. A week of representative tickets will be less glamorous and more useful.

The comparison with Fable 5.1 sharpens the strategic pressure. If Opus 5.5 offers that level of performance at 40% lower cost, the old assumption that “frontier quality requires a premium budget” is no longer stable. It also makes previous routing rules suspect. Teams may have separated easy tasks from hard tasks using a price threshold that has quietly moved. Some work previously routed to a cheaper model may now be worth sending to Opus 5.5 if fewer retries, better tool use, or less human correction lower the total cost. The reverse may be true for short, predictable requests where a higher-capability model adds no value.

Artificial Analysis’s cost-efficiency result is important for the same reason. The report’s leading index position, at 57.6 to 58, is not a promise about your product’s quality. It is evidence that capability and price are moving together in a way that should disturb settled procurement logic. Product teams often freeze model choices into an architecture because changing them feels operationally expensive. A new model at this price point raises the cost of not rechecking. Your current model may remain the best fit, but now it has to win a fresh comparison.

Don’t average away failure

The benchmark plan should preserve the failures that a blended score hides. Separate tasks that complete cleanly from tasks that require a retry. Track how often an agent produces a plausible but unusable change. Record review time, not just pass rate. Keep the task mix stable enough that the result can be compared again next month. Most importantly, calculate cost per accepted outcome. A model that saves tokens while creating more review work has not necessarily saved money. A model that costs more per attempt but reaches an accepted result in one pass may be the better product choice.

There is a counterpoint to the excitement, and it is a substantial one: these reported scores come from launch materials and an independent index, not from your production environment. Nobody outside Anthropic’s lab and the benchmark operators knows how Opus 5.5 will behave on every private codebase, tool wrapper, or workflow constraint. “Top-tier” is a reasonable description of the reported results. It is not permission to delete your existing safeguards, skip evaluation, or promise customers a new level of reliability. The model has earned a serious trial, not automatic promotion.

Still, waiting for perfect certainty is a poor response to a cost curve that has already changed. This week, I would add Opus 5.5 to the same evaluation set as the incumbent model, replay a representative batch of agent tasks, and put the total successful-run cost beside quality and time to completion. I’d then rerun the forecast for the longest workflows, where cheaper cache reads and lower token prices have the most room to matter. If the model loses, the decision is informed. If it wins, the roadmap has a new economic option.

The sharpest implication is simple: model selection is now a recurring product metric, not a one-time infrastructure choice. Opus 5.5’s arrival means your agent’s cost model may be stale even if the agent itself is working exactly as designed. The teams that benefit first won’t be the ones that repeat the 66.4% score most loudly. They’ll be the ones that measure what a successful run costs, then have the nerve to change the workflow when the number moves.

Sources