Anthropic released Opus 5 on July 24, 2026

Claude Opus 5for work that holds up

Since launch, Felo PRO users can try Claude Opus 5 free for 7 days.

Anthropic Claude Opus 5 official illustration framed in Felo blue
Visual source: Anthropic, framed for Felo without changing its original elements.

Now in Felo Search

Ask harder questions. Get a search answer that holds up.

Claude Opus 5 is now available in Felo Search for questions that need deeper investigation, stronger source synthesis, and a final answer checked against the evidence.

Try Opus 5 in Felo Search

01

Trace the real question

Work through the constraints behind an ambiguous request instead of stopping at the first plausible interpretation.

02

Weigh conflicting evidence

Bring competing sources and claims into one investigation, then identify where the evidence does and does not agree.

03

Return with fewer loose ends

Use the answer as a decision-ready starting point, with the reasoning, caveats, and open questions still visible.

Quick walkthrough

Switch to Claude Opus 5 in Felo

Open the model picker, select Claude Opus 5, then start your task in the same workspace.

Jul 24

2026 release

Anthropic's newest Opus model

$5 / $25

API input / output

Per 1M tokens, same base price as Opus 4.8

~2.5x

Fast mode speed

Available at 2x the base price on Anthropic's platform

Verify + revise

Long-running work

Planning, tool use, checks, and iteration

The Opus difference

Do not stop at the first thing that works.

Anthropic positions Opus 5 for work where sustained judgment matters: identify the underlying issue, validate the result, and revisit the work when the evidence does not hold up.

01

Trace the underlying problem

For difficult bugs and ambiguous assignments, start with evidence, constraints, and the actual root cause rather than a surface-level patch.

02

Check the work before handoff

Use a model that can inspect its own output, test assumptions, and catch an issue before a person has to discover it later.

03

Keep the thread across revisions

Bring code, source material, tool results, and new constraints into one longer task without reducing the job to disconnected prompts.

Model selection

Compare the cost of a token with the cost of getting it wrong.

Prices below are official public API rates, not Felo subscription or usage prices. They make the trade-off visible without pretending that unrelated vendor latency tests are directly comparable.

Decision factor
Claude Opus 5
GPT-5.6 Sol
Kimi K3
Claude Opus 4.8
API price / 1M tokens
$5 input / $25 output
$5 input / $0.50 cached / $30 output
$3 input / $0.30 cached / $15 output
$5 input / $25 output
Speed and capacity
Default speed, or Fast mode at about 2.5x speed for 2x base price.
OpenAI labels Sol as Fast; 1.05M-token context window.
1,048,576-token context; throughput varies by account rate-limit tier.
Prior Opus baseline. Compare its live availability and latency in your workflow.
Best fit
High-consequence, long-running work requiring verification and careful iteration.
Complex professional work with fast flagship-model response and a very large context window.
Long-horizon coding and end-to-end knowledge work with a lower published token price.
Existing Opus workflows where continuity is more important than moving immediately.
Choose it when
The task must be finished, checked, and defensible - not merely drafted.
You want OpenAI's flagship tier and its speed/context profile suits the task.
You need a 1M-token context and want to control cost through caching and reasoning effort.
You are evaluating a controlled migration from an established previous-generation baseline.

Official public API pricing checked July 25, 2026. Total task cost also depends on output length, cached input, effort settings, tool calls, retries, and human review.

Official benchmark charts

What Anthropic's cost-versus-performance charts show

These official charts plot model performance against reported evaluation cost at different effort levels. They are useful context for a cost decision, but they are vendor-published results rather than independent Felo testing.

Anthropic official Claude Opus 5 benchmark capability overview comparing Opus 5, Fable 5, Opus 4.8, and GPT-5.6 Sol across coding, knowledge work, search, computer use, business workflows, health, and biology

Official capability overview across benchmark families

Anthropic's cross-task matrix compares the published results for Opus 5, Fable 5, Opus 4.8, and GPT-5.6 Sol across coding, knowledge work, search, computer use, business workflows, and more.

Anthropic official Agentic coding by effort level chart comparing Opus 5, Fable 5, Opus 4.8, and GPT-5.6 Sol by cost per task

Agentic coding by effort level

Artificial Analysis Coding Agent Index: the official chart compares score with reported cost per task across the available effort settings.

Anthropic official Agentic computer use performance by effort level chart comparing Opus 5, Fable 5, Opus 4.8, and GPT-5.6 Sol by cost per task

Agentic computer use by effort level

OSWorld 2.0: the official chart compares computer-use performance with reported cost per task across effort settings.

Anthropic official Real-world knowledge tasks by effort level chart comparing Opus 5, Fable 5, Opus 4.8, and GPT-5.6 Sol by benchmark cost

Real-world knowledge tasks by effort level

GDPval-AA v2: the official chart compares Elo score with reported full-benchmark cost across effort settings.

Anthropic official Agentic search by effort level DeepSearchQA chart comparing Opus 5, Fable 5, and Opus 4.8 by cost per task

Agentic search by effort level

DeepSearchQA: the official chart compares search-task pass rate with reported cost per task across available effort settings.

Charts and benchmark labels are from Anthropic's Claude Opus 5 announcement. Felo added only the outer presentation frame; chart contents are preserved as published.

A better starting point

Give Opus 5 the jobs that deserve a second pass.

The model earns its place when verification changes the result. These are not generic chat prompts; they are multi-step jobs where a missed dependency, unsupported claim, or untested edge case costs time later.

Production code and QA

Investigate a regression, reproduce it, inspect the surrounding code, implement a focused fix, and test the paths a quick patch tends to miss.

Financial and document analysis

Reconcile tables, test assumptions, compare source documents, and produce a recommendation with the evidence and caveats intact.

Research that needs scrutiny

Read competing sources, expose open questions, weigh the evidence, and distinguish a confident answer from an answer that is actually supported.

Long-running agent tasks

Use tools, preserve the task constraints, recover from partial results, and keep moving through a larger assignment with deliberate checkpoints.

Try a task that has consequences

01

Bring the real context

Add the code, documents, requirements, and evidence that a useful answer must account for.

02

Ask for a verification plan

Specify what must be checked before the work is considered complete, not only what to produce.

03

Compare the final artifact

Use the same task across leading models in Felo. Review not only the answer, but the tests, caveats, and judgment behind it.

Select Claude Opus 5 when it is available in your account's live model picker.

Open Felo LLM Playground

Claude Opus 5 FAQ

Claude Opus 5 is Anthropic's Opus-tier model released on July 24, 2026. Anthropic positions it for long-running agents, complex software engineering, professional knowledge work, and tasks that benefit from verification and careful iteration.

Use the model that checks the work.

Put Claude Opus 5 beside the other leading models in Felo, then test it on the task where a plausible answer would not be enough.

Try Opus 5 in Felo

Start in Felo AI Search and select the model where available.