Skip to main content

Claude Opus 5.5 Is Now in Felo Search: Answers That Hold Up

· 8 min read
Felo Search Tips Buddy
Committed to answers at your fingertips

Claude Opus 5.5 is live in Felo Search: 1M context, $4/$20 per million tokens, and answers verified against the evidence before you act on them.

Anthropic released Claude Opus 5.5 on September 22, and the number that matters for search is not on the benchmark table. It is the price: $4 per million input tokens and $20 per million output, against Opus 5's $5 and $25. Anthropic puts the operating-cost reduction at roughly 40% on typical workloads, with output more than 30% faster.

Opus 5.5 is live in Felo Search now. Cheaper and faster would be a good enough story on its own. What makes it worth switching to for research is that it got more careful at the same time.

Claude Opus 5.5 in Felo Search: a one-million-token context holding sources, evidence, and caveats together

What Claude Opus 5.5 Actually Is​

Opus 5.5 is Anthropic's first release since the company publicly called for pacing frontier development in September. It sits at the top of the Opus line, above Claude Sonnet 5, with the Fable line reserved for a different kind of deployment.

The specification is built for work that has to be defended rather than merely produced: a one-million-token context window, a 128,000-token maximum output, and cache reads at $0.20 per million tokens that make long, iterative sessions affordable to repeat.

SpecClaude Opus 5.5
Model IDclaude-opus-5-5
MakerAnthropic
ReleasedSeptember 22, 2026
Context window1,000,000 tokens
Max output128,000 tokens
InputText, image
OutputText
Price$4 per million input, $20 per million output
Cached input$0.20 per million tokens
Prior generationOpus 5, released July 24, 2026 at $5 / $25

Two rows in that table are the reason a research workflow notices this model.

A million tokens of context. That is roughly 750,000 words, a stack of source documents you would otherwise have to summarize down before asking a question about them. Summarizing first is where errors enter: the summary is written by a model that did not know which detail you would need.

Cache reads at a fifth of the input price. A Felo Search investigation is not one call. It is a first pass, a challenge, a revision, and a synthesis, each one re-reading the same sources. Caching turns the repeated read from a cost into a rounding error.

Ask an AI search engine a question and you get an answer that reads well. That is not the hard part anymore. The hard part is knowing whether the sentence you are about to act on would survive someone checking it.

Verification is a different skill from summarization. It means noticing that two sources describe the same statistic with different denominators. It means flagging that a claim's only support is a press release. It means returning "the evidence here is thin" instead of a confident paragraph.

Anthropic published four benchmarks on the Opus 5.5 announcement, and the pattern across them is worth reading carefully:

BenchmarkClaude Opus 5.5GPT-6 Astra
Terminal-Bench 4.066.4%57.9%
FrontierCode v1.1 (Main)54.4%53.3%
AutomationBench40.0%41.4%
Terminal-Bench-Science 0.158.7%64.6%

Anthropic leads on two and trails on two, which is a more useful disclosure than a clean sweep would be. Astra keeps the edge on science-heavy terminal work and business-process automation; Opus 5.5 takes the general agentic-coding and frontier-code columns.

The methodology note matters more than the scores. Anthropic evaluated Opus 5.5 with its production safeguards enabled: the same configuration that ships to users, including the routing that sends risky requests to an older model. A safety system that only exists in the eval is not a safety system. The company also reports roughly 85% fewer attempts to bypass containment boundaries than Opus 5, and its strongest alignment result to date.

For search specifically, that configuration is the product. A model that declines to answer is annoying. A model that answers a question it should have declined, fluently and with citations, is worse.

1. Select it in the model picker. Open Felo Search, choose Claude Opus 5.5, and ask in the same box you always use.

2. Give it the whole document set, not a summary. The million-token window is most valuable when you stop pre-compressing. Upload the filings, the papers, the thread, all of it, and ask your question against the full text.

3. Ask for the caveats explicitly. "What would have to be true for this conclusion to be wrong?" and "which of these sources is weakest?" are questions a verification-trained model answers differently from a summarizer.

4. Make it show the disagreement. When sources conflict, ask it to lay out where they agree, where they diverge, and which claim rests on the thinnest evidence. That is the output you actually want from a research pass.

5. Iterate in the same thread. Cache reads are cheap; restarting is not. Keep the challenge, the correction, and the final synthesis in one session so the model still has the original evidence in front of it.

Opus 5.5 landed within a day of three other frontier releases, and Felo Search runs all four.

ModelReleased byPrice per million tokensContextBuilt for
Grok 4.7SpaceXAI$2 in / $6 out500KLong-horizon reasoning and knowledge work
Claude Opus 5.5Anthropic$4 in / $20 out1MAnswers that survive verification
GPT-6 LunaOpenAI$0.10 in / $0.50 out1.05MHigh-volume search at the lowest cost per call
GPT-6 SolOpenAI$2 in / $10 out1.05MAgent loops and repeat calls at mid-tier price

Opus 5.5 is the most expensive of the four on output, and it is the only one where the vendor published its safety configuration alongside its scores. Match it to the questions where being wrong costs more than being slow.

What to Try First​

A claim you intend to repeat. Pick a statistic you are about to put in a document and ask Opus 5.5 to trace it to its origin. The interesting output is not the number. It is the trail.

A decision with two defensible answers. Give it both options and the constraints, then ask which one the evidence actually supports and what would change the answer.

A document you are about to sign. Contracts, policies, and terms of service reward a model that reports what is missing rather than what is present.

A research thread you have already run. Re-run a question you asked a month ago and compare the caveats. That gap is what the last generation was not doing.

Claude Opus 5.5 is live in Felo Search with a one-million-token context window, $4/$20 pricing, and production safeguards enabled. It is 40% cheaper to operate than the model it replaces and it is the configuration Anthropic puts its own name behind.

Open Claude Opus 5.5 in Felo Search →


This post is also available in 简体中文, 日本語, 한국어, 繁體中文, हिन्दी, Français, العربية, Русский, اردو, Bahasa Indonesia, Deutsch, Tiếng Việt, Türkçe, Italiano, ไทย, Español, বাংলা and Português.