Claude Haiku 5.5 Is Live on Felo OpenAPI
Anthropic's cheapest, fastest model lands on Felo OpenAPI on launch day. $0.10/$0.50 per million tokens, 1M context, adaptive thinking — one API key.
Anthropic shipped Claude Haiku 5.5 on October 7, 2026. You can call it today.
Felo OpenAPI added the claude-haiku-5-5 route on launch day, at Anthropic's list price: $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. That is 90% cheaper than Haiku 4.5, which makes the fastest model in the Claude 5.5 family also the least expensive one to run.
Try Claude Haiku 5.5 on Felo OpenAPI: https://openapi.felo.ai/models/anthropic/claude-haiku-5.5

Why this release matters
Every Haiku has been cheap. This is the first one that can run an agent loop without being the weak link.
Anthropic calls Haiku 5.5 the cheapest, fastest, and most capable small model it has ever released. The benchmarks make that claim hard to argue with. Compare it to the model it replaces:
| Claude Haiku 4.5 | Claude Haiku 5.5 | |
|---|---|---|
| Input / output | $1 / $5 | $0.10 / $0.50 |
| Cache read | $0.10 | $0.01 |
| Context / max output | 200K / 64K tokens | 1M / 128K tokens |
| Effort control | — | 5 levels, default medium |
| OSWorld 2.1 (computer use) | 15.7% | 72.4% |
| Terminal-Bench 4.0 | 0.0% | 39.2% |
| GDPval-AA v2.1 | 735 | 1620 |
| Humanity's Last Exam (no tools) | 10.2% | 45.9% |
Prices per million tokens, for prompts up to 100,000 tokens. Benchmark figures as reported by Anthropic.
The capability jump is larger than the price cut. On OSWorld 2.1, a computer-use benchmark, Haiku 5.5 scores 72.4% where Haiku 4.5 managed 15.7%. Terminal-Bench 4.0 goes from 0.0% to 39.2%. Work that used to need a mid-tier model at 20 times the price now fits on the Haiku line.
The pricing, explained
Haiku 5.5 breaks with the rest of the Claude family in one way worth knowing: it is priced by prompt length.
- Prompts up to 100,000 tokens: $0.10 input, $0.50 output, $0.01 cache read — per million tokens.
- Prompts over 100,000 tokens: $0.50 input, $2.50 output, $0.05 cache read.
Every other current Claude model bills one flat rate across its full context window. Haiku 5.5 is the exception, so anything you budget in tokens is worth recomputing.
For most workloads the boundary never comes up. Anthropic says 90% of Haiku 4.5 requests sit under 100,000 tokens, so they pay the full 90% cut. And because Haiku 5.5 uses the family's newer tokenizer — the same text counts as roughly 30% more tokens than on Haiku 4.5 — Anthropic puts the average running-cost reduction at about 75% rather than a clean 90%.
The number to underline is that $0.01 cache read. Agent loops spend most of their input budget re-reading context they have already sent; at a cent per million tokens, keeping a long conversation warm costs almost nothing.

What else changed
Adaptive thinking, now with an effort dial. Haiku 5.5 is the first Haiku that accepts the effort parameter. Thinking adapts to the task by default, and five effort levels — low, medium, high, xhigh, max, with medium as the default — let you trade cost and speed against response quality. At lower effort, the model skips thinking entirely on simple requests.
1M-token context, 128K max output. Five times the context window of Haiku 4.5 and double the output ceiling. One window holds roughly 555,000 words, which is more than most codebases, transcripts, or document sets you would want to reason over in a single call.
A new tokenizer. Haiku 5.5 uses the same tokenizer as Claude 4.7 and later models. Requests and responses keep their shape, but usage counts are higher for the same text, so max_tokens limits and cost estimates tuned on Haiku 4.5 need a fresh pass.
Browser use and firmer refusals. The model supports the browser-use toolset alongside computer use, and its safety classifiers can decline a request outright — plan for a stop_reason of "refusal" in your client.
Five breaking changes before you migrate
Haiku 5.5 is not a drop-in model ID swap for code running on Haiku 4.5. Anthropic's migration guide lists the changes; these are the five that break existing requests.
- Manual thinking budgets are rejected.
thinking: {"type": "enabled", "budget_tokens": N}returns a 400. Use adaptive thinking and steer depth witheffort. Thinking text is omitted by default now, so setthinking.displayto"summarized"if you stream it. - Sampling parameters are rejected. Omit
temperature,top_p, andtop_kentirely. Any non-default value returns a 400 — includingtop_p: 1. - Assistant prefill is rejected. End
messageswith a user turn. For structured output, use structured outputs or classification tools with enum fields instead of prefill tricks. - Computer use moved to a toolset. Replace
computer_20250124withcomputer_toolset_20260801on the Claude API and Google Cloud. - Thinking blocks are account-bound and append-only. A thinking block stays valid only in the account that produced it and only while everything before it is unchanged. Keep conversations append-only when you send thinking blocks back.
One more change is silent: thinking tokens count toward max_tokens, and a response can begin with a thinking block before any text. Select content blocks by their type field, not by position, and give max_tokens room to think.
The full migration guide is on Anthropic's docs.
Calling Haiku 5.5 on Felo OpenAPI
The route is live. Here is the whole integration.
curl https://openapi.felo.ai/api/v1/chat/completions \
-H "Authorization: Bearer $FELO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-haiku-5-5",
"messages": [
{"role": "user", "content": "Classify this support ticket and extract the order ID."}
]
}'
Felo OpenAPI exposes three compatible surfaces, so you can keep the client you already use:
| Surface | Method | URL |
|---|---|---|
| Chat Completions | POST | https://openapi.felo.ai/api/v1/chat/completions |
| Responses | POST | https://openapi.felo.ai/api/v1/responses |
| Messages (Anthropic-compatible) | POST | https://openapi.felo.ai/api/v1/messages |
Because the Messages endpoint is Anthropic-compatible, Claude-style clients only need the base URL pointed at https://openapi.felo.ai/api. Third-party SDKs work the same way — set baseURL to https://openapi.felo.ai/api/v1 and keep your existing code.
Add "stream": true for server-sent events. The route supports reasoning, tool calling, JSON output, streaming, and vision — the same surface as the rest of the catalog.
If you already pay for Felo Pro, you already have access
Felo Pro includes API access to Felo OpenAPI models. If you subscribe, you do not need a second account to call Haiku 5.5: create an API key, point your agent at the endpoint, and your existing Felo credits cover the calls. The same key works across the lineup — Haiku 5.5, Sonnet 5.5, Opus 5.5, GPT-6, Grok 4.7, and the rest.
Two things to keep straight:
- Felo Pro credits are one pool. Search, research agents, slides, images, and API calls all draw from the same balance. Model calls bill by token, so heavy API use moves through credits faster than a few searches a day.
- Tool APIs are separate. Model calls are covered by your Pro access; harness tools like Web Fetch, X Search, and PPT generation run on Felo API Platform credits.
Not on Pro yet? Every Felo account gets 200 free credits per day, enough to run real requests against Haiku 5.5 before you commit.
Where Haiku 5.5 fits

Haiku 5.5 is the route to pick when the work is simple but the volume is not.
- Subagents and compaction — Anthropic points at summarization, compaction, and subagent work directly. Run Opus 5.5 or Sonnet 5.5 on the hard reasoning and let Haiku 5.5 handle the loop around it at a tenth of the price.
- High-volume classification, routing, and extraction — the per-token cost decides whether these features ship at all. At $0.10 per million input tokens and $0.50 per million output tokens, the marginal cost of one more classification stops being a budget question.
- Support bots and latency-sensitive UX — the fastest route in the lineup, now with an effort dial for trimming the last milliseconds.
- Long-document work — a 1M window at Haiku prices means whole transcripts, contract sets, or log files fit in one call without retrieval plumbing.
If you need the absolute ceiling on hard reasoning, Opus 5.5 is still the flagship and it is on Felo OpenAPI too. Sonnet 5.5 remains the balance point for most production workloads. Haiku 5.5 is the model you put everywhere the budget used to say no.
Try it now
Haiku 5.5 is live on Felo OpenAPI today, at Anthropic's list price, behind the key you already have.
Start in the Playground to test prompts without writing code, then grab an API key and point your agent at the endpoint — subagents, classifiers, and support bots are one model ID away from costing 90% less.
→ Try Claude Haiku 5.5 on Felo OpenAPI
This post is also available in 简体中文, 日本語, 한국어, 繁體中文, हिन्दी, Français, العربية, Русский, اردو, Bahasa Indonesia, Deutsch, Tiếng Việt, Türkçe, Italiano, ไทย, Español, বাংলা and Português.