Skip to main content

Claude Haiku 5.5 Is Live on Felo OpenAPI

· 9 min read
Felo Search Tips Buddy
Committed to answers at your fingertips

Anthropic's cheapest, fastest model lands on Felo OpenAPI on launch day. $0.10/$0.50 per million tokens, 1M context, adaptive thinking — one API key.

Anthropic shipped Claude Haiku 5.5 on October 7, 2026. You can call it today.

Felo OpenAPI added the claude-haiku-5-5 route on launch day, at Anthropic's list price: $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. That is 90% cheaper than Haiku 4.5, which makes the fastest model in the Claude 5.5 family also the least expensive one to run.

Try Claude Haiku 5.5 on Felo OpenAPI: https://openapi.felo.ai/models/anthropic/claude-haiku-5.5

Claude Haiku 5.5 on Felo OpenAPI: Anthropic's fastest Haiku route, 90% cheaper than Haiku 4.5

Why this release matters​

Every Haiku has been cheap. This is the first one that can run an agent loop without being the weak link.

Anthropic calls Haiku 5.5 the cheapest, fastest, and most capable small model it has ever released. The benchmarks make that claim hard to argue with. Compare it to the model it replaces:

Claude Haiku 4.5Claude Haiku 5.5
Input / output$1 / $5$0.10 / $0.50
Cache read$0.10$0.01
Context / max output200K / 64K tokens1M / 128K tokens
Effort control—5 levels, default medium
OSWorld 2.1 (computer use)15.7%72.4%
Terminal-Bench 4.00.0%39.2%
GDPval-AA v2.17351620
Humanity's Last Exam (no tools)10.2%45.9%

Prices per million tokens, for prompts up to 100,000 tokens. Benchmark figures as reported by Anthropic.

The capability jump is larger than the price cut. On OSWorld 2.1, a computer-use benchmark, Haiku 5.5 scores 72.4% where Haiku 4.5 managed 15.7%. Terminal-Bench 4.0 goes from 0.0% to 39.2%. Work that used to need a mid-tier model at 20 times the price now fits on the Haiku line.

The pricing, explained​

Haiku 5.5 breaks with the rest of the Claude family in one way worth knowing: it is priced by prompt length.

  • Prompts up to 100,000 tokens: $0.10 input, $0.50 output, $0.01 cache read — per million tokens.
  • Prompts over 100,000 tokens: $0.50 input, $2.50 output, $0.05 cache read.

Every other current Claude model bills one flat rate across its full context window. Haiku 5.5 is the exception, so anything you budget in tokens is worth recomputing.

For most workloads the boundary never comes up. Anthropic says 90% of Haiku 4.5 requests sit under 100,000 tokens, so they pay the full 90% cut. And because Haiku 5.5 uses the family's newer tokenizer — the same text counts as roughly 30% more tokens than on Haiku 4.5 — Anthropic puts the average running-cost reduction at about 75% rather than a clean 90%.

The number to underline is that $0.01 cache read. Agent loops spend most of their input budget re-reading context they have already sent; at a cent per million tokens, keeping a long conversation warm costs almost nothing.

Claude Haiku 5.5 pricing: $0.10 per million input tokens up to a 100,000-token prompt, and a 90% drop from Haiku 4.5

What else changed​

Adaptive thinking, now with an effort dial. Haiku 5.5 is the first Haiku that accepts the effort parameter. Thinking adapts to the task by default, and five effort levels — low, medium, high, xhigh, max, with medium as the default — let you trade cost and speed against response quality. At lower effort, the model skips thinking entirely on simple requests.

1M-token context, 128K max output. Five times the context window of Haiku 4.5 and double the output ceiling. One window holds roughly 555,000 words, which is more than most codebases, transcripts, or document sets you would want to reason over in a single call.

A new tokenizer. Haiku 5.5 uses the same tokenizer as Claude 4.7 and later models. Requests and responses keep their shape, but usage counts are higher for the same text, so max_tokens limits and cost estimates tuned on Haiku 4.5 need a fresh pass.

Browser use and firmer refusals. The model supports the browser-use toolset alongside computer use, and its safety classifiers can decline a request outright — plan for a stop_reason of "refusal" in your client.

Five breaking changes before you migrate​

Haiku 5.5 is not a drop-in model ID swap for code running on Haiku 4.5. Anthropic's migration guide lists the changes; these are the five that break existing requests.

  1. Manual thinking budgets are rejected. thinking: {"type": "enabled", "budget_tokens": N} returns a 400. Use adaptive thinking and steer depth with effort. Thinking text is omitted by default now, so set thinking.display to "summarized" if you stream it.
  2. Sampling parameters are rejected. Omit temperature, top_p, and top_k entirely. Any non-default value returns a 400 — including top_p: 1.
  3. Assistant prefill is rejected. End messages with a user turn. For structured output, use structured outputs or classification tools with enum fields instead of prefill tricks.
  4. Computer use moved to a toolset. Replace computer_20250124 with computer_toolset_20260801 on the Claude API and Google Cloud.
  5. Thinking blocks are account-bound and append-only. A thinking block stays valid only in the account that produced it and only while everything before it is unchanged. Keep conversations append-only when you send thinking blocks back.

One more change is silent: thinking tokens count toward max_tokens, and a response can begin with a thinking block before any text. Select content blocks by their type field, not by position, and give max_tokens room to think.

The full migration guide is on Anthropic's docs.

Calling Haiku 5.5 on Felo OpenAPI​

The route is live. Here is the whole integration.

curl https://openapi.felo.ai/api/v1/chat/completions \
-H "Authorization: Bearer $FELO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-haiku-5-5",
"messages": [
{"role": "user", "content": "Classify this support ticket and extract the order ID."}
]
}'

Felo OpenAPI exposes three compatible surfaces, so you can keep the client you already use:

SurfaceMethodURL
Chat CompletionsPOSThttps://openapi.felo.ai/api/v1/chat/completions
ResponsesPOSThttps://openapi.felo.ai/api/v1/responses
Messages (Anthropic-compatible)POSThttps://openapi.felo.ai/api/v1/messages

Because the Messages endpoint is Anthropic-compatible, Claude-style clients only need the base URL pointed at https://openapi.felo.ai/api. Third-party SDKs work the same way — set baseURL to https://openapi.felo.ai/api/v1 and keep your existing code.

Add "stream": true for server-sent events. The route supports reasoning, tool calling, JSON output, streaming, and vision — the same surface as the rest of the catalog.

If you already pay for Felo Pro, you already have access​

Felo Pro includes API access to Felo OpenAPI models. If you subscribe, you do not need a second account to call Haiku 5.5: create an API key, point your agent at the endpoint, and your existing Felo credits cover the calls. The same key works across the lineup — Haiku 5.5, Sonnet 5.5, Opus 5.5, GPT-6, Grok 4.7, and the rest.

Two things to keep straight:

  • Felo Pro credits are one pool. Search, research agents, slides, images, and API calls all draw from the same balance. Model calls bill by token, so heavy API use moves through credits faster than a few searches a day.
  • Tool APIs are separate. Model calls are covered by your Pro access; harness tools like Web Fetch, X Search, and PPT generation run on Felo API Platform credits.

Not on Pro yet? Every Felo account gets 200 free credits per day, enough to run real requests against Haiku 5.5 before you commit.

Where Haiku 5.5 fits​

Where Claude Haiku 5.5 fits: high-volume classification, extraction, and subagent work between Sonnet 5.5 and Opus 5.5

Haiku 5.5 is the route to pick when the work is simple but the volume is not.

  • Subagents and compaction — Anthropic points at summarization, compaction, and subagent work directly. Run Opus 5.5 or Sonnet 5.5 on the hard reasoning and let Haiku 5.5 handle the loop around it at a tenth of the price.
  • High-volume classification, routing, and extraction — the per-token cost decides whether these features ship at all. At $0.10 per million input tokens and $0.50 per million output tokens, the marginal cost of one more classification stops being a budget question.
  • Support bots and latency-sensitive UX — the fastest route in the lineup, now with an effort dial for trimming the last milliseconds.
  • Long-document work — a 1M window at Haiku prices means whole transcripts, contract sets, or log files fit in one call without retrieval plumbing.

If you need the absolute ceiling on hard reasoning, Opus 5.5 is still the flagship and it is on Felo OpenAPI too. Sonnet 5.5 remains the balance point for most production workloads. Haiku 5.5 is the model you put everywhere the budget used to say no.

Try it now​

Haiku 5.5 is live on Felo OpenAPI today, at Anthropic's list price, behind the key you already have.

Start in the Playground to test prompts without writing code, then grab an API key and point your agent at the endpoint — subagents, classifiers, and support bots are one model ID away from costing 90% less.

→ Try Claude Haiku 5.5 on Felo OpenAPI

→ Create your API key


This post is also available in 简体中文, 日本語, 한국어, 繁體中文, हिन्दी, Français, العربية, Русский, اردو, Bahasa Indonesia, Deutsch, Tiếng Việt, Türkçe, Italiano, ไทย, Español, বাংলা and Português.