AI/News
Anthropic's Claude Haiku 5.5 is priced 90% below Haiku 4.5 on most requests - $0.10 per million tokens in - and Sonnet 5.5's cache reads are halved the same day
Anthropic's own page: "around 75% less to run" than Haiku 4.5 on average, $0.10 in and $0.50 out per million tokens on prompts up to 100,000 tokens; OSWorld 2.1 at 72.4% against Haiku 4.5's 15.7% and GPT-6 Luna's 48.9%; Terminal-Bench 4.0 at 39.2% against 0.0%. A $100 to $500 monthly API credit for Max and Team subscribers starts this week.
By Daily Aletheia · Checked against the primary source · 8 October 2026 · 3 min read

Anthropic released Claude Haiku 5.5 on 7 October, calling it "the cheapest, fastest, and most capable small model we've ever released". It is "designed for high-volume, cost-sensitive tasks" - "summaries, compactions, database queries, and classification requests" - and "pairs well with Opus 5.5 and Sonnet 5.5 as a subagent on coding work".
The price. Per million tokens, on prompts up to 100,000 tokens: cache reads $0.01, cache writes $0.125, input $0.10, output $0.50. Over 100,000 tokens: $0.05, $0.625, $0.50 and $2.50. Haiku 4.5 was $1.00 in and $5.00 out. The page's own arithmetic, in its footnote: "priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens", and since "90% of requests fell into the former category", it "costs around 75% less to run" on average - after allowing for a new tokenizer that "uses slightly more tokens per task".
The numbers against the model it replaces, and against OpenAI's small model. OSWorld 2.1 (offline subset): 72.4%, against Haiku 4.5's 15.7% and GPT-6 Luna's 48.9%. Terminal-Bench 4.0: 39.2% against 0.0% and 16.4%. Humanity's Last Exam without tools: 45.9% against 10.2%. FrontierCode 1.1: 46.4% against GPT-6 Luna's 42.4%. GDPval-AA v2.1: 1620 against 735 and 1437. Sonnet 5.5 sits above it on every row - 83.9% on OSWorld, 70.6% on Terminal-Bench - and the page says so: "Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding tasks". It is also the first Haiku "with an adjustable effort setting".
What customers said on the page. HubSpot: "the best score we've seen on this suite yet, at 92.8% averaged over three runs". Asana: "over a 30% reduction in latency for task completions and up to 2.5x faster inference per agent turn". AlphaSense, on "about 8M calls a week": "0.84 vs. 0.76" against Haiku 4.5 on 400 queries.
Two more changes the same day. Sonnet 5.5's cache reads "now cost 50% less: $0.10 per million tokens rather than $0.20", which the page says cuts Sonnet's cost "on most agentic tasks by around 20%". And "this week" a monthly API credit arrives for subscribers: "Max 5x users will get $100 in credits per month, Max 20x users will get $200, and Team subscribers will receive up to $500, pooled across their users", usable "on any of our models".
Safeguards. Its cyber safeguards "are more restrictive than Haiku 4.5's, but somewhat less restrictive than those we've applied to other recent models"; they "still block penetration testing". Biology safeguards "are the same as for Sonnet 5, Sonnet 5.5, and Opus 5".
Where. It is "available now on all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure", as claude-haiku-5-5. The Python and TypeScript SDKs gain computer use and browser use "in beta".
What the page does not say. No context-window length, no tokens-per-second figure, and "fastest" is defined in a footnote as "at each model's standard speed" - the Opus models in Fast Mode are quicker. For anyone running a high-volume pipeline on Haiku 4.5, the move is a price cut of the order of three-quarters on the same bill, before any change in what the model can do.