# Claude Haiku 5.5 cuts short-prompt token prices 90% and brings 1M context to Anthropic's small tier

Anthropic's October 7 Haiku 5.5 launch sets API pricing at $0.10 input and $0.50 output per million tokens for prompts up to 100,000 tokens, with a 1M context window and adaptive thinking; long prompts use a higher price tier.

Haiku 5.5 resets the economics of high-volume classification, extraction and agent sub-tasks, while Anthropic also cuts Sonnet 5.5 cache-read prices and introduces API credits for Max/Team subscribers.

- Status: Active
- Published: 2026-10-08T17:56:12+13:00
- Updated: 2026-10-08T17:56:12+13:00
- Categories: Artificial Intelligence, SaaS, AI Models, Pricing & Billing, Inference & APIs
- Tags: Anthropic, API pricing, Claude Haiku 5.5, context windows, model routing
- Canonical HTML: https://beyondthe.news/dossiers/claude-haiku-5-5-pricing-context-adaptive-effort

## What changed

On October 7, 2026 Anthropic released Claude Haiku 5.5 (claude-haiku-5-5) on its API and major cloud platforms. The model has a 1M-token context window, 128K maximum output and adaptive thinking with an effort parameter. For prompts up to 100,000 tokens, API pricing is $0.10 per million input and $0.50 per million output tokens, versus Haiku 4.5's $1/$5; above 100,000 prompt tokens, rates rise to $0.50/$2.50. Anthropic estimates about 75% lower average cost per task than Haiku 4.5 after accounting for tokenization differences. It also cut Sonnet 5.5 cache-read prices in half and announced monthly API credits for Max and Team subscribers.

## Why it matters

High-volume AI products spend substantial money on short repetitive calls and delegated agent tasks. A 90% token-rate reduction at short prompts can change the economics of classification, extraction, summarization and support workloads, but a 1M context window is not a promise of flat-rate processing. The 100K pricing boundary and newer tokenizer matter to actual cost per completed task. Developers should benchmark quality, latency and token usage under their own data rather than treat vendor comparisons as universal.

## Short prompts get a steep price cut; long prompts have a different tier

Up to 100K prompt tokens, input/output rates are $0.10/$0.50 per million; above 100K, they are $0.50/$2.50. Cache writes/reads and batch discounts have their own rates. Anthropic says the typical task costs around 75% less than on Haiku 4.5, not 90%, because output/token consumption can differ.

## Context and effort controls move downmarket

Haiku 5.5 offers a 1M-token context window, up to 128K output, and adaptive thinking with adjustable effort. The API model ID is claude-haiku-5-5.

## The rest of the Claude stack also changes price

Anthropic cut Sonnet 5.5 cache-read cost from $0.20 to $0.10 per million tokens and says it lowers typical agentic cost about 20%. It also announced monthly API credits for Max/Team users, with rollout and plan-specific limits.

## Vendor benchmark comparisons need independent evaluation

Anthropic reports substantial gains over Haiku 4.5 and publishes early customer examples, but model choice for production agents depends on task success rate, retries, latency and risk rather than one benchmark score.

## Key details

- Claude Haiku 5.5 launched October 7, 2026.
- Model ID claude-haiku-5-5; context 1M tokens and max output 128K.
- Prompts <=100K tokens: $0.10 input / $0.50 output per million tokens.
- Prompts >100K tokens: $0.50 input / $2.50 output per million tokens.
- Anthropic estimates 75% lower average cost per task than Haiku 4.5, accounting for changed tokenization.
- Adaptive thinking is on by default and configurable via effort.
- Sonnet 5.5 cache reads are 50% cheaper, and Max/Team API credits were announced.
- Available through Anthropic's API, Amazon Bedrock, Google Cloud and Microsoft Foundry.

## Builder takeaways

- Segment usage by prompt length to avoid pricing long-context calls as if they were short.
- Benchmark cost per accepted result, including changed token counts, retries and cache hits.
- Test Haiku 5.5 as a narrow-task or subagent model before swapping it into high-risk orchestration.
- Update model routing and budgets for the separate 100K-token rate boundary.

## What to watch

- Independent quality and throughput measurements under realistic short-task workloads.
- How quickly Max/Team API credits become usable and whether terms change.
- Long-context cost and latency in production integrations.

## Uncertainties

- Anthropic's quality and 75% average cost reduction are vendor-produced measurements.
- API credit availability may roll out after the announcement date.
- Actual cost depends on prompt length, caching, tokenization and completion behaviour.

## Sources

- [Claude Haiku 5.5](https://www.anthropic.com/claude-haiku-5-5) — Anthropic · primary · 2026-10-07T00:00:00+13:00. Primary launch, pricing tiers, Sonnet cache change and API credit announcement.
- [Claude Haiku 5.5 model overview](https://platform.claude.com/docs/en/models/haiku-5-5/overview) — Anthropic Platform · primary documentation · 2026-10-07T00:00:00+13:00. Current model ID, context/output, tiered rates, availability and API caveats.
- [Anthropic launches third Claude 5.5 model](https://www.reuters.com/business/anthropic-launches-third-claude-55-model-expanding-ai-lineup-before-planned-ipo-2026-10-07/) — Reuters · independent reporting · 2026-10-07T00:00:00+13:00. Independent launch and price verification.

