Claude Haiku 5.5: Pricing, Features, Benchmarks, API & Is It Worth It?
Claude Haiku 5.5: Pricing, Features, Benchmarks, API & Is It Worth It?
Anthropic has released Claude Haiku 5.5, its newest low-cost AI model aimed at developers and businesses that need fast AI at high volume.
Released on:
October 7, 2026
Claude Haiku 5.5 combines:
- a 1 million-token context window;
- up to 128,000 standard output tokens;
- adaptive reasoning;
- image understanding;
- very low API pricing;
- prompt caching;
- Batch API discounts;
- computer-use and browser-oriented workloads; and
- support across Anthropic, AWS, Google Cloud and Microsoft.
Its headline API price is particularly aggressive.
For prompts up to 100,000 tokens, Anthropic charges:
$0.10 per million input tokens
and:
$0.50 per million output tokens
That places Haiku 5.5 directly against efficient models such as OpenAI’s GPT-6 Luna.
However, there is one major pricing detail developers need to understand:
Claude Haiku 5.5 costs five times more per token when the prompt exceeds 100,000 tokens.
So while the model has a 1M-token context window, using that context can materially change its economics.
Here is everything you need to know about Claude Haiku 5.5.
Claude Haiku 5.5 at a Glance
| Feature | Claude Haiku 5.5 |
|---|---|
| Developer | Anthropic |
| Release date | October 7, 2026 |
| API model ID | claude-haiku-5-5 |
| Context window | 1,000,000 tokens |
| Standard max output | 128,000 tokens |
| Batch API max output | 300,000 tokens beta |
| Input price ≤100K prompt | $0.10 / 1M tokens |
| Output price ≤100K prompt | $0.50 / 1M tokens |
| Input price >100K prompt | $0.50 / 1M tokens |
| Output price >100K prompt | $2.50 / 1M tokens |
| Cache read ≤100K | $0.01 / 1M |
| Batch discount | 50% input/output |
| Reasoning | Adaptive |
| Default effort | Medium |
| Input | Text + images |
| Output | Text |
| Reliable knowledge cutoff | June 2026 |
| Status | Active |
Anthropic positions Haiku 5.5 as its:
fastest and cheapest current Claude model
for high-volume, latency-sensitive work.
Official documentation:
Anthropic — Claude Haiku 5.5 Model Documentation
What Is Claude Haiku 5.5?
Claude Haiku 5.5 is Anthropic’s efficiency-focused model in the Claude 5.5 family.
Anthropic currently positions its major models roughly like this:
Claude Opus 5.5 → maximum capability
Claude Sonnet 5.5 → strong balance of capability and cost
Claude Haiku 5.5 → speed, scale and low cost
Haiku is therefore not primarily designed to replace Opus for the most difficult reasoning or autonomous coding tasks.
Instead, Anthropic recommends it for workloads such as:
- summarization;
- classification;
- extraction;
- routing;
- document queries;
- customer support;
- database queries;
- compaction;
- subagents; and
- other repetitive AI operations.
This makes Haiku 5.5 particularly interesting for businesses processing thousands or millions of requests.
For a broader comparison of frontier models, see our Best AI Models in 2026 guide.
Claude Haiku 5.5 Pricing
Anthropic introduced a two-tier pricing structure for Haiku 5.5.
Prompts up to 100,000 tokens
Input:
$0.10 per million tokens
Output:
$0.50 per million tokens
Cache reads:
$0.01 per million tokens
5-minute cache writes:
$0.125 per million tokens
1-hour cache writes:
$0.20 per million tokens
Prompts over 100,000 tokens
Input:
$0.50 per million tokens
Output:
$2.50 per million tokens
Cache reads:
$0.05 per million tokens
5-minute cache writes:
$0.625 per million tokens
1-hour cache writes:
$1 per million tokens
So crossing the:
100,000-token prompt threshold
increases base input and output pricing by:
5×
That is critical when calculating the cost of long-document workflows.
Claude Haiku 5.5 Pricing Table
| Usage | Prompt ≤100K | Prompt >100K |
|---|---|---|
| Input | $0.10/M | $0.50/M |
| Output | $0.50/M | $2.50/M |
| Cache read | $0.01/M | $0.05/M |
| 5-min cache write | $0.125/M | $0.625/M |
| 1-hour cache write | $0.20/M | $1.00/M |
Anthropic also offers a:
50% Batch API discount
on input and output tokens.
Official pricing:
Anthropic — Claude Haiku 5.5 Pricing
Claude Haiku 5.5 Cost Example
Suppose an application processes:
100,000 input tokens
and generates:
10,000 output tokens
while remaining in the lower pricing tier.
Input
100,000 tokens = 0.1 million
0.1 × $0.10 = $0.01
Output
10,000 tokens = 0.01 million
0.01 × $0.50 = $0.005
Total:
$0.015
That is:
1.5 cents
for this simplified example.
This excludes caching, tools and other possible charges.
What Happens Above 100K Tokens?
Suppose a request instead contains:
200,000 input tokens
and generates:
10,000 output tokens
Because the prompt is over 100K, the higher pricing tier applies.
Input
0.2 × $0.50 = $0.10
Output
0.01 × $2.50 = $0.025
Total:
$0.125
So context length can make a major difference to the economics of Haiku 5.5.
Developers should not assume the advertised:
$0.10 / $0.50
rate applies to every request across the model’s 1M context window.
Is Claude Haiku 5.5 Cheaper Than Haiku 4.5?
Yes.
Anthropic says Haiku 5.5 costs approximately:
75% less on average
to run than Haiku 4.5.
For prompts up to 100K tokens, the raw price difference is even larger.
Claude Haiku 4.5 was priced at:
$1 per million input tokens
and:
$5 per million output tokens
Haiku 5.5 costs:
$0.10 input
and:
$0.50 output
for prompts up to 100K.
That is a:
90% reduction in base token pricing
in that tier.
For longer prompts, Anthropic says Haiku 5.5’s pricing remains roughly:
50% lower
than Haiku 4.5.
One Important Tokenizer Detail
There is a complication when comparing raw token prices.
Anthropic says Claude Haiku 5.5 uses its newer tokenizer.
The same text can produce approximately:
30% more tokens
than on Claude Haiku 4.5.
The exact difference depends on content.
That means developers should benchmark:
actual cost per task
rather than comparing only advertised token prices.
Even with this tokenizer difference, Anthropic estimates Haiku 5.5 costs roughly 75% less on average than its predecessor.
Claude Haiku 5.5 Context Window
Claude Haiku 5.5 supports:
1,000,000 tokens of context
by default.
Anthropic says no special beta header is required to use the full context capacity.
This is an enormous context window for a model at Haiku’s price point.
Potential use cases include:
- long reports;
- large document collections;
- code repositories;
- business records;
- legal documents;
- financial reports;
- support histories;
- research collections;
- knowledge bases; and
- long-running agent context.
However, remember:
prompts above 100K tokens use the higher Haiku pricing tier.
So the model’s 1M context window is technically powerful but should still be used selectively.
Claude Haiku 5.5 Maximum Output
Anthropic lists the normal maximum output as:
128,000 tokens
That is comparable to larger Claude models.
Haiku 5.5 also supports:
up to 300,000 output tokens
through the Message Batches API beta.
The beta capability requires Anthropic’s applicable feature configuration.
Most applications will never need a 128K response, let alone 300K.
But the capability could matter for:
- bulk structured extraction;
- extensive document transformations;
- code generation;
- large reports; and
- long offline workflows.
Adaptive Thinking Comes to Haiku
Haiku 5.5 is the first Haiku-class model to support Anthropic’s:
adaptive thinking
with adjustable effort.
The default effort setting is:
medium
Developers can tell the model to spend more or less effort depending on the task.
That gives applications a useful trade-off between:
- intelligence;
- latency; and
- cost.
For routine classification, lower reasoning effort may be enough.
For more difficult document analysis or coding, higher effort can improve performance.
Claude Haiku 5.5 Benchmarks
Anthropic published several benchmark comparisons at launch.
Provider-published benchmark results should always be treated as evidence from the model developer rather than as completely neutral testing.
With that qualification, Anthropic reports:
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1 | 1620 | 735 | 1437 | 1840 |
| AA-Briefcase v1.1 | 1578 | 614 | 1336 | 1824 |
| OSWorld 2.1 offline | 72.4% | 15.7% | 48.9% | 83.9% |
| Humanity’s Last Exam, no tools | 45.9% | 10.2% | — | 56.9% |
| Terminal-Bench 4.0 | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1 | 46.4% | — | 42.4% | 52.1% |
| Chartography | 46.4% | 6.4% | 29.1% | 61.6% |
Source:
Anthropic — Introducing Claude Haiku 5.5
The results suggest Haiku 5.5 is a major capability jump over Haiku 4.5.
But Anthropic itself says:
Sonnet and Opus remain better choices for complex agentic coding.
Independent Benchmark Results
Artificial Analysis has also tested Claude Haiku 5.5 independently.
At its maximum tested reasoning setting, Artificial Analysis currently reports an:
Intelligence Index score of 43
for Haiku 5.5.
Other effort settings currently score approximately:
Medium: 34
High: 38
Xhigh: 41
Max: 43
Artificial Analysis also reports very high generation speed for Haiku 5.5, depending on the reasoning setting.
Independent results:
Artificial Analysis — Claude Haiku 5.5
The main lesson is that effort setting matters.
You cannot describe Haiku 5.5 with one intelligence number without also specifying how much reasoning effort the model is using.
Claude Haiku 5.5 vs GPT-6 Luna
This may be one of the most interesting low-cost AI comparisons of 2026.
Both models start at:
$0.10 per million input tokens
and:
$0.50 per million output tokens
for their lowest standard pricing tiers.
Both also offer approximately:
1 million tokens of context
and:
128K standard output
But their pricing rules differ.
| Feature | Claude Haiku 5.5 | GPT-6 Luna |
|---|---|---|
| Developer | Anthropic | OpenAI |
| Short-context input | $0.10/M | $0.10/M |
| Short-context output | $0.50/M | $0.50/M |
| Context | 1M | 1.05M |
| Max output | 128K | 128K |
| Image input | Yes | Yes |
| Reasoning | Adaptive | Yes |
| Long-context price increase | 5× above 100K prompt | Uses OpenAI long-context pricing rules |
| Main positioning | High-volume, fast Claude | Efficient high-volume GPT-6 |
Anthropic’s own benchmark table places Haiku 5.5 ahead of GPT-6 Luna on several tests, including:
- GDPval-AA;
- AA-Briefcase;
- OSWorld;
- Terminal-Bench;
- FrontierCode; and
- Chartography.
But because Anthropic published that comparison, developers should also test the models on their own workload.
For full Luna specifications, see our GPT-6 Luna pricing and features guide.
Claude Haiku 5.5 vs Claude Sonnet 5.5
Haiku is dramatically cheaper.
Claude Sonnet 5.5 currently costs:
$2 per million input tokens
and:
$10 per million output tokens
Haiku 5.5 begins at:
$0.10
and:
$0.50
respectively.
That means Sonnet’s base token rates are:
20× higher
than Haiku’s lowest tier.
But Sonnet is also substantially more capable on difficult agentic tasks.
Anthropic’s Terminal-Bench comparison shows:
Haiku 5.5: 39.2%
Sonnet 5.5: 70.6%
So the right choice depends on the workload.
Use Haiku when:
volume + speed + low cost
matter most.
Use Sonnet when:
complex task completion
matters more than raw token price.
Claude Haiku 5.5 vs Claude Opus 5.5
The difference is even clearer with Opus.
Anthropic lists Opus 5.5 at:
$4/M input
and:
$20/M output
Haiku starts at:
$0.10/M input
and:
$0.50/M output
That makes Opus:
40× more expensive per base token
at Haiku’s lowest tier.
But Opus is intended for the hardest professional and long-running agent tasks.
A useful architecture may be:
Haiku → routine subtask
Sonnet → difficult workflow
Opus → hardest reasoning / coding
This can keep AI-agent costs manageable.
Haiku 5.5 Is Designed for Subagents
One of Anthropic’s most interesting recommendations is to use Haiku 5.5 as a:
subagent
alongside larger Claude models.
Imagine an AI system where Opus manages the overall task.
Rather than asking Opus to perform every smaller operation, the system could delegate tasks such as:
- document lookup;
- classification;
- extraction;
- summarization;
- retrieval;
- data cleanup; and
- simple coding steps
to Haiku.
This architecture can potentially reduce both:
cost
and:
latency
while reserving larger models for difficult decisions.
For more examples of agentic workflows, see our AI Agents for Small Business guide.
Claude Haiku 5.5 for Customer Support
Customer support is one of the most natural Haiku use cases.
Support systems often need to process large numbers of short interactions.
Tasks can include:
- identifying customer intent;
- routing tickets;
- searching knowledge bases;
- summarizing conversations;
- drafting responses;
- classifying urgency; and
- deciding when to escalate to a human.
Anthropic describes Haiku 5.5 as especially well suited to:
live customer support
because of its speed and low cost.
Businesses should still test hallucination rates and use appropriate human escalation for consequential decisions.
Claude Haiku 5.5 for Document Processing
Another strong use case is document analysis at scale.
Possible workloads include:
- invoice extraction;
- financial report summaries;
- contract classification;
- document routing;
- form processing;
- record tagging;
- structured data extraction;
- document Q&A.
Haiku’s:
1M context window
allows very large documents or collections to fit into the same workflow.
But applications should monitor the 100K-token pricing threshold.
Is Claude Haiku 5.5 Good for Coding?
Yes—but with an important qualification.
Haiku 5.5 is dramatically more capable at coding than Haiku 4.5.
Anthropic reports:
46.4% on FrontierCode 1.1
and:
39.2% on Terminal-Bench 4.0
However, Anthropic explicitly states that:
Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding.
Haiku is better suited to coding subtasks such as:
- simple fixes;
- code classification;
- repository search;
- test generation;
- formatting;
- documentation;
- narrow subagent jobs.
For a broader coding comparison across leading frontier models, start with our Best AI Models in 2026 hub.
Is Claude Haiku 5.5 Good for Computer Use?
Anthropic reports a major improvement here.
On its OSWorld 2.1 offline-subset evaluation:
Haiku 5.5: 72.4%
versus:
Haiku 4.5: 15.7%
Anthropic is also updating its Python and TypeScript SDKs with:
computer-use and browser-use support in beta
and specifically says Haiku 5.5 is well suited to these workflows because of its combination of:
speed + price + capability
That could make Haiku especially relevant for high-volume browser agents.
Claude Haiku 5.5 API Model ID
The Claude API model ID is:
claude-haiku-5-5
Anthropic says it is a:
fixed model ID
with no date suffix.
Platform identifiers include:
Claude API
claude-haiku-5-5
Amazon Bedrock
anthropic.claude-haiku-5-5
Google Cloud
claude-haiku-5-5
Microsoft Foundry
claude-haiku-5-5
Anthropic also makes the model available through Claude Platform on AWS.
Claude Haiku 5.5 Availability
Claude Haiku 5.5 is available now.
Anthropic lists support through:
- Claude.ai;
- Claude API;
- Claude Code;
- Amazon Web Services;
- Google Cloud;
- Microsoft Foundry; and
- Claude Platform on AWS.
Anthropic says Free, Pro, Max, Team and Enterprise Claude users can select Haiku 5.5 where available.
The model is currently listed as:
Active — Latest
Anthropic says retirement will occur:
not sooner than October 7, 2027
Claude Haiku 5.5 in Claude Code
Haiku 5.5 is also available in:
Claude Code
This is especially interesting for systems that combine:
Opus or Sonnet as the main coding model
with:
Haiku as a sidekick or subagent
for cheaper supporting work.
That model-routing architecture could become increasingly common as AI-agent systems grow more sophisticated.
Prompt Caching Can Make Haiku Even Cheaper
Haiku 5.5 supports prompt caching.
For prompts up to 100K, cache reads cost:
$0.01 per million tokens
That is a:
90% reduction
compared with ordinary $0.10 input pricing.
Prompt caching can be particularly useful when an application repeatedly sends the same:
- system prompt;
- product documentation;
- policy manual;
- code context; or
- large reference material.
Instead of paying the full input rate each time, the cached content can be reused at a much lower price.
Batch API Discount
Anthropic also offers:
50% off input and output pricing
through the Batch API.
That can make Haiku 5.5 especially inexpensive for offline work such as:
- large classification jobs;
- bulk extraction;
- dataset labeling;
- document summarization;
- evaluation runs;
- content processing.
If a response does not need to arrive immediately, Batch can materially reduce cost.
Haiku 5.5 Safety
Anthropic says Haiku 5.5 performed better than Haiku 4.5 across most of its alignment evaluations.
The company reports:
- fewer examples of misaligned behavior;
- lower willingness to assist misuse;
- stronger cyber safeguards than Haiku 4.5; and
- biology safeguards comparable to several larger Claude models.
Anthropic says the model permits many defensive cybersecurity tasks but restricts penetration testing and other higher-risk techniques.
As with every frontier AI model, safety results do not mean the model cannot make errors.
Applications that allow agents to take real-world actions still need:
- permission controls;
- monitoring;
- sandboxing;
- audit logs; and
- human approval for consequential operations.
For the broader topic, see our Agentic Misalignment guide.
Claude Haiku 5.5 Advantages
The model’s strongest advantages are:
- extremely low short-prompt pricing;
- 1M context window;
- 128K output;
- adaptive reasoning;
- fast generation;
- strong small-model benchmark results;
- image input;
- prompt caching;
- Batch pricing;
- broad cloud availability;
- Claude Code support;
- computer-use potential; and
- excellent subagent economics.
Claude Haiku 5.5 Limitations
Important limitations include:
- 5× higher pricing for prompts over 100K;
- newer tokenizer can produce more tokens than Haiku 4.5;
- not as capable as Sonnet or Opus for difficult agentic coding;
- reasoning settings can increase token use and latency;
- output remains text-only;
- no model should be assumed accurate on every task;
- long context does not guarantee perfect retrieval or reasoning over every token.
The key limitation is the pricing breakpoint.
A developer seeing:
1M context + $0.10 input
might assume a one-million-token prompt costs roughly $0.10.
It does not.
Prompts above 100K use the:
$0.50/M input
tier.
Who Should Use Claude Haiku 5.5?
Haiku 5.5 deserves testing if you run:
High-volume AI applications
Large numbers of relatively focused requests.
Customer-support automation
Routing, summarization, retrieval and draft responses.
Document-processing pipelines
Extraction, classification and summarization.
AI-agent systems
Especially as an inexpensive subagent.
Browser or computer agents
Where speed and call volume matter.
Large-scale internal search
Document questions, corporate knowledge retrieval and structured extraction.
Cost-sensitive coding workflows
Especially smaller supporting tasks.
Who Should Probably Use Sonnet or Opus Instead?
Consider a larger Claude model when the work involves:
- complicated autonomous coding;
- extremely difficult reasoning;
- high-stakes professional analysis;
- long-running tasks with many dependencies;
- sophisticated agent planning;
- tasks where failure is substantially more expensive than tokens.
Haiku is optimized for:
efficient capability
not:
maximum capability at any cost.
Is Claude Haiku 5.5 Worth It?
For many high-volume applications:
yes, it is one of the most interesting small models released in 2026.
Its biggest attraction is not simply cheap tokens.
It combines:
$0.10/M input
$0.50/M output
1M context
128K output
and:
adaptive reasoning
in the same model.
Independent Artificial Analysis results also suggest a large jump in intelligence compared with previous efficiency-class models.
However, Haiku should not automatically replace larger models.
The better strategy may be:
Haiku for routine work
plus:
Sonnet/Opus for difficult work
That routing approach can substantially reduce cost while still giving applications access to stronger models when needed.
Frequently Asked Questions
What is Claude Haiku 5.5?
Claude Haiku 5.5 is Anthropic’s fastest and cheapest current Claude model, designed for high-volume and latency-sensitive AI workloads.
When was Claude Haiku 5.5 released?
Anthropic released Claude Haiku 5.5 on:
October 7, 2026
How much does Claude Haiku 5.5 cost?
For prompts up to 100K tokens:
$0.10/M input
$0.50/M output
For prompts over 100K:
$0.50/M input
$2.50/M output
What is the Claude Haiku 5.5 context window?
Claude Haiku 5.5 has a:
1 million-token context window
What is the maximum output?
The standard maximum output is:
128K tokens
The Batch API beta can support up to:
300K output tokens
What is the Claude Haiku 5.5 API model ID?
claude-haiku-5-5
Does Claude Haiku 5.5 support reasoning?
Yes.
It supports adaptive thinking and adjustable effort.
Does Claude Haiku 5.5 support images?
Yes.
Haiku 5.5 accepts text and image input and produces text output.
Is Claude Haiku 5.5 cheaper than Haiku 4.5?
Yes.
Anthropic estimates it costs approximately 75% less on average, with raw token rates 90% lower for prompts up to 100K.
Is Claude Haiku 5.5 better than GPT-6 Luna?
Anthropic’s launch evaluations put Haiku 5.5 ahead of GPT-6 Luna on several benchmarks, but both start at the same $0.10/$0.50 short-context token pricing. Independent workload testing is the best way to choose.
Is Haiku 5.5 good for coding?
Yes for many coding tasks and subagent work, but Anthropic recommends Sonnet 5.5 or Opus 5.5 for more difficult agentic coding.
Is Claude Haiku 5.5 available in Claude Code?
Yes.
Is Claude Haiku 5.5 available on AWS?
Yes.
Anthropic lists support through Amazon Bedrock and Claude Platform on AWS.
Is Claude Haiku 5.5 available on Google Cloud?
Yes.
Is Claude Haiku 5.5 available on Microsoft?
Yes.
Anthropic lists Microsoft Foundry among supported platforms.
Bottom Line
Claude Haiku 5.5 is one of the most aggressive price-performance releases in Anthropic’s current lineup.
For prompts up to 100K tokens, it costs only:
$0.10 per million input tokens
and:
$0.50 per million output tokens
while still providing:
1M context
128K output
adaptive reasoning
image input
and strong agent-oriented capabilities.
The model is dramatically more capable than Haiku 4.5 according to Anthropic’s published benchmark results, while independent Artificial Analysis currently scores Haiku 5.5 Max at:
43 on its Intelligence Index.
The biggest caveat is pricing above 100K tokens.
Once a prompt crosses that threshold, Haiku pricing rises to:
$0.50/M input
and:
$2.50/M output
So the best Haiku workloads are likely to be high-volume tasks that remain reasonably focused rather than constantly filling the entire 1M context window.
For developers building customer support, document processing, model routing, browser agents or multi-model AI-agent systems, Haiku 5.5 deserves serious testing.
It may be especially valuable when used as a fast, inexpensive subagent alongside Sonnet or Opus.
Related Reading on Elite Era Trends
Best AI Models in 2026: GPT-6 Astra vs Gemini 4 Argon vs Claude
What Is GPT-6 Luna? Pricing, Features, Context Window & How to Use It
GPT-6 Astra: Pricing, Features, Context Window & Complete Guide