October 10, 2026

Claude Haiku 5.5: Pricing, Features, Benchmarks, API & Is It Worth It?

Claude Haiku 5.5: Pricing, Features, Benchmarks, API & Is It Worth It?

Claude Haiku 5.5: Pricing, Features, Benchmarks, API & Is It Worth It?

Anthropic has released Claude Haiku 5.5, its newest low-cost AI model aimed at developers and businesses that need fast AI at high volume.

Released on:

October 7, 2026

Claude Haiku 5.5 combines:

  • a 1 million-token context window;
  • up to 128,000 standard output tokens;
  • adaptive reasoning;
  • image understanding;
  • very low API pricing;
  • prompt caching;
  • Batch API discounts;
  • computer-use and browser-oriented workloads; and
  • support across Anthropic, AWS, Google Cloud and Microsoft.

Its headline API price is particularly aggressive.

For prompts up to 100,000 tokens, Anthropic charges:

$0.10 per million input tokens

and:

$0.50 per million output tokens

That places Haiku 5.5 directly against efficient models such as OpenAI’s GPT-6 Luna.

However, there is one major pricing detail developers need to understand:

Claude Haiku 5.5 costs five times more per token when the prompt exceeds 100,000 tokens.

So while the model has a 1M-token context window, using that context can materially change its economics.

Here is everything you need to know about Claude Haiku 5.5.

Claude Haiku 5.5 at a Glance

FeatureClaude Haiku 5.5
DeveloperAnthropic
Release dateOctober 7, 2026
API model IDclaude-haiku-5-5
Context window1,000,000 tokens
Standard max output128,000 tokens
Batch API max output300,000 tokens beta
Input price ≤100K prompt$0.10 / 1M tokens
Output price ≤100K prompt$0.50 / 1M tokens
Input price >100K prompt$0.50 / 1M tokens
Output price >100K prompt$2.50 / 1M tokens
Cache read ≤100K$0.01 / 1M
Batch discount50% input/output
ReasoningAdaptive
Default effortMedium
InputText + images
OutputText
Reliable knowledge cutoffJune 2026
StatusActive

Anthropic positions Haiku 5.5 as its:

fastest and cheapest current Claude model

for high-volume, latency-sensitive work.

Official documentation:

Anthropic — Claude Haiku 5.5 Model Documentation

What Is Claude Haiku 5.5?

Claude Haiku 5.5 is Anthropic’s efficiency-focused model in the Claude 5.5 family.

Anthropic currently positions its major models roughly like this:

Claude Opus 5.5 → maximum capability

Claude Sonnet 5.5 → strong balance of capability and cost

Claude Haiku 5.5 → speed, scale and low cost

Haiku is therefore not primarily designed to replace Opus for the most difficult reasoning or autonomous coding tasks.

Instead, Anthropic recommends it for workloads such as:

  • summarization;
  • classification;
  • extraction;
  • routing;
  • document queries;
  • customer support;
  • database queries;
  • compaction;
  • subagents; and
  • other repetitive AI operations.

This makes Haiku 5.5 particularly interesting for businesses processing thousands or millions of requests.

For a broader comparison of frontier models, see our Best AI Models in 2026 guide.

Claude Haiku 5.5 Pricing

Anthropic introduced a two-tier pricing structure for Haiku 5.5.

Prompts up to 100,000 tokens

Input:

$0.10 per million tokens

Output:

$0.50 per million tokens

Cache reads:

$0.01 per million tokens

5-minute cache writes:

$0.125 per million tokens

1-hour cache writes:

$0.20 per million tokens

Prompts over 100,000 tokens

Input:

$0.50 per million tokens

Output:

$2.50 per million tokens

Cache reads:

$0.05 per million tokens

5-minute cache writes:

$0.625 per million tokens

1-hour cache writes:

$1 per million tokens

So crossing the:

100,000-token prompt threshold

increases base input and output pricing by:

5×

That is critical when calculating the cost of long-document workflows.

Claude Haiku 5.5 Pricing Table

UsagePrompt ≤100KPrompt >100K
Input$0.10/M$0.50/M
Output$0.50/M$2.50/M
Cache read$0.01/M$0.05/M
5-min cache write$0.125/M$0.625/M
1-hour cache write$0.20/M$1.00/M

Anthropic also offers a:

50% Batch API discount

on input and output tokens.

Official pricing:

Anthropic — Claude Haiku 5.5 Pricing

Claude Haiku 5.5 Cost Example

Suppose an application processes:

100,000 input tokens

and generates:

10,000 output tokens

while remaining in the lower pricing tier.

Input

100,000 tokens = 0.1 million

0.1 × $0.10 = $0.01

Output

10,000 tokens = 0.01 million

0.01 × $0.50 = $0.005

Total:

$0.015

That is:

1.5 cents

for this simplified example.

This excludes caching, tools and other possible charges.

What Happens Above 100K Tokens?

Suppose a request instead contains:

200,000 input tokens

and generates:

10,000 output tokens

Because the prompt is over 100K, the higher pricing tier applies.

Input

0.2 × $0.50 = $0.10

Output

0.01 × $2.50 = $0.025

Total:

$0.125

So context length can make a major difference to the economics of Haiku 5.5.

Developers should not assume the advertised:

$0.10 / $0.50

rate applies to every request across the model’s 1M context window.

Is Claude Haiku 5.5 Cheaper Than Haiku 4.5?

Yes.

Anthropic says Haiku 5.5 costs approximately:

75% less on average

to run than Haiku 4.5.

For prompts up to 100K tokens, the raw price difference is even larger.

Claude Haiku 4.5 was priced at:

$1 per million input tokens

and:

$5 per million output tokens

Haiku 5.5 costs:

$0.10 input

and:

$0.50 output

for prompts up to 100K.

That is a:

90% reduction in base token pricing

in that tier.

For longer prompts, Anthropic says Haiku 5.5’s pricing remains roughly:

50% lower

than Haiku 4.5.

One Important Tokenizer Detail

There is a complication when comparing raw token prices.

Anthropic says Claude Haiku 5.5 uses its newer tokenizer.

The same text can produce approximately:

30% more tokens

than on Claude Haiku 4.5.

The exact difference depends on content.

That means developers should benchmark:

actual cost per task

rather than comparing only advertised token prices.

Even with this tokenizer difference, Anthropic estimates Haiku 5.5 costs roughly 75% less on average than its predecessor.

Claude Haiku 5.5 Context Window

Claude Haiku 5.5 supports:

1,000,000 tokens of context

by default.

Anthropic says no special beta header is required to use the full context capacity.

This is an enormous context window for a model at Haiku’s price point.

Potential use cases include:

  • long reports;
  • large document collections;
  • code repositories;
  • business records;
  • legal documents;
  • financial reports;
  • support histories;
  • research collections;
  • knowledge bases; and
  • long-running agent context.

However, remember:

prompts above 100K tokens use the higher Haiku pricing tier.

So the model’s 1M context window is technically powerful but should still be used selectively.

Claude Haiku 5.5 Maximum Output

Anthropic lists the normal maximum output as:

128,000 tokens

That is comparable to larger Claude models.

Haiku 5.5 also supports:

up to 300,000 output tokens

through the Message Batches API beta.

The beta capability requires Anthropic’s applicable feature configuration.

Most applications will never need a 128K response, let alone 300K.

But the capability could matter for:

  • bulk structured extraction;
  • extensive document transformations;
  • code generation;
  • large reports; and
  • long offline workflows.

Adaptive Thinking Comes to Haiku

Haiku 5.5 is the first Haiku-class model to support Anthropic’s:

adaptive thinking

with adjustable effort.

The default effort setting is:

medium

Developers can tell the model to spend more or less effort depending on the task.

That gives applications a useful trade-off between:

  • intelligence;
  • latency; and
  • cost.

For routine classification, lower reasoning effort may be enough.

For more difficult document analysis or coding, higher effort can improve performance.

Claude Haiku 5.5 Benchmarks

Anthropic published several benchmark comparisons at launch.

Provider-published benchmark results should always be treated as evidence from the model developer rather than as completely neutral testing.

With that qualification, Anthropic reports:

BenchmarkHaiku 5.5Haiku 4.5GPT-6 LunaSonnet 5.5
GDPval-AA v2.1162073514371840
AA-Briefcase v1.1157861413361824
OSWorld 2.1 offline72.4%15.7%48.9%83.9%
Humanity’s Last Exam, no tools45.9%10.2%—56.9%
Terminal-Bench 4.039.2%0.0%16.4%70.6%
FrontierCode 1.146.4%—42.4%52.1%
Chartography46.4%6.4%29.1%61.6%

Source:

Anthropic — Introducing Claude Haiku 5.5

The results suggest Haiku 5.5 is a major capability jump over Haiku 4.5.

But Anthropic itself says:

Sonnet and Opus remain better choices for complex agentic coding.

Independent Benchmark Results

Artificial Analysis has also tested Claude Haiku 5.5 independently.

At its maximum tested reasoning setting, Artificial Analysis currently reports an:

Intelligence Index score of 43

for Haiku 5.5.

Other effort settings currently score approximately:

Medium: 34

High: 38

Xhigh: 41

Max: 43

Artificial Analysis also reports very high generation speed for Haiku 5.5, depending on the reasoning setting.

Independent results:

Artificial Analysis — Claude Haiku 5.5

The main lesson is that effort setting matters.

You cannot describe Haiku 5.5 with one intelligence number without also specifying how much reasoning effort the model is using.

Claude Haiku 5.5 vs GPT-6 Luna

This may be one of the most interesting low-cost AI comparisons of 2026.

Both models start at:

$0.10 per million input tokens

and:

$0.50 per million output tokens

for their lowest standard pricing tiers.

Both also offer approximately:

1 million tokens of context

and:

128K standard output

But their pricing rules differ.

FeatureClaude Haiku 5.5GPT-6 Luna
DeveloperAnthropicOpenAI
Short-context input$0.10/M$0.10/M
Short-context output$0.50/M$0.50/M
Context1M1.05M
Max output128K128K
Image inputYesYes
ReasoningAdaptiveYes
Long-context price increase5× above 100K promptUses OpenAI long-context pricing rules
Main positioningHigh-volume, fast ClaudeEfficient high-volume GPT-6

Anthropic’s own benchmark table places Haiku 5.5 ahead of GPT-6 Luna on several tests, including:

  • GDPval-AA;
  • AA-Briefcase;
  • OSWorld;
  • Terminal-Bench;
  • FrontierCode; and
  • Chartography.

But because Anthropic published that comparison, developers should also test the models on their own workload.

For full Luna specifications, see our GPT-6 Luna pricing and features guide.

Claude Haiku 5.5 vs Claude Sonnet 5.5

Haiku is dramatically cheaper.

Claude Sonnet 5.5 currently costs:

$2 per million input tokens

and:

$10 per million output tokens

Haiku 5.5 begins at:

$0.10

and:

$0.50

respectively.

That means Sonnet’s base token rates are:

20× higher

than Haiku’s lowest tier.

But Sonnet is also substantially more capable on difficult agentic tasks.

Anthropic’s Terminal-Bench comparison shows:

Haiku 5.5: 39.2%

Sonnet 5.5: 70.6%

So the right choice depends on the workload.

Use Haiku when:

volume + speed + low cost

matter most.

Use Sonnet when:

complex task completion

matters more than raw token price.

Claude Haiku 5.5 vs Claude Opus 5.5

The difference is even clearer with Opus.

Anthropic lists Opus 5.5 at:

$4/M input

and:

$20/M output

Haiku starts at:

$0.10/M input

and:

$0.50/M output

That makes Opus:

40× more expensive per base token

at Haiku’s lowest tier.

But Opus is intended for the hardest professional and long-running agent tasks.

A useful architecture may be:

Haiku → routine subtask

Sonnet → difficult workflow

Opus → hardest reasoning / coding

This can keep AI-agent costs manageable.

Haiku 5.5 Is Designed for Subagents

One of Anthropic’s most interesting recommendations is to use Haiku 5.5 as a:

subagent

alongside larger Claude models.

Imagine an AI system where Opus manages the overall task.

Rather than asking Opus to perform every smaller operation, the system could delegate tasks such as:

  • document lookup;
  • classification;
  • extraction;
  • summarization;
  • retrieval;
  • data cleanup; and
  • simple coding steps

to Haiku.

This architecture can potentially reduce both:

cost

and:

latency

while reserving larger models for difficult decisions.

For more examples of agentic workflows, see our AI Agents for Small Business guide.

Claude Haiku 5.5 for Customer Support

Customer support is one of the most natural Haiku use cases.

Support systems often need to process large numbers of short interactions.

Tasks can include:

  • identifying customer intent;
  • routing tickets;
  • searching knowledge bases;
  • summarizing conversations;
  • drafting responses;
  • classifying urgency; and
  • deciding when to escalate to a human.

Anthropic describes Haiku 5.5 as especially well suited to:

live customer support

because of its speed and low cost.

Businesses should still test hallucination rates and use appropriate human escalation for consequential decisions.

Claude Haiku 5.5 for Document Processing

Another strong use case is document analysis at scale.

Possible workloads include:

  • invoice extraction;
  • financial report summaries;
  • contract classification;
  • document routing;
  • form processing;
  • record tagging;
  • structured data extraction;
  • document Q&A.

Haiku’s:

1M context window

allows very large documents or collections to fit into the same workflow.

But applications should monitor the 100K-token pricing threshold.

Is Claude Haiku 5.5 Good for Coding?

Yes—but with an important qualification.

Haiku 5.5 is dramatically more capable at coding than Haiku 4.5.

Anthropic reports:

46.4% on FrontierCode 1.1

and:

39.2% on Terminal-Bench 4.0

However, Anthropic explicitly states that:

Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding.

Haiku is better suited to coding subtasks such as:

  • simple fixes;
  • code classification;
  • repository search;
  • test generation;
  • formatting;
  • documentation;
  • narrow subagent jobs.

For a broader coding comparison across leading frontier models, start with our Best AI Models in 2026 hub.

Is Claude Haiku 5.5 Good for Computer Use?

Anthropic reports a major improvement here.

On its OSWorld 2.1 offline-subset evaluation:

Haiku 5.5: 72.4%

versus:

Haiku 4.5: 15.7%

Anthropic is also updating its Python and TypeScript SDKs with:

computer-use and browser-use support in beta

and specifically says Haiku 5.5 is well suited to these workflows because of its combination of:

speed + price + capability

That could make Haiku especially relevant for high-volume browser agents.

Claude Haiku 5.5 API Model ID

The Claude API model ID is:

claude-haiku-5-5

Anthropic says it is a:

fixed model ID

with no date suffix.

Platform identifiers include:

Claude API

claude-haiku-5-5

Amazon Bedrock

anthropic.claude-haiku-5-5

Google Cloud

claude-haiku-5-5

Microsoft Foundry

claude-haiku-5-5

Anthropic also makes the model available through Claude Platform on AWS.

Claude Haiku 5.5 Availability

Claude Haiku 5.5 is available now.

Anthropic lists support through:

  • Claude.ai;
  • Claude API;
  • Claude Code;
  • Amazon Web Services;
  • Google Cloud;
  • Microsoft Foundry; and
  • Claude Platform on AWS.

Anthropic says Free, Pro, Max, Team and Enterprise Claude users can select Haiku 5.5 where available.

The model is currently listed as:

Active — Latest

Anthropic says retirement will occur:

not sooner than October 7, 2027

Claude Haiku 5.5 in Claude Code

Haiku 5.5 is also available in:

Claude Code

This is especially interesting for systems that combine:

Opus or Sonnet as the main coding model

with:

Haiku as a sidekick or subagent

for cheaper supporting work.

That model-routing architecture could become increasingly common as AI-agent systems grow more sophisticated.

Prompt Caching Can Make Haiku Even Cheaper

Haiku 5.5 supports prompt caching.

For prompts up to 100K, cache reads cost:

$0.01 per million tokens

That is a:

90% reduction

compared with ordinary $0.10 input pricing.

Prompt caching can be particularly useful when an application repeatedly sends the same:

  • system prompt;
  • product documentation;
  • policy manual;
  • code context; or
  • large reference material.

Instead of paying the full input rate each time, the cached content can be reused at a much lower price.

Batch API Discount

Anthropic also offers:

50% off input and output pricing

through the Batch API.

That can make Haiku 5.5 especially inexpensive for offline work such as:

  • large classification jobs;
  • bulk extraction;
  • dataset labeling;
  • document summarization;
  • evaluation runs;
  • content processing.

If a response does not need to arrive immediately, Batch can materially reduce cost.

Haiku 5.5 Safety

Anthropic says Haiku 5.5 performed better than Haiku 4.5 across most of its alignment evaluations.

The company reports:

  • fewer examples of misaligned behavior;
  • lower willingness to assist misuse;
  • stronger cyber safeguards than Haiku 4.5; and
  • biology safeguards comparable to several larger Claude models.

Anthropic says the model permits many defensive cybersecurity tasks but restricts penetration testing and other higher-risk techniques.

As with every frontier AI model, safety results do not mean the model cannot make errors.

Applications that allow agents to take real-world actions still need:

  • permission controls;
  • monitoring;
  • sandboxing;
  • audit logs; and
  • human approval for consequential operations.

For the broader topic, see our Agentic Misalignment guide.

Claude Haiku 5.5 Advantages

The model’s strongest advantages are:

  • extremely low short-prompt pricing;
  • 1M context window;
  • 128K output;
  • adaptive reasoning;
  • fast generation;
  • strong small-model benchmark results;
  • image input;
  • prompt caching;
  • Batch pricing;
  • broad cloud availability;
  • Claude Code support;
  • computer-use potential; and
  • excellent subagent economics.

Claude Haiku 5.5 Limitations

Important limitations include:

  • 5× higher pricing for prompts over 100K;
  • newer tokenizer can produce more tokens than Haiku 4.5;
  • not as capable as Sonnet or Opus for difficult agentic coding;
  • reasoning settings can increase token use and latency;
  • output remains text-only;
  • no model should be assumed accurate on every task;
  • long context does not guarantee perfect retrieval or reasoning over every token.

The key limitation is the pricing breakpoint.

A developer seeing:

1M context + $0.10 input

might assume a one-million-token prompt costs roughly $0.10.

It does not.

Prompts above 100K use the:

$0.50/M input

tier.

Who Should Use Claude Haiku 5.5?

Haiku 5.5 deserves testing if you run:

High-volume AI applications

Large numbers of relatively focused requests.

Customer-support automation

Routing, summarization, retrieval and draft responses.

Document-processing pipelines

Extraction, classification and summarization.

AI-agent systems

Especially as an inexpensive subagent.

Browser or computer agents

Where speed and call volume matter.

Document questions, corporate knowledge retrieval and structured extraction.

Cost-sensitive coding workflows

Especially smaller supporting tasks.

Who Should Probably Use Sonnet or Opus Instead?

Consider a larger Claude model when the work involves:

  • complicated autonomous coding;
  • extremely difficult reasoning;
  • high-stakes professional analysis;
  • long-running tasks with many dependencies;
  • sophisticated agent planning;
  • tasks where failure is substantially more expensive than tokens.

Haiku is optimized for:

efficient capability

not:

maximum capability at any cost.

Is Claude Haiku 5.5 Worth It?

For many high-volume applications:

yes, it is one of the most interesting small models released in 2026.

Its biggest attraction is not simply cheap tokens.

It combines:

$0.10/M input

$0.50/M output

1M context

128K output

and:

adaptive reasoning

in the same model.

Independent Artificial Analysis results also suggest a large jump in intelligence compared with previous efficiency-class models.

However, Haiku should not automatically replace larger models.

The better strategy may be:

Haiku for routine work

plus:

Sonnet/Opus for difficult work

That routing approach can substantially reduce cost while still giving applications access to stronger models when needed.

Frequently Asked Questions

What is Claude Haiku 5.5?

Claude Haiku 5.5 is Anthropic’s fastest and cheapest current Claude model, designed for high-volume and latency-sensitive AI workloads.

When was Claude Haiku 5.5 released?

Anthropic released Claude Haiku 5.5 on:

October 7, 2026

How much does Claude Haiku 5.5 cost?

For prompts up to 100K tokens:

$0.10/M input

$0.50/M output

For prompts over 100K:

$0.50/M input

$2.50/M output

What is the Claude Haiku 5.5 context window?

Claude Haiku 5.5 has a:

1 million-token context window

What is the maximum output?

The standard maximum output is:

128K tokens

The Batch API beta can support up to:

300K output tokens

What is the Claude Haiku 5.5 API model ID?

claude-haiku-5-5

Does Claude Haiku 5.5 support reasoning?

Yes.

It supports adaptive thinking and adjustable effort.

Does Claude Haiku 5.5 support images?

Yes.

Haiku 5.5 accepts text and image input and produces text output.

Is Claude Haiku 5.5 cheaper than Haiku 4.5?

Yes.

Anthropic estimates it costs approximately 75% less on average, with raw token rates 90% lower for prompts up to 100K.

Is Claude Haiku 5.5 better than GPT-6 Luna?

Anthropic’s launch evaluations put Haiku 5.5 ahead of GPT-6 Luna on several benchmarks, but both start at the same $0.10/$0.50 short-context token pricing. Independent workload testing is the best way to choose.

Is Haiku 5.5 good for coding?

Yes for many coding tasks and subagent work, but Anthropic recommends Sonnet 5.5 or Opus 5.5 for more difficult agentic coding.

Is Claude Haiku 5.5 available in Claude Code?

Yes.

Is Claude Haiku 5.5 available on AWS?

Yes.

Anthropic lists support through Amazon Bedrock and Claude Platform on AWS.

Is Claude Haiku 5.5 available on Google Cloud?

Yes.

Is Claude Haiku 5.5 available on Microsoft?

Yes.

Anthropic lists Microsoft Foundry among supported platforms.

Bottom Line

Claude Haiku 5.5 is one of the most aggressive price-performance releases in Anthropic’s current lineup.

For prompts up to 100K tokens, it costs only:

$0.10 per million input tokens

and:

$0.50 per million output tokens

while still providing:

1M context

128K output

adaptive reasoning

image input

and strong agent-oriented capabilities.

The model is dramatically more capable than Haiku 4.5 according to Anthropic’s published benchmark results, while independent Artificial Analysis currently scores Haiku 5.5 Max at:

43 on its Intelligence Index.

The biggest caveat is pricing above 100K tokens.

Once a prompt crosses that threshold, Haiku pricing rises to:

$0.50/M input

and:

$2.50/M output

So the best Haiku workloads are likely to be high-volume tasks that remain reasonably focused rather than constantly filling the entire 1M context window.

For developers building customer support, document processing, model routing, browser agents or multi-model AI-agent systems, Haiku 5.5 deserves serious testing.

It may be especially valuable when used as a fast, inexpensive subagent alongside Sonnet or Opus.

Best AI Models in 2026: GPT-6 Astra vs Gemini 4 Argon vs Claude

What Is GPT-6 Luna? Pricing, Features, Context Window & How to Use It

GPT-6 Astra: Pricing, Features, Context Window & Complete Guide

AI Agents for Small Business: 15 Practical Uses in 2026