Claude Haiku 5.5 pricing: every published rate, per million tokens
Claude Haiku 5.5 costs $0.10 per million input tokens, $0.50 per million output tokens, $0.01 per million cache-read tokens and $0.125 per million cache-write tokens. This page carries every published Haiku 5.5 rate, the GPT-6 Luna comparison at each rate, a cost calculator, and four worked examples of what 1,000 calls actually cost. Every figure comes from what the vendors publish; anything they leave undisclosed is marked not published rather than estimated.
Claude Haiku 5.5 pricing per million tokens: the full rate table
Four published rates on the Haiku side, two on the Luna side. GPT-6 Luna is shown in every row because the list rates turn out to be identical — the difference is one row wide.
| Rate | GPT-6 Luna | Claude Haiku 5.5 | What it prices |
|---|---|---|---|
| Input | $0.10 | $0.10 | Every fresh (uncached) input token processed by the model. |
| Output | $0.50 | $0.50 | Every token the model generates, at up to the 128K output ceiling. |
| Cache read | not published | $0.01 | Input tokens served from Anthropic's prompt cache instead of reprocessing. |
| Cache write | not published | $0.125 | The one-time cost of placing a token into the prompt cache. |
How Haiku 5.5 pricing converts from per-million to per-call
Per-million-token rates look abstract until you apply them to a real workload. The conversion is one formula, applied to both models with their own rates.
Each call's cost is the sum of three terms, all priced from the same table above:
cost = fresh input × $0.10/1M + cached input × cache-read-rate + output × $0.50/1M
For Luna the cache-read term collapses into the input term, because OpenAI publishes no cache rate — cached input is modelled at the full $0.10 throughout this page. For Haiku 5.5 the two input terms split at whatever cache hit rate your workload actually achieves: at 0% the model is priced identically to Luna, and every percentage point of cache hits after that pulls Haiku's effective input rate down toward $0.01. One-time cache-write costs are excluded from every example below and flagged where they matter.
Haiku 5.5 vs Luna 6 cost calculator
Set the workload; the calculator applies the published rates from the table above. It is arithmetic on list prices, not a quote.
input × in-rate + cached × cache-read + output × out-rate.
One-time cache-write costs are not modelled, so a real-world saving would be slightly smaller. This is
arithmetic on published rates, not a quote or a measurement.
What 1,000 calls actually cost on Haiku 5.5
Four usage shapes, priced end to end from the published rates. The pattern across all four: without cache hits the two models are arithmetically identical, and every point of cache hit rate is pure Haiku advantage.
Example 1 — short classification calls, no caching
Workload: 2,000 input tokens and 300 output tokens per call, no cache hits. At the identical published list rates ($0.10 input / $0.50 output per million tokens), each call costs 2,000 × $0.10 ÷ 1M + 300 × $0.50 ÷ 1M = $0.00035, so per 1,000 calls:
| Per 1,000 calls | Luna | Haiku |
|---|---|---|
| 1,000 × $0.00035 | $0.35 | $0.35 |
Example 2 — retrieval calls, half the input cached
Workload: 10,000 input tokens per call with a 50% cache hit rate, plus 1,000 output tokens. Luna publishes no cache rate, so all of its input prices at $0.10; Haiku's cached half prices at its published $0.01 cache-read rate. Per 1,000 calls:
| Component | Luna | Haiku |
|---|---|---|
| Input (10K × 1,000) | $1.00 | $0.55 |
| Output (1K × 1,000) | $0.50 | $0.50 |
| Total | $1.50 | $1.05 |
Example 3 — agentic coding sessions, heavy cache reuse
Workload: 60,000 input tokens per call (a codebase slice), 80% served from cache, plus 2,000 output tokens. Luna's whole input prices at $0.10; Haiku's cached 48K prices at $0.01 and its fresh 12K at $0.10. Per 1,000 calls:
| Component | Luna | Haiku |
|---|---|---|
| Fresh input (12K) | $1.20 | $1.20 |
| Cached input (48K) | $4.80 | $0.48 |
| Output (2K) | $1.00 | $1.00 |
| Total | $7.00 | $2.68 |
Example 4 — bulk document processing, no cache
Workload: 50,000 input tokens per call (a document dump), 0% cache hits, plus 500 output tokens. With nothing cached, both models price identically. Per 1,000 calls:
| Component | Luna | Haiku |
|---|---|---|
| Input (50K × 1,000) | $5.00 | $5.00 |
| Output (0.5K × 1,000) | $0.25 | $0.25 |
| Total | $5.25 | $5.25 |
Long context, cache writes and other Haiku 5.5 pricing edge cases
Does pricing change for long prompts?
No tiered or long-context surcharge is published for either model on the pages this comparison tracks: the same rates apply across the full 1M-token context window as far as the published pricing states. The most expensive single published quantity is a full 128K-token output, which prices at 128,000 × $0.50 ÷ 1M = $0.064 for one call.
When does the cache actually pay off?
The cache-write rate is 12.5× the input rate ($0.125 vs $0.10), so writing a token to cache costs more than reading it fresh once. Caching pays back from the second reuse onward: each subsequent read prices at $0.01 against $0.10 fresh, so a cached token breaks even on its second read and returns $0.09 of saving per read after that. Workloads that reuse a stable context — system prompts, codebases, document sets — amortise the write cost quickly; one-shot workloads should not cache at all.
Claude Haiku 5.5 pricing FAQ
Short answers, each one arithmetic on the published rates above.
How much does Claude Haiku 5.5 cost per million tokens?
Claude Haiku 5.5 publishes four rates: $0.10 per million input tokens, $0.50 per million output tokens, $0.01 per million cache-read tokens and $0.125 per million cache-write tokens. Input and output rates are identical to GPT-6 Luna's published list price; the cache rates are the only published prices Luna does not match.
Is Claude Haiku 5.5 cheaper than GPT-6 Luna?
At list price they cost the same: both publish $0.10 per million input tokens and $0.50 per million output tokens. The published difference is caching — Haiku 5.5 prices cached input at $0.01 per million, while OpenAI publishes no cache rate for Luna. On published-rate arithmetic, any cache hit rate above zero makes Haiku 5.5 the cheaper option: roughly 30% less at a 50% hit rate (example 2) and roughly 60% less at 80% (example 3). With no cache hits the two cost exactly the same (examples 1 and 4).
Does Claude Haiku 5.5 pricing change for long prompts?
No tiered or long-context surcharge pricing is published for Claude Haiku 5.5 on the pages this comparison tracks: the same $0.10 / $0.50 rates apply across the full 1M-token context window as far as the published pricing states. The most expensive single published quantity is a full 128K-token output, which prices at 128,000 × $0.50 ÷ 1,000,000 = $0.064. If Anthropic publishes tiered rates later, they belong in this page's rate table.
For the other half of the comparison, see Haiku 5.5 vs Luna 6 benchmarks or the full Claude Haiku 5.5 vs GPT-6 Luna comparison.