Pricing · haiku55vsluna6

Claude Haiku 5.5 pricing: every published rate, per million tokens

Claude Haiku 5.5 costs $0.10 per million input tokens, $0.50 per million output tokens, $0.01 per million cache-read tokens and $0.125 per million cache-write tokens. This page carries every published Haiku 5.5 rate, the GPT-6 Luna comparison at each rate, a cost calculator, and four worked examples of what 1,000 calls actually cost. Every figure comes from what the vendors publish; anything they leave undisclosed is marked not published rather than estimated.

01 / RATE TABLE

Claude Haiku 5.5 pricing per million tokens: the full rate table

Four published rates on the Haiku side, two on the Luna side. GPT-6 Luna is shown in every row because the list rates turn out to be identical — the difference is one row wide.

RateGPT-6 LunaClaude Haiku 5.5What it prices
Input$0.10$0.10Every fresh (uncached) input token processed by the model.
Output$0.50$0.50Every token the model generates, at up to the 128K output ceiling.
Cache readnot published$0.01Input tokens served from Anthropic's prompt cache instead of reprocessing.
Cache writenot published$0.125The one-time cost of placing a token into the prompt cache.
Only Claude Haiku 5.5 has published cache rates. Absence of a Luna cache rate means cache pricing is unknown, not that caching is free. Sources: Anthropic: anthropic.com/pricing · OpenAI: openai.com/api/pricing

02 / THE ARITHMETIC

How Haiku 5.5 pricing converts from per-million to per-call

Per-million-token rates look abstract until you apply them to a real workload. The conversion is one formula, applied to both models with their own rates.

Each call's cost is the sum of three terms, all priced from the same table above:

cost = fresh input × $0.10/1M + cached input × cache-read-rate + output × $0.50/1M

For Luna the cache-read term collapses into the input term, because OpenAI publishes no cache rate — cached input is modelled at the full $0.10 throughout this page. For Haiku 5.5 the two input terms split at whatever cache hit rate your workload actually achieves: at 0% the model is priced identically to Luna, and every percentage point of cache hits after that pulls Haiku's effective input rate down toward $0.01. One-time cache-write costs are excluded from every example below and flagged where they matter.


03 / CALCULATOR

Haiku 5.5 vs Luna 6 cost calculator

Set the workload; the calculator applies the published rates from the table above. It is arithmetic on list prices, not a quote.

100,000
0%
10,000
Cache hit rate only affects Haiku 5.5 — Luna publishes no cache rate, so its cached input is modelled at the full input rate.
GPT-6 Luna $0.00
Claude Haiku 5.5 $0.00
Computed from the published rates: input × in-rate + cached × cache-read + output × out-rate. One-time cache-write costs are not modelled, so a real-world saving would be slightly smaller. This is arithmetic on published rates, not a quote or a measurement.

04 / WORKED EXAMPLES

What 1,000 calls actually cost on Haiku 5.5

Four usage shapes, priced end to end from the published rates. The pattern across all four: without cache hits the two models are arithmetically identical, and every point of cache hit rate is pure Haiku advantage.

Example 1 — short classification calls, no caching

Workload: 2,000 input tokens and 300 output tokens per call, no cache hits. At the identical published list rates ($0.10 input / $0.50 output per million tokens), each call costs 2,000 × $0.10 ÷ 1M + 300 × $0.50 ÷ 1M = $0.00035, so per 1,000 calls:

Per 1,000 callsLunaHaiku
1,000 × $0.00035$0.35$0.35
Without caching the two are arithmetically identical — there is no price difference to find at these rates.

Example 2 — retrieval calls, half the input cached

Workload: 10,000 input tokens per call with a 50% cache hit rate, plus 1,000 output tokens. Luna publishes no cache rate, so all of its input prices at $0.10; Haiku's cached half prices at its published $0.01 cache-read rate. Per 1,000 calls:

ComponentLunaHaiku
Input (10K × 1,000)$1.00$0.55
Output (1K × 1,000)$0.50$0.50
Total$1.50$1.05
Haiku comes out 30% cheaper here ($1.05 vs $1.50). One-time cache-write costs are excluded, so a real-world saving would be slightly smaller.

Example 3 — agentic coding sessions, heavy cache reuse

Workload: 60,000 input tokens per call (a codebase slice), 80% served from cache, plus 2,000 output tokens. Luna's whole input prices at $0.10; Haiku's cached 48K prices at $0.01 and its fresh 12K at $0.10. Per 1,000 calls:

ComponentLunaHaiku
Fresh input (12K)$1.20$1.20
Cached input (48K)$4.80$0.48
Output (2K)$1.00$1.00
Total$7.00$2.68
Haiku comes out about 62% cheaper ($2.68 vs $7.00). Excluded: the one-time cache-write cost of a fresh 48K-token context, 48,000 × $0.125 ÷ 1M = $0.006 — amortised over many calls it barely moves the total, but it is not zero.

Example 4 — bulk document processing, no cache

Workload: 50,000 input tokens per call (a document dump), 0% cache hits, plus 500 output tokens. With nothing cached, both models price identically. Per 1,000 calls:

ComponentLunaHaiku
Input (50K × 1,000)$5.00$5.00
Output (0.5K × 1,000)$0.25$0.25
Total$5.25$5.25
This is the honest flip side of the cache story: for one-shot workloads with nothing worth caching, published pricing gives you no reason to prefer either model.
All four examples price the same two rate tables from section 01. Nothing here is a vendor quote, a measurement or an estimate of unpublished rates — it is multiplication and division on published numbers, shown so the per-million figures have a concrete meaning.

05 / EDGE CASES

Long context, cache writes and other Haiku 5.5 pricing edge cases

Does pricing change for long prompts?

No tiered or long-context surcharge is published for either model on the pages this comparison tracks: the same rates apply across the full 1M-token context window as far as the published pricing states. The most expensive single published quantity is a full 128K-token output, which prices at 128,000 × $0.50 ÷ 1M = $0.064 for one call.

When does the cache actually pay off?

The cache-write rate is 12.5× the input rate ($0.125 vs $0.10), so writing a token to cache costs more than reading it fresh once. Caching pays back from the second reuse onward: each subsequent read prices at $0.01 against $0.10 fresh, so a cached token breaks even on its second read and returns $0.09 of saving per read after that. Workloads that reuse a stable context — system prompts, codebases, document sets — amortise the write cost quickly; one-shot workloads should not cache at all.

No batch discounts, volume tiers, committed-use rates or fine-tuning rates are published for Claude Haiku 5.5 on the pages this comparison tracks, so none are modelled here. If a vendor publishes them, they belong in the section 01 table.

06 / FAQ

Claude Haiku 5.5 pricing FAQ

Short answers, each one arithmetic on the published rates above.

How much does Claude Haiku 5.5 cost per million tokens?

Claude Haiku 5.5 publishes four rates: $0.10 per million input tokens, $0.50 per million output tokens, $0.01 per million cache-read tokens and $0.125 per million cache-write tokens. Input and output rates are identical to GPT-6 Luna's published list price; the cache rates are the only published prices Luna does not match.

Is Claude Haiku 5.5 cheaper than GPT-6 Luna?

At list price they cost the same: both publish $0.10 per million input tokens and $0.50 per million output tokens. The published difference is caching — Haiku 5.5 prices cached input at $0.01 per million, while OpenAI publishes no cache rate for Luna. On published-rate arithmetic, any cache hit rate above zero makes Haiku 5.5 the cheaper option: roughly 30% less at a 50% hit rate (example 2) and roughly 60% less at 80% (example 3). With no cache hits the two cost exactly the same (examples 1 and 4).

Does Claude Haiku 5.5 pricing change for long prompts?

No tiered or long-context surcharge pricing is published for Claude Haiku 5.5 on the pages this comparison tracks: the same $0.10 / $0.50 rates apply across the full 1M-token context window as far as the published pricing states. The most expensive single published quantity is a full 128K-token output, which prices at 128,000 × $0.50 ÷ 1,000,000 = $0.064. If Anthropic publishes tiered rates later, they belong in this page's rate table.