Haiku 5.5 vs Luna 6 benchmarks: every published result
Thirteen benchmark rows exist across the two models' published materials. Exactly one is like-for-like — the same benchmark, the same version, reported by both vendors. One more shares a row but not a version. The remaining eleven were published by a single vendor each. This page reproduces all thirteen, links every figure back to the vendor that reported it, and reads them row by row — because the honest summary of Claude Haiku 5.5 vs GPT-6 Luna benchmarks is "one comparable row", not "who wins".
Percentage-scale benchmarks, charted
All ten percentage-scale rows on one 0–100 axis. A dashed empty track is an absence of data — that vendor published no result — not a score of zero.
Elo and index metrics
Every published benchmark row, with scope notes and sources
| Benchmark | GPT-6 Luna | Claude Haiku 5.5 | Scope / reporting note |
|---|
Haiku 5.5 vs Luna 6 benchmarks, row by row
What each row can and cannot tell you, given who reported it and on what version.
FrontierCode — the only like-for-like row
Both vendors report FrontierCode v1.1 Main: Claude Haiku 5.5 scores 46.4%, GPT-6 Luna scores 42.4%. Same benchmark, same version, same row — this is the single comparison on this page that a ranking can legitimately rest on. Two caveats survive even here: OpenAI's figure is labelled "max reasoning effort" (Luna's highest-effort configuration), and Anthropic's is a self-reported launch result rather than an independent reproduction. Both are vendor numbers; only the benchmark itself is shared.
OSWorld — one row, two different versions
The 72.4% vs 52.7% gap on OSWorld is the most misleading-looking row on the page, which is why it carries a version-mismatch flag. Luna's figure is OSWorld 2.0, offline set, v2026.08.08, at max reasoning effort. Haiku's is OSWorld 2.1, an offline subset, reported as a launch result. Different task sets, different sizes, possibly different scoring — the two numbers were not produced under the same conditions, so the 19.7-point gap supports no conclusion about which model is better at computer use.
Haiku-only rows — what Anthropic published
Anthropic reports four percentage-scale rows with no OpenAI counterpart: Humanity's Last Exam at 45.9% (no tools) and 57.4% (with tools), Terminal-Bench 4.0 at 39.2%, and Chartography at 46.4% (no tools). All are Anthropic-reported launch results. OpenAI publishes no Luna figures for these benchmarks, so there is nothing to rank them against — they establish Haiku's level, not Haiku's lead.
Luna-only rows — what OpenAI published
OpenAI reports four percentage-scale rows with no Anthropic counterpart: Agents' Last Exam V1 at 50.9%, AutomationBench 1.0.6 at 20.7%, DeepSWE v1.1 at 66.6%, and an internal Factual Error Rate evaluation at 7.6% on difficult prompts — the one inverted row on the page, where lower is better. All are labelled "max reasoning effort", and the Factual Error Rate note carries its own caveat: scores are not controlled for answer length and are not representative of typical usage. These establish Luna's level on agentic and SWE tasks; they rank against nothing on the Haiku side.
Elo and index — different scales, same rule
Three rows live on their own scales: GDPval-AA v2.1 at 1620 Elo and AA-Briefcase v1.1 at 1578 Elo (both Anthropic-reported), and the Artificial Analysis Intelligence Index v4.3.2 at 37 (OpenAI-side, captured 2026-09-22 at max reasoning effort). Elo numbers mean nothing outside their pool, and the index score is the only third-party-flavoured figure either side publishes here. None of the three has a counterpart on the other side.
Methodology: why these numbers cannot be ranked against each other
- Luna's results are all labelled "max reasoning effort." OpenAI attaches that note to every Luna benchmark, so Luna's figures describe its highest-effort configuration.
- Haiku 5.5's results are all "Anthropic-reported launch result." They are vendor-reported numbers, not independent reproductions — launch results are, by construction, the configuration the vendor chose to publish.
- OSWorld is not the same version on both sides. Luna's entry is OSWorld 2.0 (offline set, v2026.08.08); Haiku's is OSWorld 2.1 (offline subset). They sit on the same row but they were not measured on the same task set.
- Single-vendor rows have no comparator. Eleven of the thirteen rows were published by one vendor only. A number with no counterpart can set a level; it cannot produce a winner.
Cost, capabilities and the full comparison
What the models cost
Both models publish identical list rates — $0.10 input / $0.50 output per million tokens. The only published price difference is Haiku 5.5's $0.01 cache-read rate, which matters on cache-heavy workloads.
Claude Haiku 5.5 pricing — full rate table, calculator and worked examples
The complete head-to-head
Benchmarks are one of seven sections in the full comparison, which also covers context window, capability flags, release dates and knowledge cut-offs.