Benchmarks · haiku55vsluna6

Haiku 5.5 vs Luna 6 benchmarks: every published result

Thirteen benchmark rows exist across the two models' published materials. Exactly one is like-for-like — the same benchmark, the same version, reported by both vendors. One more shares a row but not a version. The remaining eleven were published by a single vendor each. This page reproduces all thirteen, links every figure back to the vendor that reported it, and reads them row by row — because the honest summary of Claude Haiku 5.5 vs GPT-6 Luna benchmarks is "one comparable row", not "who wins".

01 / THE PICTURE

Percentage-scale benchmarks, charted

All ten percentage-scale rows on one 0–100 axis. A dashed empty track is an absence of data — that vendor published no result — not a score of zero.

GPT-6 Luna Claude Haiku 5.5 Not published by the vendor † Lower score is better
All bars share a 0–100 axis. The two models sit on the same row only where both vendors published a figure; elsewhere each track shows one vendor's result and one absence.

Elo and index metrics


02 / THE TABLE

Every published benchmark row, with scope notes and sources

All thirteen published benchmark rows for Claude Haiku 5.5 and GPT-6 Luna
Benchmark GPT-6 Luna Claude Haiku 5.5 Scope / reporting note
Every published figure traces back to the vendor that reported it: rows with a Luna score link to OpenAI's model documentation, rows with a Haiku score link to Anthropic's model documentation. No third-party leaderboard numbers are mixed into this table.

03 / READING THE ROWS

Haiku 5.5 vs Luna 6 benchmarks, row by row

What each row can and cannot tell you, given who reported it and on what version.

FrontierCode — the only like-for-like row

Both vendors report FrontierCode v1.1 Main: Claude Haiku 5.5 scores 46.4%, GPT-6 Luna scores 42.4%. Same benchmark, same version, same row — this is the single comparison on this page that a ranking can legitimately rest on. Two caveats survive even here: OpenAI's figure is labelled "max reasoning effort" (Luna's highest-effort configuration), and Anthropic's is a self-reported launch result rather than an independent reproduction. Both are vendor numbers; only the benchmark itself is shared.

OSWorld — one row, two different versions

The 72.4% vs 52.7% gap on OSWorld is the most misleading-looking row on the page, which is why it carries a version-mismatch flag. Luna's figure is OSWorld 2.0, offline set, v2026.08.08, at max reasoning effort. Haiku's is OSWorld 2.1, an offline subset, reported as a launch result. Different task sets, different sizes, possibly different scoring — the two numbers were not produced under the same conditions, so the 19.7-point gap supports no conclusion about which model is better at computer use.

Haiku-only rows — what Anthropic published

Anthropic reports four percentage-scale rows with no OpenAI counterpart: Humanity's Last Exam at 45.9% (no tools) and 57.4% (with tools), Terminal-Bench 4.0 at 39.2%, and Chartography at 46.4% (no tools). All are Anthropic-reported launch results. OpenAI publishes no Luna figures for these benchmarks, so there is nothing to rank them against — they establish Haiku's level, not Haiku's lead.

Luna-only rows — what OpenAI published

OpenAI reports four percentage-scale rows with no Anthropic counterpart: Agents' Last Exam V1 at 50.9%, AutomationBench 1.0.6 at 20.7%, DeepSWE v1.1 at 66.6%, and an internal Factual Error Rate evaluation at 7.6% on difficult prompts — the one inverted row on the page, where lower is better. All are labelled "max reasoning effort", and the Factual Error Rate note carries its own caveat: scores are not controlled for answer length and are not representative of typical usage. These establish Luna's level on agentic and SWE tasks; they rank against nothing on the Haiku side.

Elo and index — different scales, same rule

Three rows live on their own scales: GDPval-AA v2.1 at 1620 Elo and AA-Briefcase v1.1 at 1578 Elo (both Anthropic-reported), and the Artificial Analysis Intelligence Index v4.3.2 at 37 (OpenAI-side, captured 2026-09-22 at max reasoning effort). Elo numbers mean nothing outside their pool, and the index score is the only third-party-flavoured figure either side publishes here. None of the three has a counterpart on the other side.

Methodology: why these numbers cannot be ranked against each other

  • Luna's results are all labelled "max reasoning effort." OpenAI attaches that note to every Luna benchmark, so Luna's figures describe its highest-effort configuration.
  • Haiku 5.5's results are all "Anthropic-reported launch result." They are vendor-reported numbers, not independent reproductions — launch results are, by construction, the configuration the vendor chose to publish.
  • OSWorld is not the same version on both sides. Luna's entry is OSWorld 2.0 (offline set, v2026.08.08); Haiku's is OSWorld 2.1 (offline subset). They sit on the same row but they were not measured on the same task set.
  • Single-vendor rows have no comparator. Eleven of the thirteen rows were published by one vendor only. A number with no counterpart can set a level; it cannot produce a winner.

04 / WHERE NEXT

Cost, capabilities and the full comparison

What the models cost

Both models publish identical list rates — $0.10 input / $0.50 output per million tokens. The only published price difference is Haiku 5.5's $0.01 cache-read rate, which matters on cache-heavy workloads.

The complete head-to-head

Benchmarks are one of seven sections in the full comparison, which also covers context window, capability flags, release dates and knowledge cut-offs.