Independent Benchmarks Are the Only Benchmarks: What the Haiqu vs QMill Face-Off Teaches Quantum Buyers
QuantufAI Labs ·
Brian Siegelwax recently did something the quantum industry needs far more of: he took two circuit-compression vendors — Haiqu and QMill — and benchmarked them head-to-head himself, on real hardware, without either company running the tests.
What the face-off found
The setup: three algorithms (a Grover's instance, quantum phase estimation on LiH, and Shor's factoring of 15 and 21), compressed by each vendor's tool, then executed on IBM's Heron-class hardware and IQM Emerald.
The findings that matter, summarized: Haiqu — which optimizes for specific backends and runs locally — produced smaller circuits in seconds to minutes, and those circuits executed correctly on real hardware. QMill — gateset-generic, running on the LUMI supercomputer or AWS — took from minutes to six hours, generally produced larger circuits in round one, and in one case the results couldn't be retrieved at all because of portal timeouts. The reviewer declared Haiqu the round-one winner while calling QMill a legitimate competitor, and deferred final judgment to a future round focused on result quality.
Two observations in the piece deserve more attention than the scoreboard.
First: as Siegelwax puts it, "everybody always beats Qiskit." When every vendor's marketing compares against the same permissive baseline, the comparison stops carrying information. Benchmark selection is itself a claim, and it's the one least often disclosed.
Second: one vendor's benchmark loss was partly operational, not algorithmic — six hours of supercomputer time produced results that a portal timeout made unretrievable. Result retrieval is part of the product. A compression ratio you cannot collect is a compression ratio of zero.
Why this matters if you're buying quantum
The article is a working demonstration that vendor-published numbers and independently-measured numbers are different genres of information. Nothing in it suggests bad faith by either vendor — it shows that the structure of vendor benchmarking (pick the algorithm, pick the baseline, pick the backend, publish the win) produces numbers that can't be compared across vendors. Only a third party running the same workloads through every tool produces comparable data.
If you are evaluating quantum tooling, the practical checklist falls out directly: demand the baseline, demand the backend, demand wall-clock time, and demand evidence the output executed — not just that it got smaller.
Where we stand — stated plainly
QuantufAI is quantum middleware: we route workloads to the hardware vendors we integrate with — IBM, IonQ, Rigetti, IQM, AQT, QuEra, and Pasqal are dispatch-reachable today through our direct integrations and the AWS Braket and Azure Quantum platforms, with Quantinuum currently plan-limited to its emulator. Compression vendors like the two in this article are complements to routing, not competitors: better circuits in, better routing decisions out.
And here is our disclosure, because this post is about exactly that: we do not publish execution benchmarks, because our real-hardware job execution is currently gated and we refuse to print numbers we have not measured. What a routing platform can measure honestly today is transpile-time circuit fit — depth and two-qubit gate counts per backend, computed before any job is dispatched and before any money is spent. That is the register this article validates: pre-execution metrics, disclosed baselines, and independent verification as the standard of proof. It's the standard we intend to be held to.
Reflects integration and availability, not benchmarked results — job execution is currently gated. Access varies by provider: Quantinuum is currently plan-limited to its emulator.
— QuantufAI Labs