A team of 51 physicists and computer scientists re-examined six widely used physics benchmarks and found that the low scores frontier models post on them mostly measure the benchmarks, not the models. After experts corrected faulty reference solutions and repaired or removed flawed questions, GPT-5.6-Sol's mean@4 score on HLE-Physics rose from 47.3% to 78.7%.
Key takeaways
- On HLE-Physics, GPT-5.6-Sol's mean@4 climbed from 47.3% to 78.7% after expert re-grading; on CMT-Benchmark it went from 61.0% to 87.2%.
- Corrected pass@4 reached 94.4% on the 54 CritPt challenges that survived expert review, a level the authors describe as near-saturation.
- The paper directly disputes the impression created by low physics scores in the Artificial Analysis Intelligence Index (2026), arguing most audited errors were grading artifacts rather than reasoning failures.
What the auditors actually did
The study, published as arXiv:2609.13009, started from a mismatch its authors kept noticing. Published benchmark results said frontier models still struggle with advanced physics. Physicists who use those same models in their research reported a different experience.
Rather than argue from anecdote, the team ran frontier models across six benchmarks and put the results in front of domain experts. Faculty and graduate researchers with relevant subfield expertise read each problem statement, each reference solution and each model response, restricting the scope to text-only problems with verifiable final answers.
Their job was to sort failures into categories: genuine model errors, grader errors, incorrect reference solutions, and questions too ambiguous or underspecified to have a defensible answer. The headline finding is that most audited cases initially scored as incorrect fell into the latter three buckets.
Why the corrected scores moved so far
Once the experts fixed erroneous reference solutions and repaired or excluded broken questions, scores rose sharply across the board. The two largest documented jumps belong to GPT-5.6-Sol: HLE-Physics from 47.3% to 78.7% and CMT-Benchmark from 61.0% to 87.2%, both measured as mean@4.
On CritPt, corrected pass@4 hit 94.4% across the 54 problems retained after review. The authors report that audited subsets of UGPhysics, PRISM-Physics and PHYBench also rose substantially. Crucially, corrected scores are computed only on the retained subsets โ the numbers describe performance on problems that survived expert scrutiny, not on the original full sets.
The measurement problem this exposes
The result lands on a live dispute about what public large language model leaderboards are measuring. If a benchmark's reference answers are wrong, a model that reasons correctly gets marked down, and the aggregate index built on top of that benchmark propagates the error to anyone reading it.
That failure mode is not unique to physics. Evaluation scaffolding has repeatedly proven to be the variable that moves scores: Nvidia's AVO harness took Claude Opus 5 from 30% to a perfect ARC-AGI-3 run without changing the model at all. Here the scaffolding under scrutiny is the grading itself.
What comes after saturation
The authors do not present the corrected numbers as vindication of model capability in general. Their conclusion is narrower and sharper: current benchmarks substantially understate what frontier models can do on well-posed, closed-ended physics problems, and those problems are close to exhausted as a source of signal.
That points the field toward harder, expert-validated evaluations โ open-ended research tasks where a verifiable final answer does not exist, and where an expert has to judge the reasoning rather than check a number. Building those is considerably more expensive than scraping problem sets, which is part of why the broken ones stayed in circulation.
FAQ
Which physics benchmarks were audited?
Six: HLE-Physics, CMT-Benchmark, CritPt, UGPhysics, PRISM-Physics and PHYBench. The audit covered text-only problems with verifiable final answers, and corrected scores were computed on the subsets retained after expert review.
Does this mean frontier models have solved physics?
No. The paper's claim is that models are near-saturation on closed-ended physics problems with verifiable answers, which is a narrow slice of physics work. The authors explicitly call for more demanding, expert-validated evaluations rather than declaring the problem solved.
Why do the original benchmark scores matter if they were wrong?
Because aggregate indices repeat them. Low physics scores featured in the Artificial Analysis Intelligence Index (2026) shaped the public read on model reasoning; if the underlying grading was faulty, that read was too.






