NEWAI & Machine LearningAnalysis
Physicists Re-Graded the Physics Benchmarks. GPT-5.6-Sol Jumped From 47.3% to 78.7%.
A 51-author audit of six physics benchmarks found most wrong answers were grader errors, not model errors. Corrected scores jumped as much as 31 points.
Seung Jung·
4m