SciCode-Verified: Benchmark Audit Reveals Language Models 40-70% More Capable Than Previously Measured
Key Takeaways
- ▸SciCode benchmark contained 263 defects affecting 91% of main problems, causing systematic underestimation of LLM capabilities
- ▸Corrected benchmark (SciCode-Verified) shows LLMs achieve 84-98% subproblem accuracy vs. previously measured 45-60%—a 40-70% improvement
- ▸78% of defects required specialist domain knowledge (physics/mathematics) to identify, not clerical review—establishing new quality standards for benchmark design
Summary
A comprehensive audit of the SciCode benchmark—the standard measure of LLM scientific-coding ability—has uncovered 263 defects that were systematically underestimating model performance. The defects, spread across 91% of main test problems, caused correct solutions to be wrongly rejected due to non-reproducible gold answers, overly tight tolerances, and self-contradictory specifications. Researchers corrected all confirmable defects to produce SciCode-Verified, with each change independently verified by a second domain expert.
Re-evaluation of twelve frontier model snapshots on the corrected benchmark reveals dramatic performance recovery: subproblem accuracy jumped from 45–60% to 84–98%, while main-problem accuracy rose from 9–27% to 69–92%. The research demonstrates that state-of-the-art language models are far more proficient in scientific coding than SciCode suggested—the bottleneck was not model capability, but the quality of the evaluation instrument itself. Critically, 78% of score-suppressing defects required specialized physics and mathematics knowledge to detect, highlighting the importance of domain-expert review in benchmark construction.
- SciCode-Verified is released as the corrected public standard with complete audit trail and independent verification for all corrections
- Finding reveals that evaluation methodology, not model capability, was the limiting factor in scientific coding benchmarks
Editorial Opinion
This research exposes a hard truth in AI benchmarking: the measurement tool itself can be the primary constraint on perceived progress. By revealing that models are 40-70% more capable than SciCode indicated, this work fundamentally challenges how we assess LLM abilities and sets a higher bar for benchmark rigor across the industry. The requirement for domain-expert auditing—not just clerical review—should become standard practice for all high-stakes evaluation frameworks, particularly those guiding government and national-laboratory AI assessments.



