ProofCouncil: AI Agent Wins Mathematical Challenge, Solves Real Open Problems
Key Takeaways
- ▸ProofCouncil achieved 60% full-solution rate on FirstProof challenge, winning best-in-class among all teams
- ▸On 30 researcher-provided open problems, the agent produced 5 complete solutions and 18 additional problems with useful progress or promising leads
- ▸Author-critic agentic architecture demonstrates that LLMs can tackle frontier mathematics requiring genuine problem-solving rather than pattern matching
Summary
Researchers have unveiled ProofCouncil, an LLM agent that uses an author-critic architecture to tackle open-ended mathematical problems. Competing in the second batch of the FirstProof challenge—a rigorous test of 10 real-world mathematical problems—ProofCouncil achieved best-in-class performance with correct or near-correct solutions for 6 of 10 problems, outperforming all other participating teams. Beyond competition results, the agent was evaluated on 30 additional open problems sourced from mathematics researchers. Of the 21 problems receiving expert human feedback, ProofCouncil produced 5 completely correct solutions, 2 promising approaches pending verification, and 8 containing substantial partial progress. The research team is releasing the underlying agent-building library as open source, democratizing the advanced workflows that power the system.
- Open-source release of agent-building library enables broader research community to build on these advances
Editorial Opinion
ProofCouncil is a genuinely significant milestone: an LLM agent that doesn't just win benchmarks but produces correct mathematics on open problems. The 60% success rate on FirstProof and demonstrated progress on researcher-curated open problems show we've moved beyond toy domains into territory where autonomous agents might actually contribute to mathematical research. Most importantly, open-sourcing the agent-building library sets a precedent for transparency and community acceleration in this crucial capability area.


