Math Reasoning Benchmarks
FrontierMath
- 资料:
../papers_md/2411.04872_frontiermath/note.md
- 译文:
../papers_md/2411.04872_frontiermath/main_zh.md
- 核心定位: 原创未公开、研究级数学题,答案自动验证,面向闭源 frontier model 的高级数学推理评测。
- 关键判断: 它把 benchmark 设计重点从“题目是否更难”推进到“难题是否原创、防污染、可自动评测、可由专家审核”。当前模型低于 2% 的表现说明它短期内更像能力上限探针,而不是细粒度模型排序工具。
Hard2Verify
- 资料:
../papers_md/2026.acl-long.1031_hard2verify/note.md
- 核心定位: 开放式前沿数学解答的 step-level verification benchmark。
- 与 FrontierMath 的关系: FrontierMath 评最终可验证答案,Hard2Verify 评证明过程中的局部正确性、支撑充分性和 first-error identification。