En
title: 英文全文 · Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math paper_id: 'ACL Anthology `2026.acl-long.1031' year: '2026'
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended
Frontier Math
Shrey Pandit* , Austin Xu* , Xuan-Phi Nguyen, Yifei Ming,
Caiming Xiong, Shafiq Joty
Salesforce AI Research
Data: https://huggingface.co/datasets/Salesforce/Hard2Verify
Code: https://github.com/SalesforceAIResearch/Hard2Verify
Abstract frontier LLM saturating new benchmarks, most re-
cently with GPT-5 Pro achieving 96.5%+ on AIME
Large language model (LLM)-based reason- 2024. As a result, recent efforts (Glazer et al.,
ing systems have recently achieved gold medal-
2024; Phan et al., 2025) have written novel, unseen
level performance in the IMO 2025 competi-
tion, writing mathematical proofs where, to re- mathematical questions to test LLMs.
ceive full credit, each step must be not only While the training approaches of closed frontier
correct but also sufficiently supported. To models remain a secret, open-source progress in
train LLM-based reasoners in such challeng- mathematical reasoning has been driven by scal-
ing, open-ended settings, strong verifiers capa- ing reinforcement learning from verifiable rewards
ble of catching step-level mistakes are neces- (RLVR) (Lambert et al., 2024), with the break-
sary prerequisites. We introduce Hard2Verify, a
through of DeepSeek-R1 (Guo et al., 2025) leading
human-annotated, step-level verification bench-
mark produced with over 500 hours of human
to an explosion of interest. This paradigm requires
labor. Hard2Verify is designed to rigorously training data with solutions that are easily verifi-
assess step-level verifiers at the frontier: Ver- able, i.e., have solutions that can be easily checked
ifiers must provide step-level annotations or against a known ground-truth by string matching
identify the first error in responses generated or symbolic checkers. Math benchmarks, for the
by frontier LLMs for very recent, challenging, most part, also adopt the verifiable setup, where
and open-ended math questions. We evaluate a model response is considered correct if its final
29 generative critics and process reward mod-
answer matches the ground-truth. Answer correct-
els, demonstrating that, beyond a few standouts,
open-source verifiers lag closed source models. ness, while a necessary condition for overall solu-
We subsequently analyze what drives poor per- tion correctness, is not sufficient: LLMs can pro-
formance in step-level verification, the impacts duce incorrect intermediate reasoning but conclude
of scaling verifier compute, as well as funda- with correct final answers (Lightman et al., 2023;
mental questions such as self-verification and Zheng et al., 2024a; Setlur et al., 2025).
verification-generation dynamics. The next frontier for LLMs is solving problems
that are hard to verify. A grand example of such a
1 Introduction problem is proving the Riemann hypothesis, where
the expected solution is not a short phrase, but a
Mathematical reasoning serves as a gold-standard multi-step proof. To verify correctness, each step
evaluation setting for benchmarking reasoning must be rigorously checked. Hints of open-ended
progress in large language models (LLMs). Over problem solving abilities already exist: advanced
the past half-decade, benchmarks have been intro- reasoning systems (OpenAI, 2025a; Google, 2025a;
duced to assess LLMs at the grade-school (Cobbe Huang and Yang, 2025) have achieved gold-level
et al., 2021), high-school (Hendrycks et al., 2021), performance in the 2025 IMO. Here, LLM outputs
university (Zhang et al., 2023), and competition were judged at the step-level by human experts who
math level (MMA, 2025; He et al., 2024a; Gao determined if steps are both correct and sufficiently
et al., 2024). However, the progress of mathemati- supported, with supporting lemmas and claims all
cal reasoning ability of LLMs has outpaced bench- appropriately stated and applied.
mark creation, with every subsequent release of a Training reasoning LLMs capable of open-ended
* Equal contribution. Correspondence:shrey.pandit@ problem solving requires scalable automatic evalu-
salesforce.com ation: Not every LLM rollout during RLVR train-
22502
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 22502–22517 July 2-7, 2026 ©2026 Association for Computational Linguistics 100 Skywork-PRM-1.5B Llama-3.3-70B-Instruct Qwen2.5-14B-Instruct Qwen2.5-Math-PRM-72B 80
Performance on
Skywork-PRM-7B Qwen2.5-7B-Instruct Qwen2.5-72B-Instruct Qwen2.5-Math-PRM-7B
GPT-5-High
60 Gemini 2.5 Pro
Hard2Verify
40
20
0
20 30 40 50 60 70 80 90 100
Performance on ProcessBench
Figure 1: Comparison of models evaluated on both ProcessBench (Zheng et al., 2024a) and our Hard2Verify benchmark. Past benchmarks do not sufficiently evaluate in the frontier-level math settings that Hard2Verify does; On the same error identification task, Qwen2.5-Math-PRM-72B performance drops from ProcessBench state-of-the-art at 78.3 to 37.3 on Hard2Verify.
ing can be audited by human experts. Rather, eval- justified or properly invoked. Third, Hard2Verify uation in open-ended settings requires step-level focuses on benchmarking verifiers in naturally oc- verifiers, typically process reward models (PRMs) curring application settings: Verifiers must assess or generative critic models. Such verifiers have al- model-written responses, which often differ dra- ready been used to provide dense process rewards matically from human-written reference answers. (Lightman et al., 2023; Shao et al., 2024; Zha et al., We benchmark 29 models spanning proprietary 2025). Furthermore, step-level verifiers are also models to open-weight models to PRMs. Com- used in many test-time scaling methods, selecting pared to past work, Hard2Verify represents a step the most promising candidate from multiple solu- up in difficulty, as shown in Fig. 1; Models capable tions or steps (Snell et al., 2024; Yu et al., 2025; of scoring 60%+ on ProcessBench (Zheng et al., Lifshitz et al., 2025; Zhou et al., 2025b). However, 2024a) are unable to crack 20% on Hard2Verify. are these step-level verifiers sufficient for pushing Our analysis reveals that this degraded performance the frontier of mathematical reasoning? is because weaker verifiers cannot identify mis- This work introduces Hard2Verify, which bench- takes, marking nearly every step as correct. We marks verifiers in assessing frontier LLM responses additionally analyze several fundamental questions to difficult, recent, and open-ended math prob- in step-level verification: How should one to scale lems. We curate challenging problems from re- verifier compute? What are the impacts of self- cent math competitions like IMO and Putnam, sam- verification? How much easier is generation than ple responses from three strong LLMs, GPT-5 verification for frontier models? (high) (OpenAI, 2025), Gemini 2.5 Pro (Google, 2025b), and Claude Sonnet 4 (thinking) (Anthropic, 2 Background and Related Work 2025), and employ PhD-level math experts to an- LLM-based verification. To meet demands for notate each model-generated step. The resulting scalable evaluation, LLM-based evaluators were benchmark is the culmination of 500+ hours of hu- first used in chat settings (Zheng et al., 2023). How- man effort encompassing three rounds of indepen- ever, as LLMs are deployed in challenging rea- dent agreement checks, yielding 1860 rigorously soning settings (Ke et al., 2025), recent work has graded steps across 200 unique model responses. shown the need for more capable reasoning eval- Beyond operating at the frontier, Hard2Verify uators (Frick et al., 2024; Tan et al., 2024; Zhou distinguishes itself from existing benchmarks for et al., 2025b). To get denser signals, focus quickly step-level annotation (Table 1). First, we empha- shifted to PRMs (Lightman et al., 2023) and syn- size collecting open-ended questions, with 78.5% thetic ways to curate step-level training data (Wang of our samples being open-ended. This way, veri- et al., 2023; Luo et al., 2024). However, Shao et al. fiers cannot “cheat” if they have seen the question (2024) showed that dense reward signals for pol- or ground-truth answer during training; rather ver- icy optimization brings limited improvement over ifiers must substantively assess step correctness. outcome-level baselines. This observation stems Second, step correctness is judged not only on cor- from the fact that PRMs only measure if a step rectness, but also based on whether all invoked could lead to a correct final answer, not whether results, such as supporting lemmas or claims, are the step is correct in any absolute sense. As a correctly stated and applied; saying “X follows result, generative verifiers (Mahan et al., 2024; from Y ” receives no credit if Y is not sufficiently Zhang et al., 2025a; Liu et al., 2025) have been 22503 Table 1: Comparison between Hard2Verify and existing step-level math benchmarks.
Question Open-Ended Natural Generator Step-Level
Annotator
Difficulty Responses? Responses? Strength Labels?
MR-GSM8K (Zeng et al., 2023) Easy ✗ ✓ Weak Human ✓
MR-MATH (Xia et al., 2025) Easy ✗ ✓ Weak Human ✓
MR-Ben (Zeng et al., 2023) Easy ✗ ✓ Weak Human ✓
ProcessBench (Zheng et al., 2024a) Easy-Hard 10.3% ✓ Weak-Medium Human ✗
PRMBench (Song et al., 2025) Easy ✗ ✗ Weak Synth. + Human Check ✓
Hard2Verify (Ours) Hard 78.5% ✓ Strong Human ✓
Step-level Breakdown Response-level Breakdown
stance and injected errors may not represent natu-
Gemini
(461) Gemini
(57) Claude
(72)
rally occurring errors in generation. Hard2Verify,
Claude
(858) in contrast, operates at the current frontier, tasking
verifiers to evaluate responses from frontier-level
GPT-5
(541) Correct
Incorrect GPT-5
(71)
LLMs to difficult, largely open-ended questions.
Figure 2: Breakdown of correct vs. incorrect steps (left) and responses (right) by model. We consider a response 3 The Hard2Verify Benchmark incorrect if any step in the response is labeled incorrect.
deployed. This allows for more precise descrip- 3.1 Design philosophy tion of evaluation criteria and increased inference- Hard2Verify is designed to test verifiers at the fron- time compute. Generative approaches are either tier of LLM-based math reasoning. At the question, reference-based (Luong et al., 2025; Ma et al., response, and annotation level, Hard2Verify is cu- 2025) or reference-free (Shao et al., 2025; Rahman rated based on the following philosophy: et al., 2025; Xu et al., 2025). The former requires detailed grading rubrics; As a result, Hard2Verify • Questions. To measure progress in step-level operates in the latter, more scalable setting. verification, we must characterize how verifiers Benchmarking step-level math verifiers. Ta- perform on extremely difficult, open-ended math ble 1 contrasts Hard2Verify with related bench- questions. Open-ended problems represent the marks. MR-GSM8K (Zeng et al., 2023) annotate next frontier of mathematical reasoning, one model responses to GSM8K (Cobbe et al., 2021) where verifiers become increasingly important in questions on a per-step basis to evaluate generative lieu of available ground-truth answers. We focus models as evaluators. MR-MATH (Xia et al., 2025) our data collection on very recent mathematical and MR-Ben (Zeng et al., 2024) are similar, us- Olympiads, prioritizing open-ended questions. ing slightly harder sources like MATH (Hendrycks • Model responses. The responses that verifiers et al., 2021) and MMLU (Hendrycks et al., 2020). evaluate must be from highly capable, frontier- The two most relevant works to Hard2Verify are level models. To push the frontier of math rea- ProcessBench (Zheng et al., 2024a) and PRM- soning, verifiers must be able to tell when the Bench (Song et al., 2025). ProcessBench uses most powerful models make potentially subtle a mix of easy (GSM8K and MATH) and hard mistakes. Moreover, such mistakes should be (OlympiadBench and Omni-MATH) questions, but naturally occurring, i.e., arise from the model mostly contains samples that are not open-ended1 . generation process. We do not inject or edit ex- Further, ProcessBench only evaluates first error isting correct model-or human-written solutions. identification ability of verifiers. PRMBench ob- This is meant to closely approximate the response tains step-level annotations by taking fully correct distribution that verifiers will see “in the wild”, human-written and model-generated solutions from as they are applied in frontier math settings. the now easy PRM800K dataset and injecting er- • Annotation process. We employ a strict view rors with an LLM, yielding responses that are not of response grading: Any step that contains a naturally occurring: Human-and model-written mistakes or is derived from a previous mistake text may have large differences in style and sub- is considered incorrect, i.e., we do not employ “Error Carried Forward” grading. This is inspired 1 The fraction of open-ended questions in ProcessBench by competitive math settings, the entire solution in Table 1 is derived by counting the number questions from must be correct to receive full points. the Omni-MATH split that are not in the rule-based Omni- MATH. All other splits are not open-ended. Based on this philosophy, we create Hard2Verify. 22504 Correct by Step Incorrect by Step 3.2 Curating hard questions 175 150 Overall Correct GPT-5 Correct Claude Sonnet 4 Correct Overall Incorrect GPT-5 Incorrect Claude Sonnet 4 Incorrect 125 Gemini 2.5 Pro Correct Gemini 2.5 Pro Incorrect
We construct our benchmark by collecting problem 100 Count 75 50
statements and official solutions (Q, Aofficial ) from 25 0 0.0 2.5 5.0 7.5 10.0 12.5 15.0 17.5 20.0 0.0 2.5 5.0 7.5 10.0 12.5 15.0 17.5 20.0
leading math competitions including the IMO, Put- Step Step
nam, and INMO; We provide a full list of sources Figure 3: Count of correct (left) and incorrect (right) in App. A. We focus question curation on recent labels by model solution step. Models tend to begin (2024 and beyond) Olympiad-level math competi- solutions correctly, but get derailed after a few steps. tions. For each Olympiad, we parse the official from annotators, and finetuned evaluation instruc- PDFs using MathPix and extract all content in tions accordingly. We then performed annotations LATEX to preserve mathematical typography and en- in batches of samples, performing spot-checks of sure stable equation rendering. We exclude image- samples as they became available. This is in addi- dependent problems and only keep questions that tion to internal processes at Turing, which include could be solved using textual information. The initial human annotation and three rounds of hu- resulting question set comprises 80 frontier-level man review, where annotations were reviewed for problems from 10 distinct Olympiads. correctness and guideline alignment. Overall, this 3.3 Response generation process represents over 500 hours of manual hu- man labor. See App. E for more annotation details. Using our curated question pool, we sample re- sponses from three frontier LLMs: GPT-5 (with 3.5 Overall dataset statistics. high reasoning), Gemini 2.5 Pro, and Claude Son- Our annotation process yields 1,860 unique model net 4 (Thinking). We employ a standardized steps annotated across 200 model solutions. 58% prompt (App. D), instructing models to produce (1,080/1,860) steps are labeled correct, while the exam-style, stepwise proofs that mirror how an remaining 780 are labeled incorrect. Fig. 2 shows Olympiad participant would structure a solution. how models perform on a step-level and problem We use the same prompt and decoding settings level. We consider a model response correct if all across models and disable access to external tools, steps in the solution are graded correct by humans. like web search or code interpreters. Each model Claude Sonnet 4 takes the most steps but gets the produces a single solution per problem, which we least percentage of steps correct, whereas GPT-5 record for annotation. These samples are challeng- and Gemini 2.5 Pro perform similarity in terms of ing; for example, Gemini 2.5 Pro takes up to 15 step-level accuracy. However, at the response level, minutes to return a solution via API access. After GPT-5 outperforms Gemini 2.5 Pro by larger mar- curating all model responses to all questions, we gins. Claude Sonnet 4, while achieving over 50% filter out responses with undesirable qualities, such step-level accuracy, fails to string correct steps to- as a small number of long, dense steps or responses gether, only producing 4 entirely correct solutions with degenerate outputs. This leaves us with a com- out of 72. Fig. 3 visualizes how errors appear as a pact but high quality set of 200 responses. function of steps, with all three models following similar trends: Errors tend to occur in the middle 3.4 Ensuring high-quality annotations of solutions, appearing after a few steps. After sampling responses to our curated ques- tions, human annotators meticulously annotate 3.6 Evaluation tasks each model solution step-by-step. We partnered Our step-level annotations allow us to construct with Turing, a research accelerator. Turing employs three distinct tasks: (1) Step-level correctness mathematical experts, with a super-majority of our (Step-Level), (2) Response-level correctness annotators having an advanced graduate level ed- (Response-Level), and (3) First error identifica- ucation in mathematics. To ensure consistent and tion (ErrorID). The Step-Level task corresponds high quality evaluations, we provided comprehen- to the setup in Song et al. (2025), whereas the sive annotation instructions as well official solu- ErrorID tasks corresponds to that of Zheng et al. tions Aofficial as references. Annotation began with (2024a). As we show in § 4, both tasks are chal- a multi-round pilot study, where we hand-annotated lenging settings for current verifiers. We provide three model responses, then worked together with our evaluation prompts in App. D. annotators to review samples, solicited feedback Step-Level. Here, the verifier is tasked with de- 22505 termining the correctness of each step. Generative Balanced F1 both serve as aggregate measures: the verifiers are prompted to output a binary yes/no former reflects average performance across both label for each step, whereas PRM step-level scores modes, while the latter penalizes imbalanced per- are converted to binary labels via a fixed threshold. formance. An ideal verifier scores highly on both. Response-Level. Here, we derive an outcome- based task from Step-Level labels which reflects 4.2 Evaluated models strict grading of open-ended math problems: For a question to be correct, all steps in the solution must We select a variety of PRMs and generative mod- be deemed correct. Therefore, if any step in the els prompted as step-level critics. For prompted solution is incorrect, the solution is wrong2 . From critics, we test a closed-source models as well as human labels, we create an overall response-level large (≥ 70B) and small-medium (<70B) open- correctness label. Likewise, we create a response- weight models. We evaluate all reasoning models level prediction from step-level verifier predictions. at the maximum reasoning level (e.g., “high” for This task is more forgiving than Step-Level: Ex- GPT-5), using suggested sampling parameters for act step labels need not match exactly for a verifier various baselines. All Qwen3 models are evaluated to agree with a human at the response level. with “thinking” on. For instruction-tuned models, ErrorID. Here, the verifier is prompted to output we use greedy decoding. For all models, we set the first step that contains a mistake, if present, the maximum number of output tokens to be 32K. or step −1, corresponding to “No error”. For The full set of models is enumerated in App. F. generative verifiers, the first error step can also For PRMs, we select Qwen2.5-Math-PRM (Zhang be derived from Step-Level labels, similar to et al., 2025b), Skywork-PRM (He et al., 2024b), the Response-Level setting. Following Process- ReasonFlux-PRM (Zou et al., 2025), and Universal- Bench, we prompt the verifier to output the step PRM (Tan et al., 2025). We tune PRM thresholds index directly, allowing us to more directly com- following Zheng et al. (2024a); See App. F.1. pare across benchmarks; we quantify differences between the direct prompting approach and convert- 4.3 Main evaluation results ing from step-level labels in § 4.3. For PRMs, we select the first step below the correctness threshold. Table 2 presents our main results, with detailed re- sults presented in App. C. Among proprietary mod- els, GPT-5 stands out in its overall ability across 4 Experiments all three tasks. Gemini 2.5 Pro follows closely for 4.1 Evaluation Metrics step-level identification, but lags in error identifi- cation. Finally, Claude Sonnet 4 with Thinking Let TPR and TNR denote the True Positive Rate lags OpenAI models and Gemini 2.5 Pro, failing to and True Negative Rate, i.e., verifier accuracy on match reasoning models from previous generations, correct and incorrect samples, respectively. We de- like o3 and o4-mini. Among larger open-weight fine Balanced Accuracy as the mean and Balanced models, the gpt-oss series are clear standouts, with F1 Score as the harmonic mean of TPR and TNR3 : gpt-oss-120B roughly matching the performance 2 TPR · TNR of o3. Larger Qwen3 models and DeepSeek-R1 Balanced F1 Score = , (1) TPR + TNR challenge for second place on Step-Level, but all We report Balanced Accuracy and Balanced F1 lag on ErrorID. Notably, Llama-3.3-70B, which Score for all tasks. The ground-truth labels achieves 58.0 on ProcessBench (Fig. 1) achieves and model predictions vary based on task. For only 2.50 on ErrorID. Among smaller models, Step-Level, we aggregate all steps and all veri- gpt-oss-20B performs extremely well on step-level fier predictions across all responses, whereas for and response-level tasks, but falters in identify- Response-Level and ErrorID, we compute met- ing errors. ByteDance Seed-OSS-36B and Qwen3- rics at the response level. Balanced Accuracy and 30B-A3B are the next best performers, but only 2 Because we are concerned with ensuring completely cor- ByteDance Seed-OSS-36B is able to outperform rect responses, we apply this procedure to all responses/ques- random guessing performance on ErrorID. Finally, tions in Hard2Verify, including non-open-ended questions. even strong PRMs perform below random guess 3 This is equivalent to the ProcessBench “F1 Score”, which differs from the typical F1 Score by using TNR instead of performance on ErrorID. For example, Qwen2.5- precision. To avoid confusion, we use “Balanced F1 Score”. Math-PRM-72B achieves only 37.28 Balanced F1. 22506 Table 2: Main evaluation results on Hard2Verify across our three evaluation tasks (§ 3.6). We report Balanced Accuracy and Balanced F1 Score. Best and second-best scores in each category marked.
Step-Level Response-Level ErrorID
Bal. Accuracy Bal. F1 Bal. Accuracy Bal. F1 Bal. Accuracy Bal. F1
Generative Critics, proprietary models
GPT-5 86.53 85.83 89.69 89.52 70.61 69.72
Gemini 2.5 Pro 83.37 83.09 85.73 85.46 52.46 52.46
Claude Sonnet 4 70.61 60.37 78.24 73.44 53.45 39.30
GPT-5-Mini 81.06 78.73 81.93 81.92 65.96 60.04
o3 78.70 75.29 83.21 82.58 60.32 57.31
o4-Mini 74.90 68.09 83.94 81.71 67.31 57.62
GPT-4.1 56.17 24.66 58.94 33.55 52.44 21.29
Generative Critics, large (≥ 70B) models
Kimi K2 61.79 42.83 65.34 51.66 49.10 31.40
DeepSeek-R1 68.92 62.30 73.95 72.75 54.23 45.35
Qwen3-235B-A22B 72.55 64.03 79.42 77.87 60.90 50.78
Qwen3-Next-80B-A3B 67.91 54.69 75.08 68.31 58.29 43.05
Qwen2.5-72B-Instruct 56.01 26.36 61.06 46.89 26.49 16.38
GLM-4.5-Air 57.40 29.40 61.78 41.00 41.97 17.81
gpt-oss-120B 78.10 74.64 83.92 83.71 63.97 60.64
Llama-3.3-70B-Instruct 54.28 18.37 57.04 28.16 49.44 2.50
Generative Critics, small/medium (< 70B) models
Qwen3-32B 63.99 51.77 67.86 63.16 51.96 26.83
Qwen3-30B-A3B 70.71 61.91 73.79 71.02 58.83 50.51
ByteDance Seed-OSS-36B 66.79 53.09 72.54 63.88 59.24 45.18
gpt-oss-20B 75.18 70.93 83.85 83.32 46.13 45.28
Qwen3-14B 65.48 52.91 74.59 70.12 53.69 37.33
Qwen3-8B 65.26 53.51 77.61 72.45 45.92 34.26
Qwen2.5-14B-Instruct 60.45 47.59 63.40 63.23 43.47 18.86
Qwen2.5-7B-Instruct 48.82 22.84 55.67 44.18 29.75 15.96
Process Reward Models, open-source models
Qwen2.5-Math-PRM-72B 55.82 35.50 66.80 64.91 41.80 37.28
Qwen2.5-Math-PRM-7B 57.56 42.37 63.08 57.57 35.03 32.50
Skywork-PRM-7B 38.52 34.12 56.77 29.81 11.56 8.36
Skywork-PRM-1.5B 40.81 12.94 52.46 20.89 8.62 7.48
ReasonFlux-PRM-7B 53.09 22.40 55.89 53.82 42.48 28.71
UniversalPRM-7B 64.17 60.27 54.74 41.46 26.08 25.97
Table 3: ErrorID performance using two prompting approaches, with ∆ = Step-Level − ErrorID. The ErrorID prompt tasks verifier to directly identify the first step with an error, as in ProcessBench (Zheng et al., 2024a). The Step-Level prompt tasks the verifier to produce step-level labels, from which the first error step is derived. Balanced Accuracy tends to improve with the Step-Level prompt, but Balanced F1 changes are mixed.
ErrorID Prompt Step-Level Prompt ErrorID Prompt Step-Level Prompt
∆Bal. Acc ∆Bal. F1
Bal. Accuracy Bal. Accuracy Bal. F1 Bal. F1
GPT-5 70.61 76.72 +6.11 69.72 75.66 +5.94
gpt-oss-120B 63.97 69.68 +5.71 60.64 64.81 +4.17
GPT-5-Mini 65.96 66.43 +0.47 60.04 63.25 +3.21
o4-Mini 67.31 67.16 -0.15 57.62 53.35 -4.27
o3 60.32 68.02 +7.70 57.31 60.61 +3.30
Gemini 2.5 Pro 52.46 66.11 +13.65 52.46 62.78 +10.32
Qwen3-235B-A22B 60.90 65.17 +4.27 50.78 55.35 +4.57
Qwen3-30B-A3B 58.83 60.19 +1.36 50.51 47.25 -3.26
DeepSeek-R1 54.23 61.53 +7.30 45.35 52.02 +6.67
gpt-oss-20B 46.13 66.44 +20.31 45.28 57.75 +12.47
ByteDance Seed-OSS-36B 59.24 58.94 -0.30 45.18 33.55 -11.63
Qwen3-Next-80B-A3B 58.29 63.37 +5.08 43.05 44.85 +1.80
Claude Sonnet 4 53.45 60.83 +7.38 39.30 38.59 -0.71
Qwen3-14B 53.69 56.56 +2.87 37.33 33.25 -4.08
Qwen3-8B 45.92 57.35 +11.43 34.26 29.09 -5.17
Kimi K2 49.10 54.26 +5.16 31.40 23.33 -8.07
Qwen3-32B 51.96 52.35 +0.39 26.83 31.09 +4.26
GPT-4.1 52.44 51.97 -0.47 21.29 11.89 -9.40
Qwen2.5-14B-Instruct 43.47 40.30 -3.17 18.86 23.04 +4.18
GLM-4.5-Air 41.97 53.24 +11.27 17.81 16.25 -1.56
Qwen2.5-72B-Instruct 26.49 48.09 +21.60 16.38 10.72 -5.66
Qwen2.5-7B-Instruct 29.75 43.01 +13.26 15.96 9.53 -6.43
Llama-3.3-70B-Instruct 49.44 50.71 +1.27 2.50 7.31 +4.81
22507
Model Performance by Reasoning Effort 82.96 85.83 80 74.64 76.34 70.93
Balanced F1 Score
100
59.69 61.83 61.46 64.20
60
80
Percentage 40 60
40 Balanced F1 20
TPR 0
20 TNR Low Medium High Low Medium High Low Medium High
gpt-oss-20B gpt-oss-120B GPT-5
G t-o o3
ss-
min PT-5
i2 Qw en o4-m t-o B 12
ss- 0
20 B
90
.5 3-2 35 ini
De 22
ep B-A B
GP Pro
T-5
Qw See
en k-R
3
Cla -30B
Qw ude -A3B 1
gpt-oss-20B, Low, N=1 (59.69) gpt-oss-20B, High (70.93)
-M en 3-N onn S
ex et 4
t-8 0B -A
ini
80
gp Qw 3B
Se en3-
ed 8B
Step-level Balanced F1
Ge gp -O
Qw 36B
en
Qw 14B
en
Qw 3-32
SS -
3-
en B
2.5 -14
70
Qw en LM- 2 G iK
2.5 4.5
-72 -A Kim B
Qw en
Lla .5- PT-4
2
ma 7B .1 B-I ir
Gns tru ct
-3. -Ins
3-7 tru
0B ct
60
-In str uc t
50
40
Figure 4: Weaker models are unable to find mistakes, 30 20
eventually considering all steps correct: TNR tends 10 0 4 6 8 10 12 14 16 toward 0 while TPR tends towards 1. N for Best-of-N
What separates strong and weak verifiers? To Figure 5: Top: Scaling inference-time compute sequen- tially leads to higher performance in GPT-5 and gpt-oss provide additional insights into variations across models Bottom: Parallel decoding has little effect on different verifiers, Fig. 4 plots the TPR and TNR step-level F1 performance for gpt-oss-20B, failing to for all generative critics models, sorted in perfor- bridge the gap vs. gpt-oss-20B at high-reasoning effort. mance from strongest (left) to weakest (right) in terms of Balanced F1 Score. A clear trend emerges: to benefit the most, while insufficiently capable Verifier performance degrades because TNR drops models fare worse. As the change in performance quickly to near 0, while TPR rises gradually to is mixed across models, we advise practitioners almost 1. This indicates that almost all steps are la- optimize prompts on a per-verifier basis. beled as correct, revealing that weaker verifiers can- not catch errors. Notably, the order of models from 5 Additional Analysis left to right approximately correlates with mathe- 5.1 How should we scale verifier matical generation ability, i.e., the ability to solve inference-time compute? extremely difficult math problems. As such, this may indicate that a baseline level of solving ability Here, we scale verifier inference-time compute se- is a necessary prerequisite for verification. App. C quentially and in parallel. We find sequential scal- shows this trend holds similarly for other tasks. ing brings substantive gains unlike parallel scaling. Sequential inference-time compute scaling. To identify errors, how should verifiers be Here we explore scaling inference-time compute se- prompted? Our ErrorID task adopts the setup of quentially by letting the verifier generate more out- ProcessBench (Zheng et al., 2024a), which prompts put tokens, focusing on the Step-Level task. We the verifier to output the index of the first step with use gpt-oss-20B, gpt-oss-120B, and GPT-5, which an error. However, the first error index can also all have low, medium, and high reasoning levels. be derived from step-level labels, like those pro- In Fig. 5 (top), we plot Balanced F1. Affording the duced in the Step-Level task. In Table 3, we com- verifier to generate more “thinking” tokens at in- pare the performance under the ErrorID Prompt ference time generally improves performance, with and Step-Level Prompt. Surprisingly, directly gpt-oss-120B improving the most from low (61.46) prompting for the given task may not yield the to high (74.64) and gpt-oss-20B likewise improv- best performance: In terms of Balanced F1, perfor- ing significantly. Gains for GPT-5 are smaller com- mance across models is mixed, with some models pared to gpt-oss models, but still significant, with exhibiting very small performance changes and 12.3% relative improvement from low to high. others exhibiting significant changes. For example, Parallel inference-time compute scaling. Here, step-level prompting significantly degrades perfor- we attempt to match the performance of gpt-oss- mance for ByteDance Seed-OSS-36B from 45.18 20B at high reasoning effort by sampling parallel to 33.55, while boosting performance for Gemini outputs from gpt-oss-20B at low reasoning effort. 2.5 Pro from 52.46 to 62.78. Overall, we find that We sample 32 responses per sample from gpt-oss- more capable models, like GPT-5 and Gemini 2.5 20B and simulate best of N from N = 4, . . . , 16 Pro, benefit the most from switching to deriving via bootstrap sampling. Concretely, for each N , we first identified error from Step-Level outputs. We sample N responses from the 32 without replace- hypothesize that requiring step-by-step annotations ment, and aggregate predicted step-level labels via requires models inspect each step more carefully, majority vote, breaking ties arbitrarily. To reduce allowing for better error identification. Models variance, we repeat this process for 10 trials for capable of performing step-level verification tend each N , and report mean and standard deviation 22508 TPR TNR 1.0 80 98
GPT-5 94.32 92.14 96.20 GPT-5 80.50 74.03 79.40 0.8
Fraction of correctly
96 70
94 60 0.6
verified steps
Gemini 2.5 Pro
Verifier
Verifier 87.55 92.86 85.09 92 80.50 74.59 78.39 Gemini 2.5 Pro 50 0.4 Verification Verification 90 is easier is harder 40 0.2
Claude Sonnet 4 Claude Sonnet 4
88
95.85 98.93 98.54 61.75 24.86 25.13 Claude Gemini GPT-5 y=x
86 30 0.0
0.0 0.2 0.4 0.6 0.8 1.0
Claude Sonnet 4 Gemini 2.5 Pro GPT-5 Claude Sonnet 4 Gemini 2.5 Pro GPT-5 Fraction of correctly generated steps
Generator Generator
Figure 6: Verifier TPR and TNR based on generator Figure 7: Each generator evaluates self-produced re- model. For strong verifiers (GPT-5, Gemini 2.5 Pro), sponses, and the fraction of steps correctly solved vs. TPR varies based on generator, with GPT-5 being the fraction of steps correctly verified for a given question most stable. Claude Sonnet 4 generates the easiest to is plotted. In general, models are more successful in catch mistakes, whereas Gemini 2.5 Pro produces the catching mistakes than generating error-free responses. hardest to catch mistakes, as measured by TNR. responses from Gemini 2.5 Pro, showing that GPT- across trials Fig. 5 (bottom). We also plot the base- 5 is more reliable in self-critique than Gemini 2.5 line gpt-oss-20B performance at low and high rea- Pro is. The fact that Gemini 2.5 Pro has the lowest soning efforts. Surprisingly, Best-of-N does not TNR on self-generated responses is consistent with meaningfully improve over sampling 1 response as recent work analyzing self-reflection (Stechly et al., N increases. An intuitive explanation for this phe- 2023, 2024; Huang et al., 2023), where LLMs were nomenon is that step-level verification is inherently shown to have difficulties correcting their own mis- a sequential task: Each step must be processed one- takes in challenging reasoning settings. In contrast, after-another. As such, affording the verifier more Claude Sonnet 4 as a relatively weaker verifier can- time to “think” about each step is more effective not identify errors in stronger model responses. than sampling multiple “rushed” judgments. 5.3 Is verifying easier than solving? 5.2 How do verifiers verify their own Here, we examine if generating a solution is eas- responses? ier than verifying the same solution. We split Hard2Verify into three subsets corresponding to We investigate the dynamics of self-verification, each of the three generator models and have the focusing on GPT-5, Gemini 2.5 Pro, and Claude generators verify their own responses. For each Sonnet 4 as verifiers. Fig. 6 plots the step-level TPR response, we record the fraction of correctly gen- and TNR performance based on response generator. erated steps (“solve rate”), as deemed by human The results notably depend on verifier strength: Ta- annotators, and the fraction of correctly verified ble 2 shows that GPT-5 and Gemini 2.5 Pro are the steps (“verification rate”), as deemed by agreement top two performers, whereas Claude Sonnet 4 is a with human annotators. In Fig. 7, we plot the veri- relatively weak proprietary verifier. We find that fication rate against the solve rate. We observe that GPT-5 and Gemini 2.5 Pro as verifiers are more the verification rate is consistently higher than the likely to consider a correct self-generated response solve rate across all models; Only on a few prob- as correct, as measured by TPR. Of the two, GPT- lems does the verifier have a more difficult time 5 exhibits the least variation in TPR across mod- verifying a problem than generating the problem. els, while Gemini 2.5 Pro performance drops from This result offers some optimism for future work 92.86 TPR on own-generated responses to as low as in verification: Because verifying a solution tends 85.09 TPR for GPT-5-generated responses. Claude to be “easier” than generating the solution, veri- Sonnet 4, on the other hand, overwhelmingly as- fiers may not necessarily need to be as powerful as signs “Correct” as a label, leading to high TPRs frontier generators to reliably identify errors. regardless of generator. Across all threee models, it is easier to identify errors from the weaker model 6 Conclusion (Claude Sonnet 4) than it is to identify errors from the stronger models. This result is consistent with We introduce Hard2Verify, a human-annotated, recent work (Zhou et al., 2025a) studying verifica- step-level benchmark aimed to assess how step- tion, which finds weaker generators produce eas- level verifiers operate in frontier settings. We focus ier to catch errors. Interestingly, both GPT-5 and our data curation on recent open-ended math prob- Gemini 2.5 Pro struggle have the lowest TNR on lems, sampling responses from frontier LLMs. The 22509 end result of over 500 hours of human annotation https://deepmind.google/discover/blog/advanced- effort is a benchmark that challenges many current version-of-gemini-with-deep-think-officially- achieves-gold-medal-standard-at-the-international- open-source verifiers, which are unable to match mathematical-olympiad/. the performance of larger, proprietary models. Google. 2025b. Gemini 2.5: Our most intelligent 7 Limitation ai model. https://blog.google/technology/google- deepmind/gemini-model-thinking-updates-march- Our work has some limitations. First, Hard2Verify 2025/. is modest in scale (200 model-generated solutions). Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, This is a consequence of careful filtering and the Abhinav Pandey, Abhishek Kadian, Ahmad Al- use of a strict cutoff date during data collection, Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, which may lead to some variance in measured per- Alex Vaughan, and 1 others. 2024. The llama 3 herd formance. Second, the benchmark is currently lim- of models. arXiv preprint arXiv:2407.21783. ited to English, Olympiad-style problems and text- Daya Guo, Dejian Yang, Haowei Zhang, Junxiao only inputs. In future work, we plan to expand Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shi- the dataset and analysis to a larger scale, include rong Ma, Peiyi Wang, Xiao Bi, and 1 others. 2025. problems in additional languages, and incorporate Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint diagram-based reasoning. arXiv:2501.12948.
Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding
References Hu, Zhen Leng Thai, Junhao Shen, Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, and 1 oth- Sandhini Agarwal, Lama Ahmad, Jason Ai, Sam Alt- ers. 2024a. Olympiadbench: A challenging bench- man, Andy Applebaum, Edwin Arbus, Rahul K mark for promoting agi with olympiad-level bilin- Arora, Yu Bai, Bowen Baker, Haiming Bao, and 1 gual multimodal scientific problems. arXiv preprint others. 2025. gpt-oss-120b & gpt-oss-20b model arXiv:2402.14008. card. arXiv preprint arXiv:2508.10925. Jujie He, Tianwen Wei, Rui Yan, Jiacai Liu, Chaojie Anthropic. 2025. Introducing claude 4. Wang, Yimeng Gan, Shiwen Tu, Chris Yuhao Liu, Liang Zeng, Xiaokun Wang, Boyang Wang, Yong- Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, cong Li, Fuxiang Zhang, Jiacheng Xu, Bo An, Yang Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Liu, and Yahui Zhou. 2024b. Skywork-o1 open se- Plappert, Jerry Tworek, Jacob Hilton, Reiichiro ries. Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word prob- Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, lems. arXiv preprint arXiv:2110.14168. Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language under- Evan Frick, Tianle Li, Connor Chen, Wei-Lin Chiang, standing. arXiv preprint arXiv:2009.03300. Anastasios N Angelopoulos, Jiantao Jiao, Banghua Zhu, Joseph E Gonzalez, and Ion Stoica. 2024. How Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul to evaluate reward models for rlhf. arXiv preprint Arora, Steven Basart, Eric Tang, Dawn Song, and Ja- arXiv:2410.14872. cob Steinhardt. 2021. Measuring mathematical prob- lem solving with the math dataset. arXiv preprint Bofei Gao, Feifan Song, Zhe Yang, Zefan Cai, Yibo arXiv:2103.03874. Miao, Qingxiu Dong, Lei Li, Chenghao Ma, Liang Chen, Runxin Xu, and 1 others. 2024. Omni- Jie Huang, Xinyun Chen, Swaroop Mishra, math: A universal olympiad level mathematic bench- Huaixiu Steven Zheng, Adams Wei Yu, Xiny- mark for large language models. arXiv preprint ing Song, and Denny Zhou. 2023. Large language arXiv:2410.07985. models cannot self-correct reasoning yet. arXiv preprint arXiv:2310.01798. Elliot Glazer, Ege Erdil, Tamay Besiroglu, Diego Chicharro, Evan Chen, Alex Gunning, Caroline Falk- Yichen Huang and Lin F Yang. 2025. Gemini 2.5 pro man Olsson, Jean-Stanislas Denain, Anson Ho, capable of winning gold at imo 2025. arXiv preprint Emily de Oliveira Santos, and 1 others. 2024. Fron- arXiv:2507.15855. tiermath: A benchmark for evaluating advanced mathematical reasoning in ai. arXiv preprint Zixuan Ke, Fangkai Jiao, Yifei Ming, Xuan-Phi Nguyen, arXiv:2411.04872. Austin Xu, Do Xuan Long, Minzhi Li, Chengwei Qin, Peifeng Wang, Silvio Savarese, and 1 others. 2025. Google. 2025a. Advanced version of gemini with A survey of frontiers in llm reasoning: Inference scal- deep think officially achieves gold-medal stan- ing, learning to reason, and agentic systems. arXiv dard at the international mathematical olympiad. preprint arXiv:2504.09037. 22510 Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying OpenAI. 2025a. Openai imo 2025 proofs. Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. https://github.com/aw31/openai-imo-2025-proofs. Gonzalez, Hao Zhang, and Ion Stoica. 2023. Effi- cient memory management for large language model OpenAI. 2025b. Openai o3 and o4-mini system card. serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Principles. Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, and 1 oth- Nathan Lambert, Jacob Morrison, Valentina Pyatkin, ers. 2025. Humanity’s last exam. arXiv preprint Shengyi Huang, Hamish Ivison, Faeze Brahman, arXiv:2501.14249. Lester James V Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, and 1 others. 2024. Tulu 3: Pushing fron- Salman Rahman, Sruthi Gorantla, Arpit Gupta, Swastik tiers in open language model post-training. arXiv Roy, Nanyun Peng, and Yang Liu. 2025. Spark: preprint arXiv:2411.15124. Stepwise process-aware rewards for reference- free reinforcement learning. arXiv preprint Shalev Lifshitz, Sheila A McIlraith, and Yilun Du. arXiv:2512.03244. 2025. Multi-agent verification: Scaling test-time compute with multiple verifiers. arXiv preprint Amrith Setlur, Chirag Nagpal, Adam Fisch, Xinyang arXiv:2502.20379. Geng, Jacob Eisenstein, Rishabh Agarwal, Alekh Agarwal, Jonathan Berant, and Aviral Kumar. 2025. Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harri- Rewarding progress: Scaling automated process veri- son Edwards, Bowen Baker, Teddy Lee, Jan Leike, fiers for LLM reasoning. In The Thirteenth Interna- John Schulman, Ilya Sutskever, and Karl Cobbe. tional Conference on Learning Representations. 2023. Let’s verify step by step. In The Twelfth Inter- national Conference on Learning Representations. Zhihong Shao, Yuxiang Luo, Chengda Lu, ZZ Ren, Jiewen Hu, Tian Ye, Zhibin Gou, Shirong Ma, and Zijun Liu, Peiyi Wang, Runxin Xu, Shirong Ma, Xiaokang Zhang. 2025. Deepseekmath-v2: To- Chong Ruan, Peng Li, Yang Liu, and Yu Wu. 2025. wards self-verifiable mathematical reasoning. arXiv Inference-time scaling for generalist reward model- preprint arXiv:2511.22570. ing. arXiv preprint arXiv:2504.02495. Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Liangchen Luo, Yinxiao Liu, Rosanne Liu, Samrat Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Phatale, Harsh Lara, Yunxuan Li, Lei Shu, Yun Zhang, YK Li, Yang Wu, and 1 others. 2024. Zhu, Lei Meng, Jiao Sun, and 1 others. 2024. Im- Deepseekmath: Pushing the limits of mathematical prove mathematical reasoning in language models reasoning in open language models. arXiv preprint by automated process supervision. arXiv preprint arXiv:2402.03300. arXiv:2406.06592. Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Ku- Minh-Thang Luong, Dawsen Hwang, Hoang H Nguyen, mar. 2024. Scaling llm test-time compute optimally Golnaz Ghiasi, Yuri Chervonyi, Insuk Seo, Junsu can be more effective than scaling model parameters. Kim, Garrett Bingham, Jonathan Lee, Swaroop arXiv preprint arXiv:2408.03314. Mishra, and 1 others. 2025. Towards robust mathe- matical reasoning. In Proceedings of the 2025 Con- Mingyang Song, Zhaochen Su, Xiaoye Qu, Jiawei Zhou, ference on Empirical Methods in Natural Language and Yu Cheng. 2025. Prmbench: A fine-grained Processing, pages 35406–35430. and challenging benchmark for process-level reward models. arXiv preprint arXiv:2501.03124. Wenjie Ma, Andrei Cojocaru, Neel Kolhe, Bradley Louie, Robin Said Sharif, Haihan Zhang, Vincent Kaya Stechly, Matthew Marquez, and Subbarao Kamb- Zhuang, Matei Zaharia, and Sewon Min. 2025. Reli- hampati. 2023. Gpt-4 doesn’t know it’s wrong: An able fine-grained evaluation of natural language math analysis of iterative prompting for reasoning prob- proofs. arXiv preprint arXiv:2510.13888. lems. arXiv preprint arXiv:2310.12397.
Dakota Mahan, Duy Van Phung, Rafael Rafailov, Kaya Stechly, Karthik Valmeekam, and Subbarao Kamb- Chase Blagden, Nathan Lile, Louis Castricato, Jan- hampati. 2024. On the self-verification limitations Philipp Fränken, Chelsea Finn, and Alon Albalak. of large language models on reasoning and planning 2024. Generative reward models. arXiv preprint tasks. arXiv preprint arXiv:2402.08115. arXiv:2410.12832. Sijun Tan, Siyuan Zhuang, Kyle Montgomery, MMA. 2025. (american invitational mathematics exam- William Y Tang, Alejandro Cuadron, Chenguang ination). https://maa.org. Wang, Raluca Ada Popa, and Ion Stoica. 2024. Judgebench: A benchmark for evaluating llm-based OpenAI. 2025. Gpt-5 system card. judges. arXiv preprint arXiv:2410.12784.
OpenAI. 2025. Introducing GPT-4.1 in the api. https: Xiaoyu Tan, Tianchu Yao, Chao Qu, Bin Li, Minghao //openai.com/index/gpt-4-1/. Accessed: 2025- Yang, Dakuan Lu, Haozhe Wang, Xihe Qiu, Wei Chu, 09-25. Yinghui Xu, and 1 others. 2025. Aurora: Automated 22511 training framework of universal process reward mod- Lunjun Zhang, Arian Hosseini, Hritik Bansal, Mehran els via ensemble prompting and reverse verification. Kazemi, Aviral Kumar, and Rishabh Agarwal. 2025a. arXiv preprint arXiv:2502.11520. Generative verifiers: Reward modeling as next-token prediction. In The Thirteenth International Confer- ByteDance Seed Team. 2025. Seed-oss open-source ence on Learning Representations. models. https://github.com/ByteDance-Seed/ seed-oss. Xiaotian Zhang, Chunyang Li, Yi Zong, Zhengyu Ying, Liang He, and Xipeng Qiu. 2023. Evaluating the Kimi Team, Yifan Bai, Yiping Bao, Guanduo Chen, performance of large language models on gaokao Jiahao Chen, Ningxin Chen, Ruijue Chen, Yanru benchmark. arXiv preprint arXiv:2305.12474. Chen, Yuankun Chen, Yutian Chen, and 1 others. 2025. Kimi k2: Open agentic intelligence. arXiv Zhenru Zhang, Chujie Zheng, Yangzhen Wu, Beichen preprint arXiv:2507.20534. Zhang, Runji Lin, Bowen Yu, Dayiheng Liu, Jin- Qwen Team. 2024. Qwen2.5: A party of foundation gren Zhou, and Junyang Lin. 2025b. The lessons of models. developing process reward models in mathematical reasoning. arXiv preprint arXiv:2501.07301. Peiyi Wang, Lei Li, Zhihong Shao, RX Xu, Damai Dai, Yifei Li, Deli Chen, Yu Wu, and Zhifang Sui. Chujie Zheng, Zhenru Zhang, Beichen Zhang, Runji 2023. Math-shepherd: Verify and reinforce llms step- Lin, Keming Lu, Bowen Yu, Dayiheng Liu, Jin- by-step without human annotations. arXiv preprint gren Zhou, and Junyang Lin. 2024a. Processbench: arXiv:2312.08935. Identifying process errors in mathematical reasoning. arXiv preprint arXiv:2412.06559. Shijie Xia, Xuefeng Li, Yixin Liu, Tongshuang Wu, and Pengfei Liu. 2025. Evaluating mathematical reason- Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan ing beyond accuracy. In Proceedings of the AAAI Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Conference on Artificial Intelligence, volume 39, Zhuohan Li, Dacheng Li, Eric Xing, and 1 others. pages 27723–27730. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in neural information pro- Austin Xu, Xuan-Phi Nguyen, Yilun Zhou, Chien- cessing systems, 36:46595–46623. Sheng Wu, Caiming Xiong, and Shafiq Joty. 2025. Foundational automatic evaluators: Scaling multi- Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, task generative evaluator training for reasoning- Chuyue Livia Sun, Jeff Huang, Cody Hao Yu, Shiyi centric domains. arXiv preprint arXiv:2510.17793. Cao, Christos Kozyrakis, Ion Stoica, Joseph E Gonza- lez, and 1 others. 2024b. Sglang: Efficient execution An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, of structured language model programs. Advances Binyuan Hui, Bo Zheng, Bowen Yu, Chang in neural information processing systems, 37:62557– Gao, Chengen Huang, Chenxu Lv, and 1 others. 62583. 2025. Qwen3 technical report. arXiv preprint arXiv:2505.09388. Yefan Zhou, Austin Xu, Yilun Zhou, Janvijay Singh, Jiang Gui, and Shafiq Joty. 2025a. Variation in verifi- Fei Yu, Yingru Li, and Benyou Wang. 2025. Scaling cation: Understanding verification dynamics in large flaws of verifier-guided search in mathematical rea- language models. arXiv preprint arXiv:2509.17995. soning. arXiv preprint arXiv:2502.00271. Aohan Zeng, Xin Lv, Qinkai Zheng, Zhenyu Hou, Bin Yilun Zhou, Austin Xu, Peifeng Wang, Caiming Xiong, Chen, Chengxing Xie, Cunxiang Wang, Da Yin, Hao and Shafiq Joty. 2025b. Evaluating judges as Zeng, Jiajie Zhang, and 1 others. 2025. Glm-4.5: evaluators: The jetts benchmark of llm-as-judges Agentic, reasoning, and coding (arc) foundation mod- as test-time scaling evaluators. arXiv preprint els. arXiv preprint arXiv:2508.06471. arXiv:2504.15253.
Zhongshen Zeng, Pengguang Chen, Shu Liu, Haiyun Jiaru Zou, Ling Yang, Jingwen Gu, Jiahao Qiu, Jiang, and Jiaya Jia. 2023. Mr-gsm8k: A meta- Ke Shen, Jingrui He, and Mengdi Wang. 2025. reasoning benchmark for large language model eval- Reasonflux-prm: Trajectory-aware prms for long uation. arXiv preprint arXiv:2312.17080. chain-of-thought reasoning in llms. arXiv preprint arXiv:2506.18896. Zhongshen Zeng, Yinhong Liu, Yingjia Wan, Jingyao Li, Pengguang Chen, Jianbo Dai, Yuxuan Yao, Rongwu A Detailed Dataset Sources Xu, Zehan Qi, Wanru Zhao, and 1 others. 2024. Mr- ben: A comprehensive meta-reasoning benchmark In table App. A we provide the distribution of the for large language models. arXiv e-prints, pages 80 problems we sourced from different Olympiads arXiv–2406. along with the date the Olympiads were conducted. Kaiwen Zha, Zhengqi Gao, Maohao Shen, Zhang- For the IMO-shortlist, we report the earliest date Wei Hong, Duane S Boning, and Dina Katabi. 2025. Rl tango: Reinforcing generator and veri- that the shortlist questions were made publicly fier together for language reasoning. arXiv preprint available, typically the calendar year after the arXiv:2505.15034. Olympiad was conducted. 22512 Contest Date of Olympiad # Questions IMO - Shortlist 2023 21 July 2024 10 IMO - Shortlist 2024 23 July 2025 29 Putnam 7 Dec 2024 12 EGMO (European Girls’ Mathematical Olympiad) 17 April 2025 6 IMO (International Mathematical Olympiad) 20 July 2025 6 BMO (British Mathematical Olympiad) 22 Jan 2025 4 CMO (Canadian Mathematical Olympiad) 6 March 2025 4 USA-JMO (Junior Mathematical Olympiad) 20 March 2025 4 INMO (Indian National Mathematical Olympiad) 19 Jan 2025 3 USAMO (United States of America Mathematical Olympiad) 20 March 2025 2 Total 80
Table 4: Distribution of questions from various Olympiads with Year-wise Splits
B Case study: Where do models and question A1, Claude Sonnet 4 as generator con- humans disagree? structs a proof by cases by invoking Weyl’s equidis- tribution theorem, but considers only a single case: We inspect outputs from a relatively strong open- “if α is not an even integer, then α = m + β with m source verifier, ByteDance Seed-OSS-36B (Team, odd and 2/3 ≤ β < 1...”. Seed-OSS-36B green- 2025) on multiple IMO-level problems and found lights this step as correct, whereas human annota- a recurring theme: The verifier incorrectly accepts tors find it incomplete: “The case analysis ignores partial or under-justified claims as correct. We the branch where m is even and 0 < β < 1/3,...”. provide two concrete examples below. These mis- Further, the theorem invocation itself is deemed matches reflect larger systematic behavior in veri- under-specified: “justification [for invoking Weyl’s fiers, revealed in § 4: Current verifiers are too gen- equidistribution theorem] should explicitly specify erous, with TPR rate tending towards 1 and TNR the estimate and the choice of n”. tending toward 0, indicating that a vast majority steps are considered correct. On IMO 2023 Shortlist, question A6, Gemini C Additional Experimental Results 2.5 Pro makes a generalized claim, but only proves the claim for a single input. Human annotators catch this mistake, noting “The equality holds only We report TPR and TNR for all evaluated models at one point ... not a polynomial identity, so coeffi- in Table 5, alongside our aggregate metrics pre- cients need not match.” Seed-OSS-36B considers sented in § 4. We also visualize TPR and TNR this step correct without mentioning the unfounded trends for the Response-Level and Step-Level generalization. Similarly, on IMO 2024 Shortlist, tasks, similar to Fig. 4. As shown in Fig. 8, TNR is the primary driver in poor Balanced F1 perfor- 100 mance: The weaker the verifier, the more it strug- 80 gles in identifying mistakes, opting to mark nearly Percentage every step as correct. 60
40 Balanced F1
TPR
20 TNR
min PT-5G o3 3-2 -min
3
Cla 5B-A Min i
i
i2
gp .5 Pr T-5
ud eS B
De nnet
ep o
Se 22 4
t-o o
ss- 1 - o4 Qw wen Q R1
en 3-8
3-3 B ek -
gp 20B
t-o ss- GP Qw Q -A3B
en w
3-N en3-
ex 14B
t-8
Se 0B-A
ed 3B
0B
Ge 20B -O
Qw S-36
en BS
2.5
Qw 14B -
Qw en
en 2.5 Kim
Qw -72B i K2 en 3-3 2B
Qw
Lla
en -Ins
2.5 -7B ct
-In
GL ruct
M- tru
st
4.5
ma -3.3 GP
-70 T-4.1
B-In -Air
D Prompts for Generation and
str uc t
100
80
Percentage
60
40
Balanced F1
Evaluation
20 TPR
TNR
0
T-5 T-512 0B Ge min o3
en i 2.5
3-2 Pr
Qw 5B-A3
en 22 o
GP o4 -M ini 3-3 B
De B-A3
ep 0
Se
gp k-R1 e B
ss- -m t-o ss
In this section we provide prompts used for gen-
ini Qw Seed -20B
en -OS
3-N S-
ex 36B
t
t-o GP Cla -80B
ud -A3
eS Bon
Qw et 4
en n
gp Qw 4Ben3-1
Kim
Qw i K2
en 3-8 B
G
Qw PT-4.1 3-3 2B
Qw Qw en
G -14
en LM- B
2.5 4.5
Qw -72B -Air
en -In
2.5
Lla 2.5- stru
ma 7B ct
-3.3 -Ins
-70 truc
B-In tstr uc t
erating responses to Olympiad-level math ques-
Figure 8: Response-Level and ErrorID tasks follow tions. We also provide the prompts used for the similar trends in TPR and TNR, with weaker verifiers Step-Level and ErrorID tasks. unable to identify mistakes.
22513
Table 5: Complete metrics for our three evaluation tasks, reporting Balanced Accuracy, Balanced F1, TPR, and TNR.
Step-Level Response-Level ErrorID
TPR TNR Bal. Accuracy Bal. F1 TPR TNR Bal. Accuracy Bal. F1 TPR TNR Bal. Accuracy Bal. F1
Generative Critics, proprietary models
GPT-5 94.35 78.72 86.53 85.83 85.71 93.67 89.69 89.52 78.57 62.66 70.61 69.72 Gemini 2.5 Pro 88.15 78.59 83.37 83.09 80.95 90.51 85.73 85.46 52.38 52.53 52.46 52.46 Claude Sonnet 4 97.50 43.72 70.61 60.37 97.62 58.86 78.24 73.44 80.95 25.95 53.45 39.30 GPT-5-Mini 94.81 67.31 81.06 78.73 80.95 82.91 81.93 81.92 85.71 46.20 65.96 60.04 o3 95.09 62.31 78.70 75.29 90.48 75.95 83.21 82.58 73.81 46.84 60.32 57.31 o4-Mini 97.50 52.31 74.90 68.09 97.62 70.25 83.94 81.71 92.86 41.77 67.31 57.62 GPT-4.1 98.24 14.10 56.17 24.66 97.62 20.25 58.94 33.55 92.86 12.03 52.44 21.29 Generative Critics, large (≥ 70B) models Kimi K2 96.02 27.56 61.79 42.83 95.24 35.44 65.34 51.66 78.57 19.62 49.10 31.40 DeepSeek-R1 90.28 47.56 68.92 62.30 83.33 64.56 73.95 72.75 76.19 32.28 54.23 45.35 Qwen3-235B-A22B 97.41 47.69 72.55 64.03 90.48 68.35 79.42 77.87 85.71 36.08 60.90 50.78 Qwen3-Next-80B-A3B 97.87 37.95 67.91 54.69 97.62 52.53 75.08 68.31 88.10 28.48 58.29 43.05 Qwen2.5-72B-Instruct 96.76 15.26 56.01 26.36 90.48 31.65 61.06 46.89 42.86 10.13 26.49 16.38 GLM-4.5-Air 97.50 17.31 57.40 29.40 97.62 25.95 61.78 41.00 73.81 10.13 41.97 17.81 gpt-oss-120B 94.54 61.67 78.10 74.64 88.10 79.75 83.92 83.71 78.57 49.37 63.97 60.64 Llama-3.3-70B-Instruct 98.43 10.13 54.28 18.37 97.62 16.46 57.04 28.16 97.62 1.27 49.44 2.50 Generative Critics, small/medium (< 70B) models Qwen3-32B 91.94 36.03 63.99 51.77 85.71 50.00 67.86 63.16 88.10 15.82 51.96 26.83 Qwen3-30B-A3B 95.65 45.77 70.71 61.91 88.10 59.49 73.79 71.02 80.95 36.71 58.83 50.51 ByteDance Seed-OSS-36B 97.04 36.54 66.79 53.09 97.62 47.47 72.54 63.88 88.10 30.38 59.24 45.18 gpt-oss-20B 93.06 57.31 75.18 70.93 90.48 77.22 83.85 83.32 52.38 39.87 46.13 45.28 Qwen3-14B 94.17 36.79 65.48 52.91 92.86 56.33 74.59 70.12 83.33 24.05 53.69 37.33 Qwen3-8B 92.96 37.56 65.26 53.51 97.62 57.59 77.61 72.45 69.05 22.78 45.92 34.26 Qwen2.5-14B-Instruct 88.33 32.56 60.45 47.59 66.67 60.13 63.40 63.23 76.19 10.76 43.47 18.86 Qwen2.5-7B-Instruct 84.44 13.21 48.82 22.84 80.95 30.38 55.67 44.18 50.00 9.49 29.75 15.96 Process Reward Models, open-source models Qwen2.5-Math-PRM-72B 89.50 22.14 55.82 35.50 55.56 78.05 66.80 64.91 55.56 28.05 41.80 37.28 Qwen2.5-Math-PRM-7B 87.13 27.99 57.56 42.37 44.44 81.71 63.08 57.57 44.44 25.61 35.03 32.50 Skywork-PRM-7B 51.55 25.50 38.52 34.12 17.65 95.89 56.77 29.81 17.65 5.48 11.56 8.36 Skywork-PRM-1.5B 74.53 7.08 40.81 12.94 11.76 93.15 52.46 20.89 11.76 5.48 8.62 7.48 ReasonFlux-PRM-7B 93.47 12.72 53.09 22.40 66.67 45.12 55.89 53.82 66.67 18.29 42.48 28.71 UniversalPRM-7B 80.00 48.35 64.17 60.27 27.78 81.71 54.74 41.46 27.78 24.39 26.08 25.97
Prompt used to generate responses to Prompt used for the Step-Level task
Olympiad questions
The following is a math problem and a solution
You are a careful, rigorous math proof (split into steps, enclosed with tags and
assistant. Provide correct, detailed, and indexed from 0):
complete proofs. [Math Problem]
Solve the following math problem formally. {problem}
Return a detailed and formal solution that
[Solution]
can be verified by a grader.
{steps}
Use start the proof with
22514
E Annotation details
Provide reasoning for your correctness
determinations. Your final verdict should
be a comma-separated list of yes and no’s, Each sample was annotated over four rounds: An
where each yes or no corresponds to a step’s
correctness, with yes meaning correct and no initial annotation round and three rounds of reviews
meaning incorrect. to resolve disagreements. A total of 52 annotators
Please use the following format to return were employed for grading, with 35 having at least
your answer: a graduate degree in mathematics or related fields.
Reasoning:
- If step 3 is the first incorrect step: 3 (2) the consequence of the result (theorem) is correctly described and applied to the
- If all steps are correct: -1 specific problem
Do not use any other formatting, including Important: Do not apply “Error carried
markdown, bold text, code blocks, or any forward” grading.
other formatting. If your formatting is
incorrect, your evaluation will be affected. If a current step is derived from a previous
step that is incorrect, consider the current
step incorrect, even if the logic/computation
of the step is correct.
22515Example: Mark the step that contains the conclusion of Step 1: 1 + 1 = 3 [Incorrect] Case 1 incorrect, as well as any subsequent steps that depend on Case 1. Step 2: We now must add 5 to Step 1’s result, which gives us 8 [Incorrect, even though the Computationally invalid: Makes an operation computation in the step is correct; It is / value computation mistake. This should be based on an incorrect Step 1] relatively easy to spot, but please verify all complex expressions, such as integrals, Extra note: trigonometric functions, etc. “Hand-waviness”: If a model produces Note: This is not an exhaustive list a “hand-wavy” argument, wherein they of errors. Verify all computations, and say that a new result follows by document any error that occurs, no matter similar logic/computation as a previously how minor. established result, then annotators must verify that the hand-wavy argument in-fact holds. This means verifying (1) The previously established result’s assumptions F Evaluated baselines are met by the new result scenario (2) The previously established computation/logic is Here we provide a comprehensive list of models applicable to the new that were evaluated on our benchmark. Example: Step N: A valid proof of Case 1, yielding • OpenAI: GPT-5, GPT-5-Mini (OpenAI, 2025), Result 1 o3, o4-Mini (OpenAI, 2025b), GPT-4.1 (Ope- Step N+1: Case 2 follows by a similar argument nAI, 2025), gpt-oss-120b, gpt-oss-20b (Agar- to Case 1, yielding Result 2. wal et al., 2025) [This is “hand-wavy”, as the exact computation is omitted by appealing to previously computed Steps] • Google: Gemini 2.5 Pro (Google, 2025b) Incorrect: A step is considered incorrect if it is: • Anthropic: Claude Sonnet 4 (Anthropic, Based in any way on an incorrect past step. 2025) Logically invalid: The model’s output contains a reasoning error or mistake. • Moonshot (Kimi): Kimi-K2-Instruct- Examples: Unfounded logical leap 0905 (Team et al., 2025) Incorrectly invoking a mathematical result or past result when assumptions/conditions are not satisfied • DeepSeek: DeepSeek-R1 (Guo et al., 2025) Incorrect application of a mathematical result when conditions are met, i.e., • Alibaba Qwen: Qwen3-235-A22B, Qwen3- mis-applying a theorem. Next-80B-A3B, Qwen3-32B, Qwen3-30B- Failing to consider/cover a scenario or A3B, Qwen3-14B, Qwen3-8B (Yang et al., case within a proof, i.e., the proof concludes without covering all scenarios and 2025), Qwen2.5-72B-Instruct, Qwen2.5-14B- is incomplete. Instruct, Qwen2.5-7B-Instruct (Team, 2024), If the top-level proof misses a case/scenario: Qwen2.5-Math-PRM-72B, Qwen2.5-Math- As this case involves text not in the model PRM-7B (Zhang et al., 2025b) output, there is no concrete step to mark as incorrect. As a result, mark the conclusion of the proof (i.e., last step) as incorrect • Zhipu GLM: GLM-4.5-Air (Zeng et al., 2025) and provide corresponding justification. If an intermediate result is stated, but • Meta: Llama-3.3-70B-Instruct (Grattafiori the derivation of the intermediate result et al., 2024) misses a case/scenario: Mark the step that states the intermediate result as incorrect (as well as any subsequent steps that depend • ByteDance: ByteDance Seed-OSS- on the intermediate result). As a concrete 36B (Team, 2025) toy example Say a model is doing Proof by Cases for all real numbers. • Skywork: Skywork-PRM-7B, Skywork-PRM- It splits its analysis into 2 cases, Case 1 (positives) and Case 2 (negatives). For 1.5B (He et al., 2024b) Case 1, it proves the claim for all positive integers, but does not consider non-integer • ReasonFlux-PRM-7B (Zou et al., 2025) reals.
• UniversalPRM-7B (Tan et al., 2025) 22516For Kimi K2, DeepSeek-R1, and GLM-4.5-Air, Failure Mode Description % we used together.ai for inference. All other Error Propagation Conflating local vs. global 50.0 open-weight baselines were run locally, hosted via Confusion correctness (e.g., accepting a locally flawed step because vLLM (Kwon et al., 2023) or SGLang (Zheng et al., the proof direction seems 2024b). right, or rejecting a locally correct step due to a prior F.1 PRM Threshold Tuning flagged step) Rigor Misassess- Accepting heuristic argu- 18.0 To decide the cutoff threshold for evaluated PRMs, ment ments as proof, or rejecting we select 100 responses at random from our bench- standard techniques as insuf- ficient mark and tune PRM performance against this sub- Surface-Level Eval- Judging based on superficial 13.7 set, following (Zheng et al., 2024a). The same 100 uation cues rather than mathemati- responses are kept fixed across all baselines, and cal substance Mathematical Mis- The verifier’s own reasoning 13.4 we sweep the threshold from 0.1 to 0.9 in incre- understanding contains errors (e.g., wrong ments of 0.05. To select the threshold, we compute counterexamples or misinter- the harmonic mean of the three task-specific Bal- pretations) Case Boundary Missing incomplete case 4.9 anced F1 Scores, prioritizing selecting a threshold Blindness analysis, or flagging irrele- that yields strong yet balanced performance. We vant edge cases find that PRM performance can vary considerably based on chosen threshold. Table 7: Failure modes observed across 284 step-level disagreements between GPT-5 and human annotators. Error Type Description % Propagated Error Locally valid logic depend- 39.3 Computational Budget and Infrastructure ing on a prior incorrect step Details Unjustified Leap Claims without sufficient jus- 31.2 tification, or conclusions that Computation Time: 12 hours do not follow from premises Incorrect Math Wrong computations, mis- 20.0 Human Annotation time: 500 hours applied theorems, or invalid transformations GPU Hardware: 8 x NVIDIA H200 (143,771 Incomplete Cases Missing cases in case analy- 5.5 MiB RAM each) sis, or overgeneralization Invalid Setup Wrong WLOG, invalid struc- 4.0 tural assumptions, or circular Table 8: Infrastructure Details while generating reasoning Hard2Verify Dataset.
Table 6: Taxonomy of incorrect steps across 455 erro- neous steps in GPT-5-generated responses. ments, suggesting that improving verifiers’ ability to separate local step validity from global proof correctness is the single most impactful direction G Error Taxonomy of Incorrect Steps for reducing verification failures.
To better characterize the kinds of mistakes frontier I Computational Cost models make, we categorize the 455 incorrect steps from GPT-5-generated responses into five classes. This section reports the computational and human Propagated Error dominates at 39.3%, consistent effort required to construct the Hard2Verify dataset, with the error-cascade patterns in Fig. 3, indicat- including model inference for solution generation, ing that a substantial share of step-level errors are data processing, and expert annotation. downstream consequences of a single earlier mis- Table 8 summarizes the overall budget and in- take rather than independent failures. frastructure used in this work. The compute time reported corresponds to the end-to-end pipeline H GPT-5 Verifier Failure Modes for dataset generation and preparation (including solution sampling and formatting), executed on To understand where even the strongest verifier dedicated GPU resources. Human annotation time falls short, we manually categorize the 284 step- reflects the cumulative effort spent by expert an- level disagreements between GPT-5 and human notators and reviewers on step-level correctness annotators into five failure modes. Error Propa- labeling and quality control. gation Confusion accounts for half of all disagree- 22517