En
title: 英文全文 · Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification paper_id: 'ACL Anthology `2026.findings-acl.1243' year: '2026'
Inference-Time Scaling of Verification: Self-Evolving Deep Research
Agents via Test-Time Rubric-Guided Verification
Yuxuan Wan† , Tianqing Fang‡ * , Zaitang Li‡ , Yintong Huo†† ,
Wenxuan Wang‡‡ , Haitao Mi‡ , Dong Yu‡ , Michael R. Lyu†
†
The Chinese University of Hong Kong, ‡ Tencent AI Lab
††
Singapore Management University ‡‡ The Renmin University of China
§ https://github.com/Tencent/CognitiveKernel-Pro
§ https://github.com/yxwan123/DeepVerifier
Abstract +8%
Performance Gain (GAIA)
+6%
Recent advances in Deep Research Agents +4%
(DRAs) are transforming automated knowledge +2%
discovery and problem-solving. While the ma- 0
jority of existing efforts focus on enhancing
-2%
policy capabilities via post-training, we pro- 0 1 2 3 4 5 6 7 8 9 10
pose an alternative paradigm: self-evolving Number of Feedback Rounds
the agent’s ability by iteratively verifying the Claude-3.7-Sonnet DeepVeri er-8B
policy model’s outputs, guided by meticu- fi
GPT-4.1 ZeroShot Performance
lously crafted rubrics. This approach gives 40% 40%
38%
36%
rise to the inference-time scaling of verifica- 34%
29%
32%
30% 28%
tion, wherein an agent self-improves by eval- 28% 27%
Accuracy
uating its generated answers to produce iter- 20%
ative feedback and refinements. We derive 11%
the rubrics based on an automatically con- 10% 8% 8%
structed DRA Failure Taxonomy, which sys-
0%
tematically classifies agent failures into five Level 1 Level 2 Level 3 Average
major categories and thirteen sub-categories.
DeepVeri er-8B CK-Pro-8B Qwen3-8B
We present DeepVerifier, a rubrics-based out- fi
come reward verifier that leverages the asym- Figure 1: Upper: Inference-time scaling of verification
metry of verification and outperforms vanilla on the full GAIA development set (n = 165). Lower:
agent-as-judge and LLM judge baselines by Performance comparison between DeepVerifier-8B fine-
12%–48% in meta-evaluation F1 score. To en- tuned on our dataset and other open-sourced models
able practical self-evolution, DeepVerifier in- after 10 rounds of verification & feedback on the full
tegrates as a plug-and-play module during test- GAIA development set.
time inference. The verifier produces detailed
rubric-based feedback, which is fed back to
the agent for iterative bootstrapping—refining 1 Introduction
responses without additional training. This
test-time scaling delivers 8%–11% accuracy Recent advances in Deep Research Agents (DRAs),
gains on challenging subsets of GAIA and powered by large language models (LLMs) and
XBench-DeepResearch when powered by ca- vision-language models (VLMs), are transform-
pable closed-source LLMs. Finally, to sup- ing automated knowledge discovery and com-
port open-source advancement, we release plex problem-solving. These systems demonstrate
DeepVerifier-4K, a curated supervised fine-
strong performance on tasks requiring coding, web
tuning dataset of 4,646 high-quality agent steps
focused on DRA verification. These examples navigation, file processing, and multi-step reason-
emphasize reflection and self-critique, enabling ing.
open models to develop robust verification ca- However, DRAs remain prone to unreliable out-
pabilities. puts stemming from incorrect actions, API failures,
hallucinations, or other errors (Song et al., 2025;
Li and Waldo, 2024), which significantly constrain
- Correspondence: yxwan@link.cuhk.edu.hk, tianq- their practical deployment (Zhang et al., 2025a). fang@tencent.com For instance, when tasked with identifying a re- 24822 Findings of the Association for Computational Linguistics: ACL 2026, pages 24822–24835 July 2-7, 2026 ©2026 Association for Computational Linguistics searcher’s earliest publication, an agent might rely soning, multimodality, web browsing, and tool use. on incomplete secondary sources and deliver an Results show DeepVerifier outperforming vanilla inaccurate result. In long-horizon tasks involving agent-as-judge and LLM judge baselines by 12– dozens of pages and hundreds of actions, online 48% in meta-evaluation F1 score. When integrated human supervision becomes infeasible. for test-time scaling with capable closed-source These challenges underscore the need for scal- LLMs (e.g., Claude-3.5-Sonnet), it yields 8–11% able, automated methods to enhance DRA reliabil- accuracy improvements across challenging GAIA ity and performance at test time (Zhu et al., 2025c; subsets and 3–6% improvements on the XBench- Hu et al., 2025a). Prior work on inference-time im- DeepSearch dataset. provement has largely emphasized scaling output Beyond test-time inference, we extend DeepVer- tokens or selection across parallel rollouts. For ex- ifier to develop DeepVerifier-4K, a high-quality ample, Zhu et al. (2025c) introduced parallel sam- supervised fine-tuning (SFT) dataset comprising pling for optimal trajectory search, while Gonzalez- 4,646 prompt-response pairs tailored for DRA veri- Pumariega et al. (2025) employed narrative-driven fication. Curated by filtering and parsing 400 ini- aggregation across iterations. Despite existence of tial agent verification trajectories, DeepVerifier-4K Reflexion (Shinn et al., 2023)-based methods use enables robust reflection and self-critique. Using textual feedback (Zhou et al., 2025a; Yuksekgonul this dataset, we fine-tune DeepVerifier-8B, a model et al., 2024) to bootstrap the agent response, the that surpasses other open-sourced models after re- generation of feedback itself is a hard task that flection on key benchmarks. Our framework thus requires sophisticated reasoning capability (Team offers a scalable solution for both DRA verification et al., 2025; Hu et al., 2025a). and high-quality dataset creation. Moreover, as A more robust test-time self-evolution pipeline, reflection-enhanced reinforcement learning gains wherein an agent iteratively improves its outputs momentum (Hübotter et al., 2026; Liu et al., 2025), through verification and feedback without addi- our taxonomy and dataset can serve as a foundation tional training, involves (1) verifying generated for reliable self-verification and reward signals in outputs, (2) producing targeted feedback upon de- RL-based agent training. In summary, our contri- tecting errors, and (3) iterating with this feedback. butions are as follows: In this paper, we advance this pipeline in two key areas. • We formalize the agent reflection pipeline for For (1) verification, we exploit the asymmetry Deep Research Agents (DRAs) and leverage the of verification to decompose complex problems asymmetry of verification to achieve superior into simpler sub-tasks, where checking correctness meta-evaluation performance. is often easier than generation (Wei, 2025). For (2) • We introduce a comprehensive DRA failure tax- feedback generation, we incorporate rubrics-based onomy, automatically constructed to categorize rewards (Gunjal et al., 2025; Huang et al., 2025) to failures systematically, and derive structured provide structured, discriminative signals, derived rubrics for outcome-based rewards. from an automatically constructed DRA failure taxonomy. We construct the the taxonomy by an- • Through extensive experiments, we demonstrate alyzing the failure trajectories on the WebAggre- the inference-time scaling of verification that gator dataset (Wang et al., 2025), categorizing fail- holds for both capable closed-source LLM APIs ures into five major classes and thirteen sub-classes. and supervised fine-tuned models; integrating Based on (1) and (2), we present DeepVerifier, an enhanced verification capabilities significantly agentic pipeline for automatically verifying the suc- boosts overall agent performance. cess of DRA output and provide feedbacks based on the rubrics. DeepVerifier decomposes intricate 2 Related Work verification challenges into verifiable information- retrieval sub-tasks (Figure 2), overcoming limita- 2.1 Deep Research Agents tions of prior holistic judging approaches. This Research on DRA has rapidly advanced, aiming decomposition principle extends naturally to report to build autonomous systems capable of multi- generation (Fan et al., 2025). We evaluate Deep- step tasks such as web navigation, data analy- Verifier on the GAIA benchmark (Mialon et al., sis, code generation, and report synthesis. Pro- 2023), which assesses core abilities including rea- prietary frameworks like OpenAI’s Deep Re- 24823 Figure 2: Overview of DeepVerifier, which decomposes complex verification problems into smaller, simpler sub-questions leveraging the asymmetry of verification, and provides corrective feedback for the DRA to retry when the answer is considered incorrect.
search (OpenAI, 2025), Google’s Gemini Deep et al., 2025). However, these works have focused Research (Google DeepMind, 2025), Perplexity’s on web navigation tasks, general reasoning tasks, Deep Research (Perplexity AI, 2025), and Moon- or software development tasks, while none have shot AI’s Kimi-Researcher (Moonshot AI, 2025a,b) studied the responses of DRAs. demonstrate strong performance on benchmarks Recent research also investigates self-evolving such as GAIA and Humanity’s Last Exam, set- LLMs (Zhou et al., 2025b; Zhang et al., 2025b; ting high standards for autonomy and multimodal Zuo et al., 2025; Zhang et al., 2025a; Feng et al., reasoning (Mialon et al., 2023; Phan et al., 2025; 2025). For example, recent methods explore code- Zhang et al., 2026). Meanwhile, open-source as-task self-play (Zhou et al., 2025b), self-aware frameworks democratize agent development. No- RL (Zhang et al., 2025b), and test-time RL (Zuo table systems include SmolAgents (Roucher et al., et al., 2025), but none address DRAs. (Zhang et al., 2025), the WebAgent family (Wu et al., 2025a; 2025a) systematically analyze failure modes of Li et al., 2025; Tao et al., 2025), OWL (Hu et al., DRAs, but do not provide an automated framework 2025b), OAgents (Zhu et al., 2025a,c), and CK- for detecting failures or improving agents based Pro (Fang et al., 2025b), among others (Wu et al., on these findings. In contrast, we (1) construct an 2025b; Bahdanau et al., 2024; Tang et al., 2025; agent failure taxonomy, (2) introduce a verification- Zhang et al., 2024; Fang et al., 2025a). While asymmetry–based framework to automatically de- most efforts are being put into Agent Foundation tect failures, and (3) extend it to self-evolving veri- Model Training using Supervised Finetuning (Ope- fication, demonstrating a clear verification scaling nAI, 2025; Hu et al., 2025b; Wu et al., 2025a; Hu effect. et al., 2025c) and Reinforcement Learning (Li et al., 2025, 2026; Yu et al., 2025; Wang et al., 2026; Fang 3 DRA Failure Taxonomy et al., 2026a,b; Zhu et al., 2026), DRA verification To exploit the asymmetry of verification and de- and its scaling effect remain underexplored. compose complex problems into simpler sub-tasks, 2.2 Test-Time Scaling of Agents we first investigate the common failures of DRA and construct a DRA Failure Taxonomy. To avoid Many works apply Test-Time-Scaling (Choi et al., data leakage or contamination and ensure general- 2023; Snell et al., 2024) to enhance the quality of ization, we select the WebAggregatorQA dataset agent responses. (Zhu et al., 2025b) proposes Best- to construct the taxonomy, and evaluate the frame- of-N selection, majority vote, etc. However, such work on three distinct dataset: GAIA, BrowseC- test-time-scaling methods remain prone to the same omp, and XBench-DeepSearch to demonstrate the set of failures in different roll-outs, meaning that effectiveness and generalization of the method. errors arising in one run also tend to recur in other runs, rendering the overall result unreliable. Other Trajectory Collection To construct the taxon- works explored using LLMs or agents as judges to omy, we first collect problem-solving trajectories evaluate agent responses (He et al., 2024; Pan et al., from a representative deep research agent. Table 1 2024; Lù et al., 2025; Zhuge et al., 2024; Yang summarizes the resulting corpus, which is substan- 24824 Table 1: Statistics of collected trajectories. Steps refers iteration, we construct a new version of the tax- to the actions (planning, searching, clicking, etc.) per- onomy by comparing and merging similar labels, formed by agents and sub-agents. Number of tokens is removing inadequate categories, refining unclear calculated by the GPT-4o tokenizer. definitions based on the results of previous itera- Trajectory Stat Min Max Avg Total tions, and discussing the results of the last iteration. As a result, we obtain a classification scheme illus- Steps 2.0 156.0 33.3 2,997 Tokens 18.7K 60.0M 8.2M 738M trated in Figure 3. The more frequent the subclass, Correct/Incorrect - - - 0.96 the wider the branch. Unique Tasks - - - 90 Analysis Figure 3 shows that DRA failures are dominated by Finding Sources, with the largest tial (2,997 agent actions), diverse (90 distinct tasks; flows corresponding to errors such as consulting trajectories range from 2 to 156 steps), and nearly the wrong evidence and relying on generic searches, balanced (correct/incorrect ratio of 0.96). We use highlighting that upstream information acquisition Cognitive Kernel-Pro (Fang et al., 2025b), a high- is the most frequent point of collapse. Reasoning performing fully open-source multi-module DRA failures are the next most common, driven by pre- framework, with Claude-3.7-Sonnet as the back- mature conclusions, misinterpretation, and halluci- bone model due to its strong performance in this nated or overconfident claims, indicating that even setting. Trajectories are generated by running the when information is present, agents often make in- agent on WebAggregatorQA (Wang et al., 2025), correct inferential leaps. Problem Understanding, a benchmark that exercises core DRA capabilities Action Errors, and Max Step Reached account for including multi-step reasoning, multimodal inputs, the remaining failures, often cascading from early web browsing, and general tool-use proficiency. mistakes into long, unproductive trajectories. Error Points Collection For each trajectory that 4 DeepVerifier produces an incorrect final answer, we annotate the underlying failure points. We use the human We present an overview of the DeepVerifier frame- reference solution traces provided by WebAggre- work in Figure 2. We adopt a three-stage multi- gatorQA as a grounding signal, and recruit two module framework in our agent implementation. research staff annotators to independently inspect This framework consists of a decomposition agent, the agent’s execution and identify deviations from a verification agent, and a judge agent. The follow- the reference reasoning and evidence-gathering pro- ing sections describe each module in detail. cess. Each annotator records a set of error points, i.e., concrete, localized mistakes such as missing 4.1 Decomposition Module critical evidence, using an invalid source, or mis- The decomposition agent leverages previous tra- interpreting an instruction, along with the support- jectories and the DRA failure taxonomy to exploit ing trajectory step(s). We then reconcile the two the asymmetry of verification. Instead of asking annotation sets through a merge procedure: dupli- the verification agent to re-solve the entire complex cated items are consolidated, and distinct items are task (e.g., "Given a query, an unverified answer, retained in the final list. We calculate that on av- and the agent’s trajectory, verify the correctness of erage, 63.0% of the error points of one annotator the answer"), which often results in high error rates overlapped with the other’s, indicating a relatively similar to those of the original agent execution, high agreement rate between the annotators. This the decomposition agent breaks the problem into process yields 555 error points. Full annotation smaller, more manageable sub-questions. These guidelines are provided in Appendix A. sub-questions target specific vulnerabilities in the previous solution, such as “Does source X state Taxonomy Construction To gain further insight claim Y?” or “What is the exact figure for Y in the into the failures, we construct a taxonomy based on latest report X?” The workflow of the decomposi- the error points. In particular, we conduct an itera- tion agent comprises three steps. tive analysis and labeling process with two annota- tors with multiple years of AI research experience Trajectory Summarization. Agent trajectories from our institute. The initial labels are determined average 8.2M tokens, far exceeding any LLM’s by clustering a subset of 50 error points. In each context window. Moreover, concise descriptions of 24825 Figure 3: DRA failure taxonomy that categorizes 555 agent failures into five major classes and thirteen subclasses.
rollout steps can improve test-time scaling (Fang 4.2 Verification Agent and Judge Module et al., 2025b; Gonzalez-Pumariega et al., 2025). We Verification The verification module retrieves an- therefore instruct the decomposition agent to first swers to the follow-up questions sequentially. In produce a compact, step-indexed synopsis of the our experiment, we use the CK-Pro agent (Fang trajectory. For each step, it records the source vis- et al., 2025b) as the verification agent, a modu- ited and the concrete information retrieved (facts, lar multi-agent framework capable of web search, numbers, quotes). The summary is descriptive, not screenshotting, and code execution. interpretive, enabling downstream checks without reloading the full trace. Judge The judge agent evaluates the unverified answer based on the trajectory summary, potential error list, follow-up questions, and their answers. Potential Error Identification. Given the sum- It begins by providing a concise explanation, fol- mary and our failure taxonomy in the system lowed by a score between 1 and 4, where: 1 = prompt, the decomposition agent scans for behav- entirely incorrect, 2 = mostly incorrect, 3 = mostly iors that align with known failure modes . It pro- correct, 4 = entirely correct. duces paired findings of the form ⟨behavior⟩ ⇒ ⟨potential error + taxonomy label⟩ with a brief jus- 5 Enhancing Deep Research Agents with tification. These structured pairs localize where Scalable Verification and how failures likely arise. Test-Time Scaling with Reflection and Feedback. Beyond verification, our framework enhances the Follow-Up Question Formulation. Finally, the test-time scaling performance of DRAs through re- decomposition agent drafts high-leverage follow- flection. By integrating DeepVerifier into the DRA, up questions targeted at the flagged vulnerabilities. the agent can review and evaluate its previous ac- Each question is answerable via external evidence tions. Specifically, we modify the judge agent’s and designed to decisively confirm or refute a risky prompt to: 1) provide actionable instructions for claim. the agent to retry tasks and avoid repeating mis- By focusing only on essential, potentially faulty takes, and 2) suggest correct answers if they are claims, this process allows the verification agent to already available within the given information (e.g., build on well-grounded conclusions, ignore trivial previous trajectories or follow-up answers). After details, and check only for suspicious or unsup- completing each task, the agent verifies its own ported assertions. Detailed prompts of each step outputs using DeepVerifier, collects feedback, and are shown in Appendix B. uses it to guide further retries. This process re- 24826 peats until a satisfactory answer is reached or a Table 2: Ablation study on GAIA-Web. “− Verifica- predefined retry limit is exceeded. tion” corresponds to a decomposition-only (LLM judge) baseline; “− Decomposition” corresponds to a vanilla Training Reflection Ability in Agent Foundation agent-as-judge baseline (CK-Pro as judge). Metrics are Models. Many open-source models, lacking fine- precision/recall of rejection (values ×100). tuning for reflection, show limited test-time scaling Method Precision Recall Accuracy F1 capabilities (Fang et al., 2025b). To address this, DeepVerifier 75.00 71.43 75.56 73.17 we propose a deep verification training dataset that - Verification 100.00 14.29 60.00 25.00 leverages existing datasets and DeepVerifier to im- - Decomposition 86.96 47.62 72.22 61.54 prove the reflection and test-time scaling abilities of open-source LLMs. Base Trajectory Collection. We first collect DeepVerifier-4K and the CK-Pro-8B training set 400 answers and trajectories from agents solving from (Fang et al., 2025b) to train reflection abili- tasks that require significant online exploration and ties in open-source models while preserving their information gathering. These tasks are sampled foundational capabilities. The training parameters from the WebAggregatorQA dataset (Wang et al., are set as follows: 2025), which tests agents on information aggrega- Baselines and Metrics. We use the LLM judge tion across 10+ domains. Using the CK-Pro agent proposed by (Lù et al., 2025) as the LLM verifier with Claude-3.7-Sonnet as the backbone model, we baseline, and the CK-Pro Agent (Fang et al., 2025b) record the answers and corresponding trajectories. as the agent verifier baseline. Detailed prompts are Verification Trajectory and SFT Data Collection. shown in Appendix B. For verification tasks, we Next, we use DeepVerifier with Claude-3.7-Sonnet calculate the standard precision, recall, accuracy, to verify the collected base trajectories and answers, and F1 score to measure the correctness of the eval- saving the verification trajectories. We filter the uation, where true positive is defined as a verifier true positive and true negative verifications—those assigning “reject” label to a wrong answer, and that correctly accept true answers and correctly a true negative is defined as a verifier assigning reject false ones. After balancing these trajecto- “accept” label to an correct answer. In the scaling ries, we convert them into prompt-response pairs, experiment, we treat a score of less than or equal resulting in DeepVerifier-4K, a dataset of 4,646 2 as incorrect, and greater or equal to 3 as correct. high-quality pairs. We stop the feedback loop as soon as the verifier judge the answer as correct. 6 Experiment Setup Research Questions We investigate the follow- Models and Benchmarks. We mainly use ing research questions (RQs) to demonstrate the Claude-3.7-Sonnet as the backbone model of Deep- effectiveness of our method: Verifier and other methods. To evaluate the gener- alization ability of our method, we also compare 1. RQ1: Is DeepVerifier effective in verification? the performance on GPT-4.1 and Qwen3-8B. We 2. RQ2: Can DeepVerifier help improve the per- evaluate baselines and our methods primarily on formance of DRA via test-time scaling? the GAIA-web dataset, which is a subset of the GAIA dataset filtered for tasks that require web 3. RQ3: Can DeepVerifier-4K help improve the browsing following (He et al., 2024). To ensure reflection ability of open-sourced models? generalization, we also extend evaluations on the full GAIA dataset (Mialon et al., 2023), XBench- 7 Results & Analysis DeepSearch (Chen et al., 2025), and BrowseC- 7.1 RQ1: Effectiveness of DeepVerifier omp (Wei et al., 2025). XBench-DeepSearch is a Chinese benchmark for search/tool-use, and We conduct an ablation study using the trajecto- BrowseComp measures agents’ ability to retrieve ries of the CK-Pro agent with a Claude-3.7-Sonnet extremely hard-to-find and entangled information. backbone on the GAIA-Web dataset, as described in Table 1. Each method, using the same back- Training Configurations. To demonstrate the bone model, is evaluated on its ability to verify the effectiveness of our approach on open-sourced correctness of these cases. As shown in Table 2, models, we fine-tune Qwen3-8B on a mixture of DeepVerifier achieves superior performance across 24827 recall, accuracy, and F1 score. Removing the veri- grounding process. Meanwhile, the reasoning fication module or decomposition module exhibits and file-operation subset also exhibits improve- high precision (100% and 86.96%, respectively) in ment across rounds, demonstrating that the reflec- detecting erroneous cases, but their recall and accu- tive feedback mechanism generalizes beyond web- racy remain unsatisfactory. Closer analysis reveals based scenarios. GPT-4.1 shows a similar trend, that these judges are effective at catching obvious improving from 29.5% to 32.5% (best), confirming mistakes, such as execution failures, but often over- cross-model generalization (Figure 1). look subtler reasoning or factual errors, accepting many incorrect answers as correct. This limitation Performance on other DRA datasets. Results arises because removing the verification module in Table 4 show that the scaling effect remains renders the judge fail to identify secondary-source consistent despite the multi-lingual nature of dependence, overconfident claims, or hallucinated DeepSearch and the extreme difficulty of BrowseC- facts supporting incorrect responses. Meanwhile, omp: XBench-DeepSearch improves from 41.0 (0 removing the decomposition does not affect the rounds) to 47.0 (best, +6.0), and ends at 44.0 (+3.0 judge’s access to external sources, but we observe at 10 rounds); BrowseComp improves from 5.0 to that without proper decomposition, the agent tends 10.0 (best, +5.0), and ends at 9.0 (+4.0). to check every step by re-solving the entire task, leaving them vulnerable to the same reasoning er- rors as the original agent. In contrast, DeepVerifier Analysis of the Scaling Trend Performance typ- decomposes complex verification into smaller, tar- ically peaks in early feedback rounds due to our it- geted sub-questions that directly test specific vul- erative setting and the verifier’s imperfect precision nerabilities, making it more robust against faulty and recall. In each round, the verifier enables many reasoning and unsupported claims. incorrect cases to be fixed (incorrect→correct), but also occasionally rejects correct answers, causing Answer to RQ1: DeepVerifier is effective in regressions (correct→incorrect). Table 5 shows DRA verification, achieving a balanced preci- that the incorrect→correct transition is stronger sion–recall tradeoff and yields a 12% - 48% im- but decays quickly, whereas the correct→incorrect provement in F1 score and highest accuracy com- transition is weaker but persists across rounds; their pared to ablated versions. interplay produces the observed peak around the 7.2 RQ2: Improving the Performance of DRA fourth round. Via Reflective Test-Time Scaling We evaluate whether DeepVerifier can enhance the Inference Cost. While iterative verification in- performance of Deep Research Agents through re- troduces additional compute, DeepVerifier is rel- flective test-time scaling by integrating it into the atively efficient: the decomposition module nar- CK-Pro agent with Claude-3.7-Sonnet and measur- rows verification to ≤3 targeted follow-up ques- ing accuracy across feedback rounds on the GAIA tions rather than re-solving the full task, accuracy dataset. As shown in Table 3, accuracy consistently gains peak around round 3–4 enabling practical improves with additional feedback iterations, reach- early stopping, and the loop terminates as soon as ing its peak at the fourth round. This demonstrates the verifier accepts the answer. Compared to broad that iterative reflection and verification feedback ef- search-based scaling (e.g., Best-of-N with full re- fectively help the agent refine reasoning and correct execution), this yields a favorable accuracy–cost previous errors. tradeoff without additional training. Performance on the GAIA dataset. The over- all accuracy on GAIA-Full increases from ap- Answer to RQ2: DeepVerifier effectively scales proximately 52% to 59%, with peak value reach- DRA performance through structured reflection: ing 60.1%, marking the best performance gain as feedback rounds increase, the agent progres- of 8%. The GAIA-Web subset shows the great- sively enhances its accuracy, achieving over 8% est improvement, rising from 52% to above 62%, performance gains on Claude-3.7-Sonnet without with peak value reaching 63.5%, indicating that additional training or external supervision. The web-based, retrieval-heavy tasks benefit most from scaling behavior also generalizes to other models DeepVerifier ’s targeted verification and evidence- and datasets. 24828 Table 3: Accuracy(%) on different subsets of the GAIA dataset with different rounds of feedback using DeepVerifier (DV) across different backbone models.
# Feedback Rounds
GAIA Split Model Final Gain Best Gain
0 2 4 6 8 10
Claude-3.7 51.11 58.89 63.33 62.22 61.11 62.22 11.11 12.22
Web GPT-4.1 28.89 32.22 31.11 32.22 31.11 31.11 2.22 3.33
DV-8B 26.67 31.11 31.11 32.22 33.33 33.33 6.67 6.67
Claude-3.7 53.57 53.57 56.21 54.92 54.92 54.92 1.35 2.64
File/Reasoning/Others GPT-4.1 30.67 33.33 33.33 33.33 33.33 33.33 2.67 2.67
DV-8B 26.81 30.85 30.85 30.85 30.85 30.85 4.04 4.04
Claude-3.7 52.22 56.49 60.12 58.93 58.32 58.93 6.71 7.90
Full GPT-4.1 29.51 32.53 31.92 32.53 31.92 31.92 2.41 3.01
DV-8B 26.73 30.99 30.99 31.60 32.21 32.21 5.48 5.48
Table 4: Accuracy(%) across different datasets versus feedback rounds using DeepVerifier with Claude-3.7-Sonnet backbone.
Dataset 0 1 2 3 4 5 6 7 8 9 10 Final Gain Best Gain
DeepSearch 41.0 42.0 47.0 41.0 45.0 44.0 43.0 44.0 42.0 44.0 44.0 3.0 6.0
BrowseComp 5.0 8.0 10.0 10.0 9.0 9.0 9.0 9.0 9.0 9.0 9.0 4.0 5.0
Table 5: Transition rates between consecutive feedback rounds.
Feedback Round 1 2 3 4 5 6 7 8 9 10
Incorrect to Correct Ratio (%) 18.99 9.33 6.94 8.45 0.00 1.45 0.00 0.00 1.45 0.00
Correct to Incorrect Ratio(%) 12.79 4.44 4.30 1.06 3.03 0.00 0.00 1.03 0.00 0.00
7.3 RQ3: Enhancing Reflection Ability of Answer to RQ3: Incorporating DeepVerifier ’s Open-Sourced Models reflection ability through fine-tuning significantly improves the reasoning and verification perfor- mance of Deep Research Agents. The fine-tuned We further investigate whether incorporating reflec- DeepVerifier-8B model achieves a 5.5% accuracy tion ability through SFT can improve the reason- gain compared to its non-reflective version and the ing and verification performance of Deep Research Qwen3-8B model. Agents. We fine-tune Qwen3-8B on a mixture of DeepVerifier-4K and the CK-Pro training set (Fang 8 Conclusion et al., 2025b), which we name DeepVerifier-8B, and use this model as the backbone for CK-Pro In this paper, we address the challenge of silent Agent with DeepVerifier as the reflection module, and repeated failures in Deep Research Agents by measuring accuracy after 10 feedback rounds on systematically leveraging the asymmetry of ver- the GAIA dataset. As shown in Figure 1, models ification. We construct a human-annotated fail- fine-tuned with the DeepVerifier-4K dataset exhibit ure taxonomy, introduce a taxonomy-guided vul- notable performance gains when equipped with nerability localization mechanism that transforms reflection. Specifically, DeepVerifier-8B, which verification from holistic re-solving into targeted is trained with both the CK-Pro dataset and the evidence checking, and demonstrate consistent im- DeepVerifier-4K reflective data, achieves the high- provements across models and datasets. We also est accuracy of 32.2% after reflection, representing release DeepVerifier-4K to empower the open com- a 5.5% improvement over its non-reflective result. munity to build more trustworthy agents. Our In contrast, CK-Pro-8B, trained only on the CK-Pro framework offers a practical solution for scalable dataset, achieves a smaller gain of 2.6 points, while DRA verification, and we believe it can meaning- Qwen3-8B, which lacks both CK-Pro and Deep- fully aid the growing body of work on reflection- Verifier training, shows minimal improvement. enhanced reinforcement learning for agents. 24829 9 Limitations Tianqing Fang, Zhisong Zhang, Xiaoyang Wang, Rui Wang, Can Qin, Yuxuan Wan, Jun-Yu Ma, Ce Zhang, Dependence on model capability. DeepVerifier re- Jiaqi Chen, Xiyun Li, Hongming Zhang, Haitao Mi, lies on models’ ability to follow rubrics precisely, and Dong Yu. 2025b. Cognitive kernel-pro: A frame- work for deep research agents and agent foundation perform careful cross-checking, and express struc- models training. Preprint, arXiv:2508.00414. tured feedback. When the underlying model is weak or lacks sufficient tool-use ability, feedback Yangyi Fang, Jiaye Lin, Xiaoliang Fu, Cong Qin, Haolin Shi, Chaowen Hu, Lu Pan, Ke Zeng, and Xunliang quality can degrade and yield noisy test-time gains. Cai. 2026a. How to allocate, how to learn? dynamic DeepVerifier-4K can help alleviate this limitation rollout allocation and advantage modulation for pol- for open-sourced models via SFT. icy optimization. arXiv preprint arXiv:2602.19208. Test-time cost and latency. While DeepVerifier Yangyi Fang, Jiaye Lin, Xiaoliang Fu, Cong Qin, Haolin can minimize the redundant problem-solving steps, Shi, Chang Liu, and Peilin Zhao. 2026b. Proximity- iterative verification still introduces extra inference based multi-turn optimization: Practical credit as- steps (and often additional tool calls), increasing signment for llm agent training. arXiv preprint arXiv:2602.19225. runtime and token usage. Kaituo Feng, Manyuan Zhang, Hongyu Li, Kaixuan Acknowledgments Fan, Shuang Chen, Yilei Jiang, Dian Zheng, Peiwen Sun, Yiyuan Zhang, Haoze Sun, and 1 others. 2025. This research is supported by the Research Grants Onethinker: All-in-one reasoning model for image Council of the Hong Kong Special Administrative and video. arXiv preprint arXiv:2512.03043. Region, China (No. CUHK 14209124) under the Gonzalo Gonzalez-Pumariega, Vincent Tu, Chih-Lun General Research Fund. Lee, Jiachen Yang, Ang Li, and Xin Eric Wang. 2025. The unreasonable effectiveness of scaling agents for computer use. References Google DeepMind. 2025. Gemini deep research — your personal research assistant. Dzmitry Bahdanau, Nicolas Gontier, Gabriel Huang, Ehsan Kamalloo, Rafael Pardinas, Alex Piché, Anisha Gunjal, Anthony Wang, Elaine Lau, Vaskar Torsten Scholak, Oleh Shliazhko, Jordan Prince Nath, Yunzhong He, Bing Liu, and Sean Hendryx. Tremblay, Karam Ghanem, Soham Parikh, Mitul 2025. Rubrics as rewards: Reinforcement learn- Tiwari, and Quaizar Vohra. 2024. Tapeagents: a ing beyond verifiable domains. arXiv preprint holistic framework for agent development and opti- arXiv:2507.17746. mization. arXiv preprint arXiv:2412.08445. Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Kaiyuan Chen, Yixin Ren, Yang Liu, Xiaobo Hu, Hao- Yong Dai, Hongming Zhang, Zhenzhong Lan, and tong Tian, Tianbao Xie, Fangfu Liu, Haoye Zhang, Dong Yu. 2024. Webvoyager: Building an end-to- Hongzhang Liu, Yuan Gong, and 1 others. 2025. end web agent with large multimodal models. arXiv xbench: Tracking agents productivity scaling with preprint arXiv:2401.13919. profession-aligned real-world evaluations. arXiv preprint arXiv:2506.13651. Chen Hu, Haikuo Du, Heng Wang, Lin Lin, Min- grui Chen, Peng Liu, Ruihang Miao, Tianchi Yue, Sehyun Choi, Tianqing Fang, Zhaowei Wang, and Wang You, Wei Ji, Wei Yuan, Wenjin Deng, Xiaojian Yangqiu Song. 2023. KCTS: knowledge-constrained Yuan, Xiaoyun Zhang, Xiangyu Liu, Xikai Liu, Yan- tree search decoding with token-level hallucination ming Xu, Yicheng Cao, Yifei Zhang, and 48 others. detection. In Proceedings of the 2023 Conference 2025a. Step-deepresearch technical report. Preprint, on Empirical Methods in Natural Language Process- arXiv:2512.20491. ing, EMNLP 2023, Singapore, December 6-10, 2023, Mengkang Hu, Yuhang Zhou, Wendong Fan, Yuzhou pages 14035–14053. Association for Computational Nie, Bowei Xia, Tao Sun, Ziyu Ye, Zhaoxuan Jin, Linguistics. Yingru Li, Qiguang Chen, Zeyu Zhang, Yifeng Wang, Qianshuo Ye, Bernard Ghanem, Ping Luo, and Guo- Tianyu Fan, Xinyao Niu, Yuxiang Zheng, Fengji Zhang, hao Li. 2025b. Owl: Optimized workforce learning Chengen Huang, Bei Chen, Junyang Lin, and Chao for general multi-agent assistance in real-world task Huang. 2025. Understanding deepresearch via re- automation. ports. arXiv preprint arXiv:2510.07861. Minda Hu, Tianqing Fang, Jianshu Zhang, Junyu Ma, Tianqing Fang, Hongming Zhang, Zhisong Zhang, Zhisong Zhang, Jingyan Zhou, Hongming Zhang, Kaixin Ma, Wenhao Yu, Haitao Mi, and Dong Yu. Haitao Mi, Dong Yu, and Irwin King. 2025c. Web- 2025a. Webevolver: Enhancing web agent self- cot: Enhancing web agent reasoning by reconstruct- improvement with coevolving world model. arXiv ing chain-of-thought in reflection, branching, and preprint arXiv:2504.21024. rollback. arXiv preprint arXiv:2505.20013. 24830 Zenan Huang, Yihong Zhuang, Guoshan Lu, Zeyu Qin, Aymeric Roucher, Albert Villanova del Moral, Thomas Haokai Xu, Tianyu Zhao, Ru Peng, Jiaqi Hu, Zhan- Wolf, Leandro von Werra, and Erik Kaunismäki. ming Shen, Xiaomeng Hu, and 1 others. 2025. Re- 2025. Smolagents: a smol library to build great inforcement learning with rubric anchors. arXiv agentic systems. preprint arXiv:2508.12790. Noah Shinn, Federico Cassano, Ashwin Gopinath, Jonas Hübotter, Frederike Lübeck, Lejs Behric, An- Karthik Narasimhan, and Shunyu Yao. 2023. Re- ton Baumann, Marco Bagatella, Daniel Marta, Ido flexion: Language agents with verbal reinforcement Hakimi, Idan Shenfeld, Thomas Kleine Buening, Car- learning. Advances in Neural Information Process- los Guestrin, and Andreas Krause. 2026. Reinforce- ing Systems, 36:8634–8652. ment learning via self-distillation. arXiv preprint arXiv:2601.20802. Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Ku- mar. 2024. Scaling llm test-time compute optimally Eric Li and Jim Waldo. 2024. Websuite: Systemati- can be more effective than scaling model parameters. cally evaluating why web agents fail. arXiv preprint arXiv preprint arXiv:2408.03314. arXiv:2406.01623. Kevin Song, Anand Jayarajan, Yaoyao Ding, Qidong Kuan Li, Zhongwang Zhang, Huifeng Yin, Liwen Su, Zhanda Zhu, Sihang Liu, and Gennady Pekhi- Zhang, Litu Ou, Jialong Wu, Wenbiao Yin, Zheng- menko. 2025. Aegis: Taxonomy and optimizations wei Tao, Xinyu Wang, Weizhou Shen, Junkai Zhang, for overcoming agent-environment failures in llm Dingchu Zhang, Xixi Wu, Yong Jiang, Ming Yan, agents. arXiv preprint arXiv:2508.19504. Pengjun Xie, Fei Huang, and Jingren Zhou. 2025. Websailor: Navigating super-human reasoning for Jiabin Tang, Tianyu Fan, and Chao Huang. 2025. Au- web agent. toagent: A fully-automated and zero-code framework for llm agents. arXiv preprint arXiv:2502.05957. Mukai Li, Qingcheng Zeng, Tianqing Fang, Zhenwen Liang, Linfeng Song, Qi Liu, Haitao Mi, and Dong Zhengwei Tao, Jialong Wu, Wenbiao Yin, Junkai Zhang, Yu. 2026. Verified critical step optimization for LLM Baixuan Li, Haiyang Shen, Kuan Li, Liwen Zhang, agents. CoRR, abs/2602.03412. Xinyu Wang, Yong Jiang, Pengjun Xie, Fei Huang, and Jingren Zhou. 2025. Webshaper: Agentically Zihan Liu, Shun Zheng, Xumeng Wen, Yang Wang, data synthesizing via information-seeking formaliza- Jiang Bian, and Mao Yang. 2025. Deep self-evolving tion. reasoning. arXiv preprint arXiv:2510.17498. Tongyi DeepResearch Team, Baixuan Li, Bo Zhang, Xing Han Lù, Amirhossein Kazemnejad, Nicholas Dingchu Zhang, Fei Huang, Guangyu Li, Guoxin Meade, Arkil Patel, Dongchan Shin, Alejandra Zam- Chen, Huifeng Yin, Jialong Wu, Jingren Zhou, and 1 brano, Karolina Stánczak, Peter Shaw, Christopher J. others. 2025. Tongyi deepresearch technical report. Pal, and Siva Reddy. 2025. Agentrewardbench: Eval- arXiv preprint arXiv:2510.24701. uating automatic evaluations of web agent trajecto- ries. arXiv preprint arXiv:2504.08942. Rui Wang, Ce Zhang, Junyu Ma, Jianshu Zhang, Hon- gru Wang, Yi Chen, Boyang Xue, Tianqing Fang, Grégoire Mialon, Clémentine Fourrier, Craig Swift, Zhisong Zhang, Hongming Zhang, Haitao Mi, Dong Thomas Wolf, Yann LeCun, and Thomas Scialom. Yu, and Kam-Fai Wong. 2025. Explore to evolve: 2023. Gaia: a benchmark for general ai assistants. Scaling evolved aggregation logic via proactive on- ArXiv, abs/2311.12983. line exploration for deep research agents. Moonshot AI. 2025a. Kimi-k2. Tianyi Wang, Long Li, Hongcan Guo, Yibiao Chen, Moonshot AI. 2025b. Kimi-researcher: End-to-end rl Yixia Li, Yong Wang, Yun Chen, and Guanhua Chen. training for emerging agentic capabilities. 2026. Anchored policy optimization: Mitigating exploration collapse via support-constrained rectifi- OpenAI. 2025. Introducing deep research. Technical cation. CoRR, abs/2602.05717. report, OpenAI. Jason Wei. 2025. Asymmetry of verification and veri- Jiayi Pan, Yichi Zhang, Nicholas Tomlin, Yifei Zhou, fier’s law. Accessed: 2025-10-30. Sergey Levine, and Alane Suhr. 2024. Autonomous evaluation and refinement of digital agents. arXiv Jason Wei, Zhiqing Sun, Spencer Papay, Scott McK- preprint arXiv:2404.06474. inney, Jeffrey Han, Isa Fulford, Hyung Won Chung, Alex Tachard Passos, William Fedus, and Amelia Perplexity AI. 2025. Introducing perplexity deep re- Glaese. 2025. Browsecomp: A simple yet challeng- search. ing benchmark for browsing agents. arXiv preprint arXiv:2504.12516. Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Jialong Wu, Baixuan Li, Runnan Fang, Wenbiao Yin, Mohamed Shaaban, John Ling, Sean Shi, and 1 oth- Liwen Zhang, Zhengwei Tao, Dingchu Zhang, Zekun ers. 2025. Humanity’s last exam. arXiv preprint Xi, Gang Fu, Yong Jiang, Pengjun Xie, Fei Huang, arXiv:2501.14249. and Jingren Zhou. 2025a. Webdancer: Towards 24831 autonomous information seeking agency. arXiv He Zhu, Tianrui Qin, King Zhu, Heyuan Huang, Yeyi preprint arXiv:2505.22648. Guan, Jinxiang Xia, Yi Yao, Hanhao Li, Ningning Wang, Pai Liu, Tianhao Peng, Xin Gui, Xiaowan Li, Jialong Wu, Wenbiao Yin, Yong Jiang, Zhenglin Wang, Yuhui Liu, Yuchen Eleanor Jiang, Jun Wang, Chang- Zekun Xi, Runnan Fang, Linhai Zhang, Yulan He, wang Zhang, Xiangru Tang, Ge Zhang, and 5 others. Deyu Zhou, Pengjun Xie, and Fei Huang. 2025b. 2025a. Oagents: An empirical study of building Webwalker: Benchmarking llms in web traversal. effective agents. CoRR, abs/2501.07572. King Zhu, Hanhao Li, Siwei Wu, Tianshun Xing, De- Yiliu Yang, Yilei Jiang, Qunzhong Wang, Yingshui Tan, hua Ma, Xiangru Tang, Minghao Liu, Jian Yang, Jia- Xiaoyong Zhu, Sherman SM Chow, Bo Zheng, and heng Liu, Yuchen Eleanor Jiang, Changwang Zhang, Xiangyu Yue. 2025. Quadsentinel: Sequent safety for Chenghua Lin, Jun Wang, Ge Zhang, and Wangchun- machine-checkable control in multi-agent systems. shu Zhou. 2025b. Scaling test-time compute for llm arXiv preprint arXiv:2512.16279. agents. Preprint, arXiv:2506.12928. Wenhao Yu, Zhenwen Liang, Chengsong Huang, Kis- han Panaganti, Tianqing Fang, Haitao Mi, and Dong King Zhu, Hanhao Li, Siwei Wu, Tianshun Xing, Dehua Yu. 2025. Guided self-evolving llms with minimal Ma, Xiangru Tang, Minghao Liu, Jian Yang, Jiaheng human supervision. CoRR, abs/2512.02472. Liu, Yuchen Eleanor Jiang, and 1 others. 2025c. Scal- ing test-time compute for llm agents. arXiv preprint Mert Yuksekgonul, Federico Bianchi, Joseph Boen, arXiv:2506.12928. Sheng Liu, Zhi Huang, Carlos Guestrin, and James Zou. 2024. Textgrad: Automatic" differentiation" via Mingchen Zhuge, Changsheng Zhao, Dylan Ashley, text. arXiv preprint arXiv:2406.07496. Wenyi Wang, Dmitrii Khizbullin, Yunyang Xiong, Zechun Liu, Ernie Chang, Raghuraman Krishnamoor- Dingling Zhang, He Zhu, Jincheng Ren, Kangqi Song, thi, Yuandong Tian, Yangyang Shi, Vikas Chan- Xinran Zhou, Boyu Feng, Shudong Liu, Jiabin Luo, dra, and Jürgen Schmidhuber. 2024. Agent-as-a- Weihao Xie, Zhaohui Wang, and 1 others. 2025a. judge: Evaluate agents with agents. arXiv preprint How far are we from genuinely useful deep research arXiv:2410.10934. agents? arXiv preprint arXiv:2512.01948.
Hangfan Zhang, Siyuan Xu, Zhimeng Guo, Huaisheng Yuxin Zuo, Kaiyan Zhang, Li Sheng, Shang Qu, Ganqu Zhu, Shicheng Liu, Xinrun Wang, Qiaosheng Zhang, Cui, Xuekai Zhu, Haozhan Li, Yuchen Zhang, Xin- Yang Chen, Peng Ye, Lei Bai, and Shuyue Hu. 2025b. wei Long, Ermo Hua, Biqing Qi, Youbang Sun, The path of self-evolving large language models: Zhiyuan Ma, Lifan Yuan, Ning Ding, and Bowen Achieving data-efficient learning via intrinsic feed- Zhou. 2025. Ttrl: Test-time reinforcement learning. back. arXiv preprint arXiv:2510.02752. In Advances in Neural Information Processing Sys- tems (NeurIPS 2025). NeurIPS 2025 poster. Hongming Zhang, Xiaoman Pan, Hongwei Wang, Kaixin Ma, Wenhao Yu, and Yu Dong. 2024. Cog- nitive kernel: An open-source agent system towards A Annotation Instructions generalist autopilots. CoRR, abs/2409.10277. This instruction is used for the human annotator Jianshu Zhang, Chengxuan Qian, Haosen Sun, Haoran for summarizing the error points in each erroneous Lu, Dingcheng Wang, Letian Xue, and Han Liu. 2026. trajectory. Progresslm: Towards progress reasoning in vision- language models. arXiv preprint arXiv:2601.15224. Instruction for Error Points Annotation Yifei Zhou, Sergey Levine, Jason Weston, Xian Li, and Sainbayar Sukhbaatar. 2025a. Self- You are given a human execution of a task challenging language model agents. arXiv preprint (which is the ground truth) and an LLM arXiv:2506.01716. agent execution of the same task (which is different from the ground truth). Please Yifei Zhou, Sergey Levine, Jason Weston, Xian Li, and Sainbayar Sukhbaatar. 2025b. Self-challenging lan- compare and explain how LLMs executions guage model agents. In Advances in Neural Informa- are different from human executions, focus- tion Processing Systems (NeurIPS 2025). NeurIPS ing on finding sources, locating information 2025 poster. in the source, drawing observations from Bin Zhu, Qianghuai Jia, Tian Lan, Junyang Ren, Feng sources, problem understanding, etc. Then Gu, Feihu Jiang, Longyue Wang, Zhao Xu, and Wei- summarize the reasons why the LLM made hua Luo. 2026. Marco deepresearch: Unlocking the errors in bullet points with short sen- efficient deep research agents via verification-centric tences based on the comparison. design. arXiv preprint arXiv:2603.28376. 24832 B Agent Prompts Follow-Up Questions
Assume a web-capable research agent
exists. Propose the **fewest** source-
question pairs needed to verify ‘answer‘,
B.1 Decomposition Module using ‘task‘, the [Trajectory Summary], and [Potential Errors].
**Required format (up to 3 pairs):**
Trajectory Summary Prompt
Additional Source 1: source
Summarize each step in the trajectory. For Additional Question 1: a yes-no question
every step, list the online sources visited by based on the source
the agent and the key info obtained from Additional Source 2: source
each source. Additional Question 2: a yes-no question
based on the source
**Required format (repeat “Step N” blocks ...
as needed):** Here are the inputs:
Step 1: [Answer] [Trajectory Summary] [Potential
Source 1: source visited by the agent Errors]
Info 1: information obtained from the
source
Source 2: source visited by the agent
Info 2: information obtained from the
B.2 Verification & Judge Module
source
Step 2:
Verification Agent Prompt
...
Here is the trajectory: Here is a source and question pair. Answer
[Trajectory] the question based on the source.
Source: source
Question: question
Return a brief explanation and concise an-
swer to the question based on the source
Error Identification without any additional text.
Identify suspicious behaviors and map
each to **one** potential error from the
list below. If none, return exactly: ‘No Judge Agent Prompt
potential errors found‘.
You are given a task description, an unveri-
**Potential error list:** fied answer, a summary of how the agent ob-
[Failure Taxonomy] tained the unverified answer, and additional
answers provided by another research agent
**Required format (or the single-line “No regarding the additional questions. Decide
potential errors found”):** if the unverified answer is correct by first
Suspicious Behavior 1: short description providing a concise explanation, then re-
Potential Error 1: one item from the list turning a score between 1 and 4, where: 1
Suspicious Behavior 2: short description = completely incorrect 2 = mostly incorrect
Potential Error 2: one item from the list 3 = mostly correct 4 = completely correct
... Your response should **exactly follow**
Here is the trajectory summary: this format, with no additional content:
[Trajectory Summary] Explanation: explanation Score: score
24833
Corrective Feedback Prompt evidence is missing or unsupported.
You are given a task description, a wrong • Reasoning: The agent’s inference chain
answer given by an agent, a summary of should be logically sound, free from prema-
how the agent obtained the wrong answer, ture conclusions, misinterpretation, or halluci-
and additional answers provided by another nated claims.
research agent regarding the additional ques-
– Excellent: All conclusions follow di-
tions. Now, the agent will try to solve the
rectly and correctly from the retrieved
task again. Based on these inputs, you need
evidence, with no overconfident or un-
to help the agent retrieve the correct an-
supported claims.
swer by first providing a brief reflection and
then providing **no more than three instruc- – Good: Reasoning is largely sound, with
tions**. Note that 1) the agent will strictly only minor inferential gaps or slight over-
follow your instruction; if it cannot get the statements.
correct answer again, which means your in- – Needs Improvement: Noticeable reason-
struction is not useful, then you will be pun- ing errors are present, such as premature
ished. 2) point out necessary sources and conclusions or misinterpretation of evi-
actions to avoid the agent making the same dence.
mistakes again. 3) The agent is good at – Poor: Hallucinated claims, contradictory
understanding clear, concise, and accurate reasoning, or conclusions that conflict
instructions rather than long or complex in- with retrieved evidence.
structions; the latter will confuse it. 4) You
can also suggest the answer to the question • Problem Understanding and Decomposi-
in the instructions if you can determine the tion: The agent should correctly interpret the
answer from available information. Your task and maintain alignment with the original
response should strictly follow this format goal throughout execution.
without any other content: – Excellent: The task is fully and cor-
Reflection: brief reflection rectly understood; all sub-goals are well-
Instruction 1: instruction Instruction 2: in- defined and consistently pursued.
struction ...
– Good: The task is mostly understood,
with minor misalignment or unnecessary
C Verification Rubrics sub-goal expansion. The agent then assesses both the trajectory and the – Needs Improvement: Partial misunder- predicted answer according to the following rubrics standing of instructions leads to goal drift derived from the five major failure categories in this or incomplete task coverage. taxonomy: – Poor: The agent fundamentally misinter- prets the task, producing outputs irrele- • Finding Sources: The agent should consult vant to the original query. specific, authoritative sources and avoid rely- ing on generic or secondary evidence. • Action Execution: Each action should be cor- rectly formatted and directed at the appropri- – Excellent: All key claims are grounded ate modality or interface. in targeted, high-quality sources directly relevant to the query. – Excellent: All actions are executed cor- – Good: Most claims are supported by ap- rectly with no UI, format, or modality propriate sources, with minor reliance on errors. secondary references. – Good: Actions are mostly correct, with – Needs Improvement: Several key claims isolated and non-critical execution er- rely on generic or tangential sources, un- rors. dermining answer reliability. – Needs Improvement: Recurring minor – Poor: Frequent use of vague, unveri- errors in action formatting or modality fied, or inappropriate sources; critical selection hinder progress. 24834 – Poor: Frequent execution failures (e.g., UI errors, wrong tool or modality) that block task completion.
• Trajectory Efficiency: The agent should reach a valid answer within a reasonable num- ber of steps, avoiding unproductive loops caused by early errors. – Excellent: The task is completed well within the step budget, with no unneces- sary detours. – Good: The task is completed within bud- get, with minor inefficiencies that do not affect the final answer. – Needs Improvement: Early errors cas- cade into extended, partially unproduc- tive trajectories, though a result is even- tually produced. – Poor: The step limit is reached without a valid answer, indicating irrecoverable cascading failure.
24835