论文笔记¶
共 18 篇精读笔记,按领域分组(每篇可属多个领域),组内按评分降序、笔记时间降序排列。
验证与奖励建模¶
verifier · 13 篇
- ★★★★★ 2026-07-08 » (2026) LLM-as-a-Verifier: A General-Purpose Verification Framework
- ☆☆☆☆☆ 2026-07-06 » (2024) Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
- ☆☆☆☆☆ 2026-07-03 » (2026) Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
- ☆☆☆☆☆ 2026-07-03 » (2026) Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification
- ☆☆☆☆☆ 2026-07-03 » (2026) Variation in Verification: Understanding Verification Dynamics in Large Language Models
- ☆☆☆☆☆ 2026-07-03 » (2025) Critique to Verify: Accurate and Honest Test-Time Scaling with RL-Trained Verifiers
- ☆☆☆☆☆ 2026-07-03 » (2026) Scaling Agentic Verifier for Competitive Coding
- ☆☆☆☆☆ 2026-07-02 » (2021) Training Verifiers to Solve Math Word Problems
- ☆☆☆☆☆ 2026-07-02 » (2022) Solving math word problems with process- and outcome-based feedback
- ☆☆☆☆☆ 2026-07-02 » (2023) Let's Verify Step by Step
- ☆☆☆☆☆ 2026-07-02 » (2024) OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning
- ☆☆☆☆☆ 2026-07-02 » (2025) Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
- ☆☆☆☆☆ 2026-07-02 » (2025) Logical Reasoning with Outcome Reward Models for Test-Time Scaling
数学推理¶
math · 7 篇
- ☆☆☆☆☆ 2026-07-16 » (2025) FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
- ☆☆☆☆☆ 2026-07-03 » (2026) Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
- ☆☆☆☆☆ 2026-07-02 » (2021) Training Verifiers to Solve Math Word Problems
- ☆☆☆☆☆ 2026-07-02 » (2022) Solving math word problems with process- and outcome-based feedback
- ☆☆☆☆☆ 2026-07-02 » (2023) Let's Verify Step by Step
- ☆☆☆☆☆ 2026-07-02 » (2024) OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning
- ☆☆☆☆☆ 2026-07-02 » (2025) Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
逻辑推理¶
logic · 1 篇
- ☆☆☆☆☆ 2026-07-02 » (2025) Logical Reasoning with Outcome Reward Models for Test-Time Scaling
竞技编程¶
coding · 1 篇
- ☆☆☆☆☆ 2026-07-03 » (2026) Scaling Agentic Verifier for Competitive Coding
创意写作 / 叙事¶
narrative · 1 篇
- ☆☆☆☆☆ 2026-07-21 » (-) Lightweight Latent Reasoning for Narrative Tasks
Agent¶
agent · 3 篇
- ☆☆☆☆☆ 2026-07-13 » (-) SAO 精读:让长时程 Agent 的异步 RL 真正训得动
- ☆☆☆☆☆ 2026-07-03 » (-) MaineCoon 精简解析
- ☆☆☆☆☆ 2026-07-02 » (-) MemSlides 精简解析
多模态¶
multimodal · 1 篇
- ☆☆☆☆☆ 2026-07-03 » (-) MaineCoon 精简解析