跳转至

论文笔记

共 42 篇(21 篇已精读,21 篇暂无笔记)。点击标题进入论文页;行尾「笔记」「译文」直达对应内容(缺失则置灰)。

Verification

verification · 11 篇

★★★★★2607.05391v2
LLM-as-a-Verifier: A General-Purpose Verification Framework
Stanford University;UC Berkeley;NVIDIA Research
Trajectory Scoring
2026-07-28
★★★★★2510.13744v1
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
ACL 2026(Long) · Salesforce AI Research
EvaluationMath
2026-07-03
★★★★★2601.15808v2
Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification
Findings of ACL 2026 · The Chinese University of Hong Kong;Tencent AI Lab;Singapore Management University;Renmin University of China
Deep ResearchRubric-guided
2026-07-03
★★★★★2509.17995
Variation in Verification: Understanding Verification Dynamics in Large Language Models
ICLR 2026 · Salesforce AI Research;Dartmouth College;University of Illinois Urbana-Champaign
Evaluation
2026-07-03
★★★★★2509.23152
Critique to Verify: Accurate and Honest Test-Time Scaling with RL-Trained Verifiers
The Hong Kong University of Science and Technology (Guangzhou);The Hong Kong University of Science and Technology;ETH AI Center, ETH Zurich;University of California, Merced;Sun Yat-sen University;MBZUAI
MathCritique-based
2026-07-03
★★★★★2602.04254
Scaling Agentic Verifier for Competitive Coding
Renmin University of China;Qwen Team, Alibaba Group
CodingCounterexample Search
2026-07-03
★★★★★2110.14168
2026-07-02
★★★★★2211.14275
2026-07-02
★★★★★2305.20050
Let's Verify Step by Step
ICLR 2024 · OpenAI
MathProcess Reward Model
2026-07-02
★★★★★2311.09724
OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning
NAACL Findings · The Chinese University of Hong Kong, Shenzhen;Shenzhen Research Institute of Big Data
MathValue Model
2026-07-02
★★★★★2508.19903
Logical Reasoning with Outcome Reward Models for Test-Time Scaling
EMNLP 2025 · Monash University;VinUniversity
LogicOutcome Reward Model
2026-07-02

创意写作 / 叙事

narrative · 4 篇

★★★★★2601.09609
DPWriter:面向创意写作的多样化规划分支强化学习
中国人民大学;快手科技(一作在快手实习期间完成)
RL-trainedPlanning
2026-07-27
★★★★2504.11900
Finding Flawed Fictions:通过情节漏洞检测评估语言模型的复杂推理
University of Washington,Paul G. Allen School of Computer Science & Engineering
Evaluation
2026-07-28
★★★★2603.05890
Lost in Stories:LLM 长篇故事生成中的一致性错误
Microsoft Beijing;Singapore University of Technology and Design
Evaluation
2026-07-28
★★★★★2512.02240
Lightweight Latent Reasoning for Narrative Tasks
University of Edinburgh;CIFAR Fellow
Latent Reasoning
2026-07-21

评测与 Benchmark

evaluation · 1 篇

★★★★★2411.04872v7
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Epoch AI;贡献数学家来自 King's College London、MIT、University of Siegen、ICMC/USP、UC Berkeley、Cornell、Harvard、University of Leicester、University of Bristol、Rutgers、ETH Zurich 等机构
Math
2026-07-16

Agent RL

agent_rl · 1 篇

★★★★★2607.07508
SAO 精读:让长时程 Agent 的异步 RL 真正训得动
Tsinghua University
Asynchronous RLSingle Rollout
2026-07-13

RL / Reward Learning

general_rl · 2 篇

★★★★★2410.18451
2026-07-06
★★★★★2502.06781
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
Shanghai AI Laboratory;Shanghai Jiao Tong University;MMLab, The Chinese University of Hong Kong;HKGAI under InnoHK
VerificationMath
2026-07-02

视频生成

video_generation · 1 篇

★★★★★2606.17800
MaineCoon 精简解析
Catnip AI Team
Real-time GenerationAudio-visual
2026-07-03

幻灯片生成

slides · 1 篇

★★★★★2606.17162
MemSlides 精简解析
Beijing University of Posts and Telecommunications;Tsinghua University;Shanghai Jiao Tong University
PersonalizationHierarchical Memory
2026-07-02

待精读

pending · 21 篇

★★★★★2506.00103
★★★★★2506.05760
★★★★★2508.18642
★★★★★2508.21476
★★★★★2601.07149
★★★★★2601.17226
★★★★★2602.01511
★★★★★2603.15061
★★★★★2604.03004
★★★★★2604.05517
★★★★★2604.27453