论文笔记¶
共 42 篇(21 篇已精读,21 篇暂无笔记)。点击标题进入论文页;行尾「笔记」「译文」直达对应内容(缺失则置灰)。
Verification¶
verification · 11 篇
★★★★★2607.05391v2
LLM-as-a-Verifier: A General-Purpose Verification Framework
Trajectory Scoring
★★★★★2510.13744v1
★★★★★2601.15808v2
★★★★★2509.17995
★★★★★2509.23152
★★★★★2602.04254
Scaling Agentic Verifier for Competitive Coding
CodingCounterexample Search
★★★★★2110.14168
Training Verifiers to Solve Math Word Problems
MathOutcome Verifier
★★★★★2211.14275
Solving math word problems with process- and outcome-based feedback
MathProcess vs Outcome
★★★★★2305.20050
Let's Verify Step by Step
MathProcess Reward Model
★★★★★2311.09724
★★★★★2508.19903
Logical Reasoning with Outcome Reward Models for Test-Time Scaling
LogicOutcome Reward Model
创意写作 / 叙事¶
narrative · 4 篇
★★★★★2601.09609
DPWriter:面向创意写作的多样化规划分支强化学习
RL-trainedPlanning
★★★★★2504.11900
★★★★★2603.05890
Lost in Stories:LLM 长篇故事生成中的一致性错误
Evaluation
★★★★★2512.02240
Lightweight Latent Reasoning for Narrative Tasks
Latent Reasoning
评测与 Benchmark¶
evaluation · 1 篇
★★★★★2411.04872v7
Agent RL¶
agent_rl · 1 篇
★★★★★2607.07508
SAO 精读:让长时程 Agent 的异步 RL 真正训得动
Asynchronous RLSingle Rollout
RL / Reward Learning¶
general_rl · 2 篇
★★★★★2410.18451
★★★★★2502.06781
视频生成¶
video_generation · 1 篇
★★★★★2606.17800
MaineCoon 精简解析
Real-time GenerationAudio-visual
幻灯片生成¶
slides · 1 篇
★★★★★2606.17162
MemSlides 精简解析
PersonalizationHierarchical Memory
待精读¶
pending · 21 篇
★★★★★2502.10391
★★★★★2503.05244
★★★★★2504.03622
★★★★★2506.00103
★★★★★2506.05760
★★★★★2508.18642
★★★★★2508.21476
★★★★★2509.02534
★★★★★2511.10643
★★★★★2511.22570
★★★★★2601.07149
★★★★★2601.09858
★★★★★2601.17226
★★★★★2602.01511
★★★★★2603.15061
★★★★★2604.03004
★★★★★2604.05517
★★★★★2604.19071
★★★★★2604.27453
★★★★★2606.26300
★★★★★2606.29354