How to design reward function in Multi-Agent RL多智能体 RL 中如何设计奖励函数
Setup设置
The baseline is single-agent EasyPPO with its Frontier-CS settings: its Qwen3.5-9B SFT initialization, trained on 200 FrontierSmith problems and validated on the 172 Frontier-CS algorithmic problems, scored 0–100 as the mean per-test-case score by a local judge that follows the official rules; the last C++ block of an answer is the program judged, and every prompt says so. We replace the single agent with a team, one lead and four subagents as in DeepSeek-V4.1-Flash's Agent Team mode: the subagents are helpers that may test programs against the judge and report ideas, analysis or code, and only the lead's final submission scores the team. We compare ways to split the team's reward; compute is counted in generated tokens and latency in tokens on the critical path. The first four questions below asked the same thing on competition math with FlashREINFORCE. Code: the team game, judge and scripts, the miles fork and the math runs; the experiment record links the exact commit of each finished run.
基线是单智能体 EasyPPO,用它的 Frontier-CS 设置:它的 Qwen3.5-9B SFT 初始化,在 200 道 FrontierSmith 题上训练,在 172 道 Frontier-CS 算法题上验证,由按官方规则实现的本地评测器打 0–100 分,即各测试用例得分的平均;评测的是回答里最后一个 C++ 代码块,每个 prompt 都写明了这一点。我们把单智能体换成团队,按 DeepSeek-V4.1-Flash 的 Agent Team 模式设为 1 个 lead 加 4 个 subagent:subagent 是帮手,可以把程序交给评测器自测,再汇报思路、分析或代码;只有 lead 的最终提交决定团队成绩。我们比较团队奖励的不同分法;算力按生成的 token 数计,延迟按关键路径上的 token 数计。下面前 4 个问题是在竞赛数学上用 FlashREINFORCE 问的同一件事。代码:团队游戏、评测器和脚本、miles 分支和数学实验;实验记录里链接了每个跑完的 run 的准确提交。
--- baseline: one agent, one judged program (simplified)+++ ours: a team game with judge feedback (simplified) def episode(problem):- return judge(problem, agent.solve(problem)).score+ # the lead writes one task per subagent# lead 给每个 subagent 写一条任务+ tasks = lead.plan(problem, n=4)+ # each subagent may test up to 3 programs, sees each result, then reports# 每个 subagent 最多自测 3 个程序,看到每次结果后写报告+ reports = [sub.explore(problem, t, judge, tests=3) for t in tasks]+ # the lead reads every report and submits the team's program# lead 读完所有报告,提交团队的程序+ final = judge(problem, lead.solve(problem, reports))+ # only the final submission scores the team; the split is what we compare# 只有最终提交决定团队成绩;比较的是奖励怎么分+ return split(final, reports)
Experiments实验
Answered已有结论Does FlashREINFORCE's quick-start recipe reproduce the paper on R1-Distill-1.5B?FlashREINFORCE 的快速入门配方能否在 R1-Distill-1.5B 上复现论文?
Yes: our run follows the paper's trust-region arm through training and ends only slightly below it.
能复现:我们的 run 一路跟住论文的信任域组,结束时只略低一点。
One agent, asynchronous REINFORCE with a sequence trust region: the quick-start recipe, run verbatim.单个智能体,带序列信任域的异步 REINFORCE:快速入门配方原样运行。
The two arms the paper reports: with the trust region, and importance sampling only.论文报告的两组:带信任域的,以及只做重要性采样的。
Answered已有结论Does score centering remove the train/inference mismatch, and does it buy held-out accuracy?Score centering 能否消除训练/推理不一致?它能提高保留集准确率吗?
It removes the mismatch, but held-out accuracy does not improve.
能消除不一致,但保留集准确率没有提高。
Subtracts the expected score under the sampler's top-32 head, so the importance weight can be one; the trust region stays. Measured against the FlashREINFORCE recipe above.减去采样器 top-32 头下的期望得分,因此重要性权重可以取 1;信任域保留。对照是上面的 FlashREINFORCE 配方。
The same code, seed and weight with centering off.代码、种子和权重都相同,只关掉 centering。
Answered已有结论Does talking help a team of agents on math, and which credit makes it help?在数学题上,智能体之间交流有没有帮助?哪种功劳分配让它有帮助?
Untrained talking hurts, because agents copy each other; only paying each turn for its own gain raised held-out accuracy.
未经训练的交流有害,因为智能体会互相照抄;只有按每一轮自己的增益给奖励,保留集准确率才提高。
Four copies of one policy answer, read each other's notes, and answer again.同一策略的 4 个副本先作答,读彼此的笔记,再作答一次。
The same four answers at the same tokens, without the notes.同样 4 次作答、同样的 token,但不给笔记。
A fifth agent reads four solvers' answers and gives the team's answer.第 5 个智能体读 4 个解答者的答案,给出团队答案。
| Study研究 | Setup做法 | Result against the pre-registered bar对照事先登记门槛的结果 |
|---|---|---|
| 1 | four agents answer, read each other's notes, answer again; scored if any is right4 个智能体作答,读彼此笔记后再答;任一答对即得分 | no gain (−0.97 points, t −1.45)无提升(−0.97 分,t −1.45) |
| 2 | the same game on a hard subset同一游戏换到难题子集 | untrained talking lowers team coverage by 3.9 points: agents copy each other未训练时交流让团队覆盖率降 3.9 分:智能体互相照抄 |
| 3 | coverage or plurality vote, broadcast or Shapley credit覆盖率或多数投票,广播或 Shapley 分配 | both below the bar (+0.02 and +1.58; bar +2.3)都低于门槛(+0.02 和 +1.58;门槛 +2.3) |
| 4 | a cleaner note protocol更干净的笔记协议 | coverage still −4.44 points: the copying is intrinsic覆盖率仍低 4.44 分:照抄是内在的 |
| 5 | credit per turn: round 2 is paid only for its gain over round 1逐轮功劳:第 2 轮只按相对第 1 轮的提升给奖励 | +6.2 points over silent training on AIME24 (8 of 8 points); vote and own credit stay below the barAIME24 上比静默训练高 6.2 分(8 个点全部更高);投票和个人功劳都低于门槛 |
| 6 | a trained reader that aggregates four solvers训练一个汇总 4 个解答者的 reader | no better than an untrained reader (+0.3 ± 2.4; bar +2)不比未训练的 reader 好(+0.3 ± 2.4;门槛 +2) |
| 7 | a 27B lead with up to three teammates, trained as in DeepSeek-V4.1-Flash27B 的 lead 带最多 3 个队友,照 DeepSeek-V4.1-Flash 训练 | not met: +0.29 [−1.52, +2.15] with the collaboration bonus, +1.20 [−0.88, +3.30] without未达到:加协作奖励 +0.29 [−1.52, +2.15],不加 +1.20 [−0.88, +3.30] |
Answered已有结论Does DeepSeek-style lead-and-teammates training beat one agent at the same budget?DeepSeek 式的 lead 加队友训练,能否胜过同预算的单智能体?
No: the collaboration bonus makes the lead delegate more often, but the team does not beat the same lead working alone.
不能:协作奖励让 lead 更常委派,但团队没有胜过同一个 lead 单干。
A Qwen3.8-27B lead delegates to up to three teammates and answers; task reward plus a collaboration bonus minus latency, as in DeepSeek-V4.1-Flash.Qwen3.8-27B 的 lead 把子任务委派给最多 3 个队友,再给出答案;奖励是任务分加协作奖励减延迟,照 DeepSeek-V4.1-Flash。
One agent with the same model and budget.同一个模型、同样预算的单个智能体。
Answered已有结论Can a local judge stand in for the official Frontier-CS judge?本地评测器能否代替官方的 Frontier-CS 评测器?
Yes: it scores like the official judge on classic and interactive problems, and contestant code cannot read the test answers.
能:它在普通题和交互题上的打分都与官方评测器一致,而且选手代码读不到测试答案。
The official rules re-implemented, with contestant code sandboxed and every case's output deleted as soon as it is scored.按官方规则重新实现,选手代码在沙箱里运行,每个测试点判完立即删除输出。
Frontier-CS's own engine, run on the same node.Frontier-CS 自己的评测引擎,在同一节点上运行。
Answered已有结论Which starting point can RL train from on Frontier-CS?在 Frontier-CS 上,RL 能从哪个起点开始训练?
EasyPPO's SFT initialization: it finishes most of its answers, while the base model with reasoning mostly runs out of tokens.
EasyPPO 的 SFT 初始化:它大部分回答都能写完,而原版模型开启思考时大多用完了 token。
Qwen3.5-9B with its reasoning on.开启思考的 Qwen3.5-9B。
The same weights answering directly.同一份权重,直接作答。
Qwen3.5-9B fine-tuned on DeepSeek-V3.1 solutions by the EasyPPO authors.EasyPPO 作者用 DeepSeek-V3.1 的解答微调过的 Qwen3.5-9B。
Testing检验中Does our EasyPPO port train a single agent on Frontier-CS?我们移植的 EasyPPO 能否在 Frontier-CS 上训练单智能体?
One agent writes one program per problem; its judged score is the reward; a critic, trained alone for the first 30 rollouts.一个智能体每题写一个程序,评测分数即奖励;带 critic,前 30 个 rollout 只训练 critic。
Testing检验中Does a team without feedback beat one agent?没有反馈的团队能否胜过单智能体?
Three teammates write programs; the lead reads them without any scores and writes the final one.3 个队友各写程序;lead 看不到任何分数,读完后写最终程序。
A single agent given the team's token budget.给单个智能体和团队同样的 token 预算。
Testing检验中With judge feedback, which reward split makes a team beat one agent at equal tokens?有评测反馈时,哪种奖励分法能让团队在相同 token 下胜过单智能体?
The lead writes four tasks; each subagent may test up to three programs, sees each result and reports; the lead's final submission is the team's. Measured against the EasyPPO single agent above.lead 写 4 条任务;每个 subagent 最多自测 3 个程序,看到每次结果后写报告;lead 的最终提交就是团队的提交。对照是上面的 EasyPPO 单智能体。
One agent submits five times in one chat, seeing each result; only the fifth counts.一个智能体在同一对话里提交 5 次,每次都看到结果;只有第 5 次算数。
One agent submits five times independently; the best counts.一个智能体独立提交 5 次,取最高分。
Testing检验中Why do most programs fail to compile, and does a prompt, the base weights or a lower temperature fix it?为什么大部分程序编译不过?换 prompt、换权重或降低温度能否解决?
The SFT initialization with the problem prompt alone.SFT 初始化,只给题目 prompt。
Declare every name, match every type, close every brace.声明每个名字,类型一一对应,括号全部闭合。
The same checklist appended to the problem.同一份清单附在题目之后。
The base model, reasoning on, with the system message.原版模型开启思考,加上同一条 system 消息。
The SFT initialization sampled cooler.SFT 初始化,以较低温度采样。
Takeaways核心结论
Acknowledgements致谢
The author is pleased to acknowledge that the work reported on in this post was substantially performed using the Princeton Research Computing resources at Princeton University. Princeton Research Computing is a consortium of groups including the Princeton Institute for Computational Science and Engineering (PICSciE) and Research Computing at Princeton University.
本文报告的工作主要使用普林斯顿大学 Princeton Research Computing 的计算资源完成。Princeton Research Computing 是由普林斯顿计算科学与工程研究所(PICSciE)和普林斯顿大学 Research Computing 等团队组成的联合体。
Experiment record实验记录
Citation
@misc{chai2026marl,
title = {How to design reward function in Multi-Agent RL},
author = {Chai, Wenhao},
year = {2026},
howpublished = {Blog post},
url = {https://wenhaochai.com/blogs/multi-agent-reward.html}
}