Experiment Template
- Blog prose for an outside reader (load
writing:style): no run names, arm codes or flags; say what each setup does. - The title names what the experiment measures, never its result; the card uses it. Cards run newest first, the grey template card last.
- English and Chinese for every text element.
- No standfirst and no Takeaways until right before publishing; the Takeaways sentence is also the description and the card text.
- While runs go on, the draft is a dashboard: every running or planned figure is already on the page with its axes, legend and a status word (Running, Queued, Planned), and fills in when its runs finish. A figure whose runs are dropped is deleted. Publish only when no placeholder is left.
- Sections in order: Cover, Setup, Experiments, Takeaways, Related work, Acknowledgements, Citation.
- 写给外部读者的博客(先加载
writing:style):不出现 run 名、分组代号或开关名,说清每种做法做了什么。 - 标题写实验衡量什么,不写结果;卡片用同一标题。卡片按新到旧排,灰色模板卡片放最后。
- 每个文字元素都有中英两版。
- 导语和 Takeaways 留到发布前再写;Takeaways 那一句同时用作页面简介和卡片文字。
- 作业还在跑时,草稿就是看板:运行中或计划中的图先放上页面,画好坐标轴和图例,面板中间写状态(运行中、排队中、计划中),作业跑完就填上数据。不再跑的图直接删掉。没有占位图了才发布。
- 各节顺序:Cover、Setup、Experiments、Takeaways、Related work、Acknowledgements、Citation。
Cover封面
- Cover: a schematic of the setup, never a result, used only as the card picture. Inline SVG 1600×1000, no text, Anthropic style: ivory
#F0EEE6, slate#141413outlines at stroke 7, wobble filter 0.012 / 4.5 withfilterUnits="userSpaceOnUse"and a unique id, fills offset +12 in oat, manilla, kraft and clay. - Animation: the same picture in motion, one part per setup, dots for the current part, no text or caption.
render(t)+render_anim.py, placed asfigure.ar-video#overview, first on the page. - Card text: the Takeaways sentence, word for word; empty while Takeaways is empty.
- 封面:实验做法的示意图,不画结果,只用作卡片图。内联 SVG,1600×1000,无文字,Anthropic 风格:象牙色底
#F0EEE6,深石板色#141413轮廓,线宽 7,抖动滤镜 0.012 / 4.5、filterUnits="userSpaceOnUse"、id 不重复,色块偏移 +12,用燕麦、马尼拉、牛皮纸和陶土色。 - 动画:同一张图动起来,比较几种做法就分几段,小点表示当前段,没有文字和图注。用
render(t)加render_anim.py,放成figure.ar-video#overview,页面第一项。 - 卡片文字:和页面的 Takeaways 一字不差;Takeaways 没写时卡片也留空。
Setup设置
[The baseline exactly, its name linked to the original recipe: the recipe and the sizes it covers. What we change, and the metric.] Code
【准确写出基线,名字链接到原始配方:配方和它覆盖的规模。改了什么、用什么指标衡量。】代码
--- baseline [function] (simplified)+++ ours (simplified; the Code link has the real code) [unchanged line]-[baseline line]+ # [what this step does]+[our line]
- One or two sentences: the baseline by its exact name, linked to the original recipe at a fixed commit, with the full range of sizes it covers; what we change; the metric, with the compute axis defined once. Sizes and tokens of the runs go to the figures. End with a link to the exact public commit of the runs.
- One diff of the core change and of any switch the figures compare, rewritten for reading: baseline red, ours green, simplified alike, one comment per changed step in each language, a header saying it is simplified. Rendered by
make_diff_html.py(Pygments, inline mode) between thediff:SLUGmarkers. - Controls that need no code change go in the captions of the figures that use them.
- 一两句话:写准基线名字,链接到原始配方的固定提交,并写出它覆盖的全部规模;写改了什么;写指标,并在这里定义一次算力轴。本实验跑的规模和 token 数归图。最后链接到实验所用的公开提交。
- 一个 diff,只放核心改动和图中比较的开关,改写成最好读的样子:基线标红,我们的标绿,两边同样简化,每个改动步骤配一行中英注释,文件头注明已简化。由
make_diff_html.py(Pygments,inline 模式)生成在diff:SLUG标记之间。 - 不改代码的对照组写在用到它的图的图注里。
- One branch per experiment on the public fork
wenhaochai/marin-pli, named after the slug, carrying the history as it was run. - In a fresh clone,
scrub.sh(git filter-repo) replaces only the cluster user, group directory and Slurm account withUSER,GROUPandgroup, in contents, archives, file names and messages. Scan every blob and archive member until keys and identifiers are zero. The private original stays onarchive. A.gitignoreblock keeps credentials and job output out. - The owner approves each public push; the Code link goes live after it.
- 每个实验一个分支,推到公开的
wenhaochai/marin-pli,分支名用 slug,历史保持实际运行时的样子。 - 在新克隆里用
scrub.sh(git filter-repo)只把集群用户名、组目录和 Slurm 账户换成USER、GROUP和group,文件内容、压缩包、文件名和提交信息都换。扫描所有文件和压缩包成员,直到密钥和标识都为 0。私有原始历史留在archive。.gitignore挡住凭据和作业输出。 - 每次公开推送都要作者批准,推送后代码链接才生效。
Experiments实验
- Before the first figure, a schematic in the cover style: one panel per setup the page uses, controls included, each with its name and one sentence; tags mark the main setup (Main) and the setup the main one is measured against (Baseline).
- The first figure compares every setup on the raw metric with the baseline; a setup equal to the baseline by construction is left out. Each later figure makes one point, with the data processed around it. Once the page settles on a setup, later scales and variants follow that setup only.
- The compute axis is backbone compute 6ND: N the parameters of the layers counted, without the embedding and the heads; D the training tokens.
- A model outside the baseline's named sizes gives its shape in the title (layers, width) and its MLP width, heads and parameter counts (backbone, embedding, heads) in the caption.
- Charts are
figure.vizdrawn byassets/article-charts.jsfrom a generated file inassets/data/, never typed numbers, under thewriting:plotrules. Titles and captions in both languages, a caption in two lines at most; no running text in Experiments that lists numbers or restates what a figure shows: each figure's title and caption carry its point.
- 第一张图之前放一张封面风格的示意图:页面用到的每种做法一格,包括对照组,每格配名字和一句话;用标签标出主做法(主做法)和它所对照的做法(基线)。
- 第一张图用原始指标比较所有做法和基线;按构造等于基线的做法不画。之后每张图只讲一件事,数据围绕它处理。页面定下主做法以后,更大规模和变体实验只跑这一种做法。
- 算力轴用骨干算力 6ND:N 是计入的各层参数量,不算 embedding 和头;D 是训练 token 数。
- 基线命名规模之外的模型,标题写形状(层数、宽度),图注写 MLP 宽度、头数和参数量(骨干、embedding、头)。
- 图用
figure.viz,由assets/article-charts.js根据assets/data/里脚本生成的文件绘制,数字不手填,遵守writing:plot。标题和图注中英双语,图注最多两行;实验部分不写罗列数字、复述图中结果的正文:每张图的结论由它的标题和图注说清。
Takeaways核心结论
[Empty while drafting. Right before publishing: one precise sentence with the result, negatives included; the same sentence is the page's description.]
【起草时留空。发布前再写:一句准确的话给出结果,负面结果也写;同一句话用作页面简介。】
Acknowledgements致谢
The author is pleased to acknowledge that the work reported on in this post was substantially performed using the Princeton Research Computing resources at Princeton University. Princeton Research Computing is a consortium of groups including the Princeton Institute for Computational Science and Engineering (PICSciE) and Research Computing at Princeton University.
本文报告的工作主要使用普林斯顿大学 Princeton Research Computing 的计算资源完成。Princeton Research Computing 是由普林斯顿计算科学与工程研究所(PICSciE)和普林斯顿大学 Research Computing 等团队组成的联合体。
Citation
@misc{chai2026SLUG,
title = {[Title]},
author = {Chai, Wenhao},
year = {2026},
howpublished = {Blog post},
url = {https://wenhaochai.com/blogs/SLUG.html}
}