Pretrain Practice›Quantifying the Per-layer Contribution of LLMs
Oct 3, 2026

Quantifying the Per-layer Contribution of LLMs

By Wenhao Chai作者 Wenhao Chai

Setup设置

The baseline is the MuonH Qwen3 recipe of the Marin speedrun, from 130m to 1.2B. We add a next-token loss after every layer and judge each layer by its final training loss against backbone compute 6ND, where N counts the parameters of the layers up to that one, without the embedding and the heads, and D the training tokens. Code

基线是 Marin speedrun 里的 MuonH Qwen3 配方,规模从 130m 到 1.2B。我们在每一层后面加一个预测下一个词的损失,用每层的最终训练损失来衡量;横轴是骨干算力 6ND,N 是到这一层为止各层的参数量,不算 embedding 和头,D 是训练 token 数。代码

--- baseline loss (simplified)+++ per-layer losses (simplified; the Code link has the real code) def loss_fn(model, tokens):     targets = tokens[1:]     x = model.embed(tokens[:-1])-    for block in model.blocks:+    loss = 0.0+    for k, block in enumerate(model.blocks[:-1]):         x = block(x)-    return cross_entropy(model.head(model.norm(x)), targets)+        # probes: read the layer, but send no gradient into the model+        h = stop_gradient(x) if probe else x+        # every earlier layer has its own norm and output head, and predicts the next token+        logits = model.layer_heads[k](model.layer_norms[k](h))+        loss += cross_entropy(logits, targets)+    x = model.blocks[-1](x)+    # the usual loss at the last layer, at the same weight+    return loss + cross_entropy(model.head(model.norm(x)), targets)
--- 基线损失(简化)+++ 逐层损失(简化;真实代码见“代码”链接) def loss_fn(model, tokens):     targets = tokens[1:]     x = model.embed(tokens[:-1])-    for block in model.blocks:+    loss = 0.0+    for k, block in enumerate(model.blocks[:-1]):         x = block(x)-    return cross_entropy(model.head(model.norm(x)), targets)+        # 探针:读取这一层,但梯度不传回模型+        h = stop_gradient(x) if probe else x+        # 前面每一层都有自己的 norm 和输出头,也预测下一个词+        logits = model.layer_heads[k](model.layer_norms[k](h))+        loss += cross_entropy(logits, targets)+    x = model.blocks[-1](x)+    # 最后一层仍是原来的损失,权重相同+    return loss + cross_entropy(model.head(model.norm(x)), targets)

Experiments实验

Shared head共享头

Every layer predicts through the model's output head, and every layer's loss trains that head.每一层都经过模型的输出头预测,每一层的损失都训练这个头。

Shared head, stop-grad共享头 stop-grad

Every layer predicts through the output head, but only the last layer's loss trains it.每一层都经过输出头预测,但只有最后一层的损失训练这个头。

Separate heads独立头 Main主做法

Every layer predicts through a head of its own.每一层都用自己的头预测。

Probes only只加探针 Baseline基线

Every layer predicts through a head of its own, but no gradient flows back into the model: it trains exactly like the baseline.每一层都用自己的头预测,但梯度不传回模型:模型和基线训练得完全一样。

Baseline by depth各深度的基线

An ordinary model with fewer layers, 2 to 6, trained on its own with only its output head. Layer k is compared with the model of k layers.层数更少的普通模型,2 到 6 层,单独训练,只有自己的输出头。第 k 层和 k 层的模型比较。

Pre-LN

Normalize, transform, add back to the stream.先归一化,再变换,加回残差流。

Sandwich-LN Baseline基线

A second normalization on the branch output before it is added: the baseline's own design.分支输出加回之前再归一化一次:基线自己的设计。

LayerNorm Scaling

The normalized input is scaled by 1/√ℓ, more in deeper layers.归一化后的输入乘以 1/√ℓ,越深的层缩得越多。

DeepNorm

No norm before the branch; the stream is scaled up by (2L)^¼ and normalized after the add.分支前不归一化;残差流乘以 (2L)^¼,相加后再归一化。

KEEL

Norms before and after, with a residual gain equal to the number of sublayers.前后都归一化,残差增益等于子层数。

Hyper-Connections

Four residual streams, mixed by learned weights, read into the branch and written back.四条残差流,由可学习的权重混合,读入分支再写回。

mHC

Hyper-connections with the stream mixing kept doubly stochastic.超连接,但流之间的混合矩阵保持双随机。

AttnRes (Full)

A layer's input is a learned softmax mix of every earlier output and the embedding.每层的输入是此前所有层输出和 embedding 的 softmax 加权混合。

AttnRes (Block)

The same mix over 8 block outputs and the running sum of the current block.同样的混合,只在 8 个块的输出和当前块的累加和上做。

MoDA

Attention also reads the same token's keys and values from earlier layers.注意力还读取同一 token 在之前各层的 key 和 value。

Takeaways核心结论

Acknowledgements致谢

The author is pleased to acknowledge that the work reported on in this post was substantially performed using the Princeton Research Computing resources at Princeton University. Princeton Research Computing is a consortium of groups including the Princeton Institute for Computational Science and Engineering (PICSciE) and Research Computing at Princeton University.

本文报告的工作主要使用普林斯顿大学 Princeton Research Computing 的计算资源完成。Princeton Research Computing 是由普林斯顿计算科学与工程研究所(PICSciE)和普林斯顿大学 Research Computing 等团队组成的联合体。

Citation

@misc{chai2026perlayer,
  title        = {Quantifying the Per-layer Contribution of LLMs},
  author       = {Chai, Wenhao},
  year         = {2026},
  howpublished = {Blog post},
  url          = {https://wenhaochai.com/blogs/per-layer-contribution.html}
}
Pretrain Practice Back to top回到顶部