Experiments›How Vocabulary Size Shapes Overfitting
Oct 7, 2026

How Vocabulary Size Shapes Overfitting

By Wenhao Chai作者 Wenhao Chai

Setup设置

The baseline is the MuonH Qwen3 recipe of the Marin speedrun, which covers 130m to 1.2B parameters, with the Marin tokenizer of 128K tokens. We cut the tokenizer to the first 64K, 32K, 16K or 8K tokens its BPE training learned. The smallest vocabulary is the 256 bytes, the tokenizer left with no merges at all. We train each model on the same text: all of fineweb-edu-10B, or a random part of it repeated 2.4 or 8 times. A smaller vocabulary cuts the same text into more tokens, so its runs take proportionally more steps. We compare Paloma bits per byte, the loss in bits divided by the UTF-8 bytes of the text, so tokenizers of different sizes share one scale. The runs use this code.

基线是 Marin speedrun 的 MuonH Qwen3 配方,覆盖 130m 到 1.2B 参数,使用 128K 个词的 Marin tokenizer。我们把 tokenizer 截到它的 BPE 训练最先学到的 64K、32K、16K 或 8K 个词。最小的词表是 256 个字节,也就是一个合并都不剩的 tokenizer。我们让每个模型训练同样的文本:全部 fineweb-edu-10B,或其中随机的一部分重复 2.4 遍或 8 遍。词表越小,同样的文本被切成越多的词,所以它的 run 按比例多训练一些步。我们比较 Paloma 的每字节比特数,即以 bit 计的损失除以文本的 UTF-8 字节数,所以不同大小的词表在同一个尺度上比较。所有 run 使用这份代码。

--- baseline (simplified)+++ truncated tokenizer (simplified; the Code link has the real code) tok = load_tokenizer("marin-tokenizer")+# keep the first K tokens BPE learned and the merges among them+tok.vocab = {t: i for t, i in tok.vocab.items() if i < K}+tok.merges = [(a, b) for a, b in tok.merges if a + b in tok.vocab] tokens = tok.encode(text)-model = Qwen3(vocab_size=tok.vocab_size)-steps = recipe.steps+model = Qwen3(vocab_size=K)+# the same text is more tokens: scale the steps by the token ratio+steps = round(recipe.steps * len(tokens) / len(marin_tokens)) train(model, tokens, steps)+# compare tokenizers per byte of text, not per token+bits_per_byte = paloma_loss_nats * tokens_per_doc / (log(2) * utf8_bytes_per_doc)
--- 基线(简化)+++ 截短的 tokenizer(简化;真实代码见“代码”链接) tok = load_tokenizer("marin-tokenizer")+# 保留 BPE 最先学到的 K 个词,以及它们之间的 merge+tok.vocab = {t: i for t, i in tok.vocab.items() if i < K}+tok.merges = [(a, b) for a, b in tok.merges if a + b in tok.vocab] tokens = tok.encode(text)-model = Qwen3(vocab_size=tok.vocab_size)-steps = recipe.steps+model = Qwen3(vocab_size=K)+# 同样的文本变成更多的词:步数按词数比例放大+steps = round(recipe.steps * len(tokens) / len(marin_tokens)) train(model, tokens, steps)+# 按文本的字节比较不同 tokenizer,而不是按词+bits_per_byte = paloma_loss_nats * tokens_per_doc / (log(2) * utf8_bytes_per_doc)

Experiments实验

Answered已有结论Does a smaller vocabulary make repeated data hurt less?更小的词表能减轻重复数据的伤害吗?

Yes, when the data repeats many times: at 8 passes each halving of the vocabulary lowers the cost of repetition, and the truncated tokenizers end about level with the Marin tokenizer, which they trail clearly on fresh data. The byte vocabulary loses almost nothing to 8 passes. At 2.4 passes repetition costs every tokenizer little, and the vocabulary size makes no clear difference.

是,在数据重复很多次时成立:重复 8 遍时,词表每减半一次,重复带来的损失就更小,截短的 tokenizer 最终和 Marin tokenizer 大致持平,而在不重复的数据上它们明显落后。字节词表重复 8 遍几乎没有损失。重复 2.4 遍时,每种 tokenizer 受的损失都很小,词表大小没有明显影响。

Marin tokenizerMarin tokenizer Baseline基线

The MuonH Qwen3 recipe, unchanged: the softmax scores all 128K tokens.MuonH Qwen3 配方,不做改动:softmax 给全部 128K 个词打分。

Truncated tokenizer截短的 tokenizer

The Marin tokenizer cut to the first 64K, 32K, 16K or 8K tokens its BPE training learned: the softmax scores fewer tokens, and the same text becomes more tokens.把 Marin tokenizer 截到它的 BPE 训练最先学到的 64K、32K、16K 或 8K 个词:softmax 要打分的词更少,同样的文本被切成更多的词。

Testing检验中Does it hold at other model sizes?在其他规模的模型上也成立吗?

So far, yes. At 130m, 300m and 520m the 8K truncation loses less to 8 passes than the Marin tokenizer. The saving shrinks as the model grows, and at 520m the two tokenizers end level at 8 passes. The 520m byte pair is still queued.

目前看是。在 130m、300m 和 520m 上,8K 版本重复 8 遍的损失都比 Marin tokenizer 小。模型越大,省下的损失越少;在 520m 上,两种 tokenizer 重复 8 遍后最终持平。520m 的字节那一对还在排队。

Answered已有结论Is the trend real?这个趋势是真的吗?

Yes. With a second seed, each tokenizer's cost moves far less than the gap between the Marin tokenizer and the 8K truncation, and counting the bytes of each repeated part confirms that both tokenizers repeat the same amount of text.

是。换第二个种子后,每种 tokenizer 的代价变化都远小于 Marin tokenizer 与 8K 版本之间的差距;实际数出的每个重复部分的字节数也确认两种 tokenizer 重复的文本量相同。

Answered已有结论Do the input rows of rare tokens overfit?少见词的输入行会过拟合吗?

Yes, and they carry more than half of the gap to the 8K truncation. Without their own input rows, the rare tokens make the model worse on all data and better at 8 passes; that 8-pass model beats both the Marin tokenizer's and the 8K truncation's.

会,而且它们占了与 8K 版本之间差距的一半以上。去掉少见词自己的输入行后,模型在全部数据上变差,在重复 8 遍时变好;这个重复 8 遍的模型比 Marin tokenizer 和 8K 版本重复 8 遍的模型都好。

Answered已有结论Do the output rows of rare tokens overfit?少见词的输出行会过拟合吗?

Probably, though we could only show it by correlation. The cost of the Marin tokenizer is far highest where the target is a rare token, and the 8K truncation has no such peak. A causal test failed: a model whose output rows are built from the 8K pieces' rows, by their mean or by their sum, falls far behind the baseline on all of the data, so its cost of repetition says nothing about the rare rows alone.

很可能会,但我们只能用相关性说明。目标词是少见词时,Marin tokenizer 的代价远高于其他位置,而 8K 版本没有这样的峰。因果检验没有成功:输出行由 8K 片段行组成的模型,不论取平均还是求和,在全部数据上都远远落后于基线,所以它重复数据的代价不能单独说明少见词的输出行。

We tried the causal test of Q4 on the output side: the model still scores all 128K tokens, but each token's output row is built from the rows of its 8K pieces, so the rare tokens have no output rows of their own. With the mean of the pieces' rows, a token can never score above its best piece, so the model cannot rate ' walking' above ' walk'. With the sum, ' walking' scores ' walk' plus 'ing', which has no such cap. Both models train far worse than the baseline on all of the data: after 900 steps the sum is still about 0.6 nats per token behind, about as far as the mean. A composed input table costs little by comparison. An output head composed this way cannot fit the data, so we stopped the test there.

我们在输出端也试了 Q4 那样的因果检验:模型仍然给全部 128K 个词打分,但每个词的输出行由它的 8K 片段行组成,所以少见词没有自己的输出行。取片段行的平均时,一个词的分数不会超过它得分最高的片段,所以模型不能认为 ' walking' 比 ' walk' 更可能。求和时,' walking' 的分数等于 ' walk' 加 'ing',没有这个上限。两个模型在全部数据上都比基线差很多:训练 900 步后,求和版每个词仍然落后约 0.6 nats,和平均版差不多。相比之下,组合输入表的损失很小。这样组合出来的输出头拟合不了数据,所以我们停止了这个检验。

Answered已有结论Does a window of fewer bytes cost less?覆盖字节更少的上下文窗口,代价更小吗?

Not for the Marin tokenizer. A window that holds about the bytes of the 8K truncation's costs as much as the full window. For the byte vocabulary the short window explains part of its low cost: at 130m a window four times longer raises the byte cost, which stays far below the Marin tokenizer's.

对 Marin tokenizer 来说不会。一个装下的字节和 8K 版本窗口差不多的窗口,代价和完整窗口一样。对字节词表,短窗口能解释它代价低的一部分原因:在 130m 上,长四倍的窗口提高了字节词表的代价,但它仍远低于 Marin tokenizer 的代价。

Answered已有结论Do more optimizer steps cost less?更多的优化步数,代价更小吗?

No. With a third more steps over the same text, the Marin tokenizer's cost of 8 passes barely moves.

不会。在同样的文本上多走三分之一的步数,Marin tokenizer 重复 8 遍的代价几乎不变。

Testing检验中Does a smaller vocabulary still help when the model is overtrained?模型过度训练时,更小的词表还能减轻重复的伤害吗?

Takeaways核心结论

Acknowledgements致谢

The author is pleased to acknowledge that the work reported on in this post was substantially performed using the Princeton Research Computing resources at Princeton University. Princeton Research Computing is a consortium of groups including the Princeton Institute for Computational Science and Engineering (PICSciE) and Research Computing at Princeton University.

本文报告的工作主要使用普林斯顿大学 Princeton Research Computing 的计算资源完成。Princeton Research Computing 是由普林斯顿计算科学与工程研究所(PICSciE)和普林斯顿大学 Research Computing 等团队组成的联合体。

Experiment record实验记录

Citation

@misc{chai2026vocab,
  title        = {How Vocabulary Size Shapes Overfitting},
  author       = {Chai, Wenhao},
  year         = {2026},
  howpublished = {Blog post},
  url          = {https://wenhaochai.com/blogs/vocab-overfitting.html}
}
Experiments Back to top回到顶部