Tokenization Bias: The Hidden Flaw Breaking Language Models
Papers:
2410.09303
,
2503.20083
arXiv:2410.09303
arXiv:2503.20083
Pipeline Overview
Byte-Token Lemma
FIM Performance
Cross-Tokenizer Ensembles
Tokenization Bias Pipeline Overview
Tokenized Model
Vocabulary: BPE / SentencePiece
Outputs token probabilities
Byte-level Model
Vocabulary: Raw bytes (256)
Outputs byte probabilities
Input Prompt
"def foo(aa"
Mid-token boundary (FIM)
Token Probabilities
P('and')=0.4, P('aa')=0.0, ...
Byte Probabilities
P('a')=0.25, P('b')=0.1, ...
Byte-Token Representation Lemma
Marginalize token probs → byte probs
Sum token probs by leading byte
Byte-Token Representation Lemma
Tokens grouped by leading byte
Next Byte Probabilities
Hover tokens to see their probability and leading byte grouping. Toggle token probability distribution below.
Show Uniform Token Distribution
FIM Coding Benchmark Accuracy
0%
20%
40%
60%
Baseline
Token Healing
Byte-level Correction
Toggle: Absolute / Relative Accuracy
Show Relative Improvement
Cross-Tokenizer Ensemble Agreement Heatmap
Low Agreement
High Agreement
Model A
Model B
Ensemble
References
2410.09303 - Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model Ensembles
2503.20083 - Cross-Tokenizer Knowledge Distillation
Neural Machine Translation of Rare Words with Subword Units — Sennrich et al., 2016
SentencePiece — Kudo & Richardson, 2018
Efficient Training of Language Models to Fill in the Middle — Bavarian et al., 2022
ByT5 — Xue et al., 2022
MegaByte — Yu et al., 2023
Podcast Website