Attention-Informed Mapping (AIM) for cross-lingual tokenizer adaptation
Fertility measures the average number of tokens produced per word. High-resource languages like English achieve near 1:1 ratios, while underserved languages like Georgian fragment into 6-8 tokens per word, consuming context budget and degrading performance.
AIM uses teacher-student distillation on attention distributions, not logits or hidden states. The teacher model runs on the original tokenizer, recording inter-token attention patterns. The student model with the new tokenizer learns to reproduce these distributions.
Interactive visualization showing how attention distributions are transferred from teacher to student.
MATT (AIM warm-up + continual pretraining) compared against random initialization, WECHSEL, and FOCUS across six languages.
| Language | Script | Fertility | Recovery |
|---|---|---|---|
| English | Latin | 1.2 | 94% |
| German | Latin | 1.8 | 89% |
| Japanese | Mixed | 3.4 | 78% |
| Arabic | Arabic | 4.1 | 74% |
| Swahili | Latin | 5.2 | 68% |
| Ukrainian | Cyrillic | 6.7 | 63% |
Episode 37 examined Cartridges' SCI (Sampled Chunk Initialization), which independently arrived at the same core insight: random initialization fails because the model's own learned activations are the right teacher signal.
Evolution from semantic matching (WECHSEL, FOCUS) to attention-aware initialization (MATT).
| Method | Year | Approach | Initialization Target | Model Access |
|---|---|---|---|---|
| WECHSEL | 2022 | FastText + Bilingual Dict | Embeddings only | None (external) |
| FOCUS | 2023 | FastText on new vocab | Embeddings only | None (external) |
| Tik-to-Tok | 2023 | Freeze attention, adapt FFN | Embeddings + FFN | Full model |
| ZETT | 2024 | Zero-shot LM objective | Full model | Full model |
| MATT (AIM) | 2025 | Attention distillation | Attention layers | Attention only |