Model-Aware Tokenizer Transfer for Multilingual LLMs

Attention-Informed Mapping (AIM) for cross-lingual tokenizer adaptation

AGH University of Krakow October 2025 Authors: Haltiuk & Smywińśki-Pohl
arXiv:2510.21954

Tokenizer Fertility: The Core Problem

Fertility measures the average number of tokens produced per word. High-resource languages like English achieve near 1:1 ratios, while underserved languages like Georgian fragment into 6-8 tokens per word, consuming context budget and degrading performance.

Low Fertility (1-2 tokens/word)
Medium Fertility (3-4 tokens/word)
High Fertility (6-8 tokens/word)

Impact on Context Budget

Attention-Informed Mapping (AIM)

AIM uses teacher-student distillation on attention distributions, not logits or hidden states. The teacher model runs on the original tokenizer, recording inter-token attention patterns. The student model with the new tokenizer learns to reproduce these distributions.

Attention Distillation Heatmap

Interactive visualization showing how attention distributions are transferred from teacher to student.

Performance Comparison: MATT vs Baselines

MATT (AIM warm-up + continual pretraining) compared against random initialization, WECHSEL, and FOCUS across six languages.

Language-Specific Breakdown

Language Script Fertility Recovery
English Latin 1.2 94%
German Latin 1.8 89%
Japanese Mixed 3.4 78%
Arabic Arabic 4.1 74%
Swahili Latin 5.2 68%
Ukrainian Cyrillic 6.7 63%

Connection to Cartridges Architecture

Episode 37 examined Cartridges' SCI (Sampled Chunk Initialization), which independently arrived at the same core insight: random initialization fails because the model's own learned activations are the right teacher signal.

Parallel Insights: MATT vs Cartridges

Cross-Tokenizer Cartridge Sharing Strategies

Tokenizer Transfer Methods Timeline

Evolution from semantic matching (WECHSEL, FOCUS) to attention-aware initialization (MATT).

Method Feature Comparison

Method Year Approach Initialization Target Model Access
WECHSEL 2022 FastText + Bilingual Dict Embeddings only None (external)
FOCUS 2023 FastText on new vocab Embeddings only None (external)
Tik-to-Tok 2023 Freeze attention, adapt FFN Embeddings + FFN Full model
ZETT 2024 Zero-shot LM objective Full model Full model
MATT (AIM) 2025 Attention distillation Attention layers Attention only

Knowledge Localization: Where Does LLM Knowledge Live?

Key References