Every recent open-weight model loses 11–39% relative accuracy on facts from 2023–2024
Same model, same weights — just facts from closer to the training cutoff score worse than facts from years earlier. Kyutai's sequential-order model breaks that pattern.
KairosQA accuracy by fact year, per model
Relative accuracy loss: 2023–24 vs 2020–21
Old facts vs. recent facts, head to head
One design choice: shuffle everything, or feed it in order
Standard pretraining pools every Common Crawl snapshot, dedupes and quality-filters each one, then globally shuffles before minibatching — erasing any timestamp signal. Kyutai's sequential model changes only the order.
dactory filtering pipeline
Quality weights by domain
Why order matters: learning-rate decay makes late data sticky
Updates late in training, when the learning rate is low, reshape the weights far more durably than updates from early on. Putting recent facts last exploits the same mechanism that normally causes catastrophic forgetting — pointed at retention instead.
Learning-rate schedule over a training run
Branching cooldown: 8 checkpoints from one run
Is KairosQA just an easy, gameable benchmark?
Two stress tests the hosts pushed on: does accuracy survive harder multiple-choice, and does it track genuine subject familiarity rather than memorized distractor patterns.
Accuracy vs. number of cloze choices
F1 by subject popularity
Not a free lunch — three cracks worth sitting with
A data-quality confound in the headline gap, a forgetting cost on older facts, and a result that mostly evaporates on an external benchmark the authors didn't build.