Append-Only Diary vs Rewritten Notebook
The same token stream arrives in both cases. The only difference is the state geometry: growing exact token rows on the left, or a small bank of rewritten slots on the right.
This companion page treats TRELLIS as a serving-architecture idea: long-context pain comes from the KV cache growing linearly at inference, and TRELLIS answers with a fixed slot bank that is rewritten online by a tiny test-time learner.
Mock deployment geometry showing why the episode frames KV growth, not only quadratic training attention, as the live bottleneck.
The same token stream arrives in both cases. The only difference is the state geometry: growing exact token rows on the left, or a small bank of rewritten slots on the right.
Raw KV preserves exact token identity. TRELLIS trades that for bounded memory and an active write rule.
The episode places TRELLIS in a long arc from segment recurrence to explicit bounded-memory control.
Bounded-memory methods usually move up and left: lower state growth, weaker exact per-token reuse.
Step through the local update. The slow network stays fixed; only memory moves at test time.
The two-pass story from the episode becomes concrete here: one compact address space for keys, one compact store for values.
Visual equation: memory is updated by a decayed gradient step on reconstruction loss, not by appending another exact K/V row.
The episode’s real claim is selective survival. These mock matrices show how repeated facts and distractors compete for the same fixed slot bank.
Rows are memory slots. Columns are latent channels after repeated writes.
Rows are tokens. Columns are slots receiving the strongest local updates.
Decay should let recurring facts persist while stale distractors fade faster.
These charts are realistic mock data. They visualize the narrative in the transcript rather than attempt a digitized copy of the paper’s figures.
The orange full-attention comparator intentionally fades after about 8k because the episode calls out that limitation in the from-scratch comparison.
The episode reports about a four-point average gain and about six points at longer contexts.
Four model sizes, roughly 2.4B to 30B tokens, and evaluation spanning language modeling, recall, reasoning, and time series.
The episode’s caution is simple: bounded memory can help quality scaling, but exact-cache reuse systems answer different deployment questions.
Higher means stronger bounded-memory behavior. Farther right means stronger exact token reuse and inspectability.
Hotter cells mean the method or paper directly covers that axis. Cooler cells mean the answer is weak or absent.
Compact paper links for the bounded-memory lineage, the test-time-learning frame, and the adjacent KV systems alternatives around the episode.
Public callbacks on test-time memory, context compression, and long-sequence alternatives.