arXiv:2512.23852 COLM 2025 arXiv v1 · Dec 29 2025 Transcript IDs: pending

TRELLIS and Bounded-Memory Transformer KV Compression

This companion page treats TRELLIS as a serving-architecture idea: long-context pain comes from the KV cache growing linearly at inference, and TRELLIS answers with a fixed slot bank that is rewritten online by a tiny test-time learner.

State budget
Fixed memory slots
Context can grow while persistent memory stays capped.
Update rule
Decay + gradient rewrite
Memory behaves like fast weights instead of a token diary.
Reported scale
125M → 1B
From-scratch runs across four model sizes.
Missing table
TTFT / decode / throughput
The episode stresses that serving-system tradeoffs remain open.

Memory Growth at Serving Time

Mock deployment geometry showing why the episode frames KV growth, not only quadratic training attention, as the live bottleneck.

append-only KV cache
prior bounded-slot family
TRELLIS fixed state

Append-Only Diary vs Rewritten Notebook

The same token stream arrives in both cases. The only difference is the state geometry: growing exact token rows on the left, or a small bank of rewritten slots on the right.

Serving Path Comparison

Raw KV preserves exact token identity. TRELLIS trades that for bounded memory and an active write rule.

Lineage Map

The episode places TRELLIS in a long arc from segment recurrence to explicit bounded-memory control.

Memory vs Exactness

Bounded-memory methods usually move up and left: lower state growth, weaker exact per-token reuse.

TRELLIS as a Tiny Online Learner

Step through the local update. The slow network stays fixed; only memory moves at test time.

Fast-Weight Regression Loop

The two-pass story from the episode becomes concrete here: one compact address space for keys, one compact store for values.

Visual equation: memory is updated by a decayed gradient step on reconstruction loss, not by appending another exact K/V row.

What Forgetting Looks Like

The episode’s real claim is selective survival. These mock matrices show how repeated facts and distractors compete for the same fixed slot bank.

Slot Feature Energy

Rows are memory slots. Columns are latent channels after repeated writes.

Write Intensity Over Time

Rows are tokens. Columns are slots receiving the strongest local updates.

Retention Curves

Decay should let recurring facts persist while stale distractors fade faster.

Episode-Faithful Results Sketch

These charts are realistic mock data. They visualize the narrative in the transcript rather than attempt a digitized copy of the paper’s figures.

Context-Length Sweep

The orange full-attention comparator intentionally fades after about 8k because the episode calls out that limitation in the from-scratch comparison.

Long-Retrieval Margin

The episode reports about a four-point average gain and about six points at longer contexts.

Scale and Training Budget

Four model sizes, roughly 2.4B to 30B tokens, and evaluation spanning language modeling, recall, reasoning, and time series.

The Missing Deployment Table

The episode’s caution is simple: bounded memory can help quality scaling, but exact-cache reuse systems answer different deployment questions.

Method Map

Higher means stronger bounded-memory behavior. Farther right means stronger exact token reuse and inspectability.

Which Questions Are Answered?

Hotter cells mean the method or paper directly covers that axis. Cooler cells mean the answer is weak or absent.

References

Compact paper links for the bounded-memory lineage, the test-time-learning frame, and the adjacent KV systems alternatives around the episode.

Related AI Post Transformers Episodes

Public callbacks on test-time memory, context compression, and long-sequence alternatives.