This companion page visualizes the paper’s central trade: replace an ever-growing Transformer KV cache with a fixed bank of memory slots, reconstruct each incoming token from that bank, and write back only the residual that still looks new. The page is about geometry, budgets, and system tradeoffs, not prose recap.
Standard attention stores an addressable trace of every past token. Lattice pays approximation instead of cache growth by forcing the history into a fixed slot bank.
Step through four token situations. The update is small when the current slots already span the token, and larger when a genuinely new direction survives reconstruction.
Exact Lattice values stated in the transcript anchor the loss view. The surrounding comparison numbers and matched-budget curves are stylized mock data for visual intuition.
Different operators care about different frontiers. A serve-heavy stack may prefer cache engineering; an edge or streaming stack may prefer fixed-state memory even with softer exact recall.