AI Post Transformers — Episode Companion Visualization
CXL Computational Memory Offloading for Lower Runtime
Partial offloading to CXL-based computational memory only helps if host and
device stop waiting on each other. This companion breaks down Remote Polling,
Bulk Synchronous Flow, and the paper's asynchronous back-streaming protocol,
KAI — visually.
Suyeon Lee, Kangkyu Park, Kwangsik Shin, Ada Gavrilovska — Georgia Tech & SK hynix, Dec 2025 (pre-publication)
arXiv:2512.04449Systems / Near-Memory ComputingCXL.io vs CXL.memHost ↔ CCM Coordination
Host ↔ Fabric ↔ CCM
Partial offloading keeps the application on the host CPU, but routes selected memory-intensive slices to near-memory compute on the CCM. Hover the boxes and paths.
Hover a box or path above for details.
Host domainCCM domainPartial offload path
Why "cheap launch" isn't enough: the roofline view
Graph analytics, filtering, vector distance, and KV-cache reads are low arithmetic-intensity, bandwidth-bound workloads — they hit the fabric's roofline long before compute saturates. Hover the points.
Hover a workload point above for details.
Who waits on whom
Same timeline, three coordination strategies. Green = busy/useful work, orange = chatty poll wait, dim-red = blocked stall, purple = bulk barrier transfer.
Busy / useful workChatty poll waitBlocked stallBulk barrier transferIdle (nothing to do yet)
Moving the polling point local
KAI splits the host DMA region into a metadata ring and a payload ring. The CCM writes both via DMA; the host polls its own local metadata ring instead of pinging the CCM. Hover the rings and arrows.
Hover a ring cell or arrow above for details.
Out-of-order streaming, in-order consumption
The CCM sends whichever chunk finishes first — arrival order is scrambled. The host-side interface reconstructs logical order without forcing either scheduler to lock-step. Hover a chunk.
Hover a numbered chunk above to trace its path.
End-to-end runtime, normalized to RP
KAI's headline number — up to 50.4% runtime reduction — shows up on Vector Distance (ANN). Gains vary by how much a workload's output streams in useful chunks.
Average CCM idle-time reduction is 22.11× vs RP and host idle-time reduction is 3.85× vs RP, with similar CCM-side gains against BS. Overlap collapses idle regions on both sides — until the fabric itself saturates.
Hover a cell above for the exact figure.
Overlap helps when there's latent parallel slack — chunked, streamable output and fabric headroom. Once the link saturates, no protocol recovers bandwidth that isn't there.