AI Post Transformers — Episode Companion Visualization

CXL Computational Memory Offloading for Lower Runtime

Partial offloading to CXL-based computational memory only helps if host and device stop waiting on each other. This companion breaks down Remote Polling, Bulk Synchronous Flow, and the paper's asynchronous back-streaming protocol, KAI — visually.

Suyeon Lee, Kangkyu Park, Kwangsik Shin, Ada Gavrilovska — Georgia Tech & SK hynix, Dec 2025 (pre-publication)
arXiv:2512.04449 Systems / Near-Memory Computing CXL.io vs CXL.mem Host ↔ CCM Coordination

Host ↔ Fabric ↔ CCM

Partial offloading keeps the application on the host CPU, but routes selected memory-intensive slices to near-memory compute on the CCM. Hover the boxes and paths.
Hover a box or path above for details.
Host domain CCM domain Partial offload path

Why "cheap launch" isn't enough: the roofline view

Graph analytics, filtering, vector distance, and KV-cache reads are low arithmetic-intensity, bandwidth-bound workloads — they hit the fabric's roofline long before compute saturates. Hover the points.
Hover a workload point above for details.

Who waits on whom

Same timeline, three coordination strategies. Green = busy/useful work, orange = chatty poll wait, dim-red = blocked stall, purple = bulk barrier transfer.
Busy / useful work Chatty poll wait Blocked stall Bulk barrier transfer Idle (nothing to do yet)

Moving the polling point local

KAI splits the host DMA region into a metadata ring and a payload ring. The CCM writes both via DMA; the host polls its own local metadata ring instead of pinging the CCM. Hover the rings and arrows.
Hover a ring cell or arrow above for details.

Out-of-order streaming, in-order consumption

The CCM sends whichever chunk finishes first — arrival order is scrambled. The host-side interface reconstructs logical order without forcing either scheduler to lock-step. Hover a chunk.
Hover a numbered chunk above to trace its path.

End-to-end runtime, normalized to RP

KAI's headline number — up to 50.4% runtime reduction — shows up on Vector Distance (ANN). Gains vary by how much a workload's output streams in useful chunks.
Remote Polling (baseline = 1.0) Bulk Synchronous Flow KAI

Idle-time reduction (×)

Average CCM idle-time reduction is 22.11× vs RP and host idle-time reduction is 3.85× vs RP, with similar CCM-side gains against BS. Overlap collapses idle regions on both sides — until the fabric itself saturates.
Hover a cell above for the exact figure.
Overlap helps when there's latent parallel slack — chunked, streamable output and fabric headroom. Once the link saturates, no protocol recovers bandwidth that isn't there.

References