AI Post Transformers — Episode Companion

Move the Query, Not the Cache: MLA Rewrites GPU Fabric Attention Routing

arXiv:2606.01502 Ma, Eitzinger, Köstler, Wellein — NHR@FAU, 2026 Cross-Instance MLA Attention Routing

References