AI Post Transformers — Episode Companion
Move the Query, Not the Cache: MLA Rewrites GPU Fabric Attention Routing
arXiv:2606.01502
Ma, Eitzinger, Köstler, Wellein — NHR@FAU, 2026
Cross-Instance MLA Attention Routing
References