Hal Turing and Dr. Ada Shannon open the episode by confronting a structural flaw that has been hiding in plain sight since the transformer era began: tokenization bias. The episode centers on "Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model Ensembles" by Buu Phan, Brandon Amos, Itai Gat, Marton Havasi, Matthew Muckley, and Karen Ullrich (ICLR 2025), which formally proves that a tokenized model and a byte-level model can be statistically equivalent and still produce wildly different predictions for the same next character. The hosts trace the origins of the problem through BPE's introduction by Rico Sennrich, Barry Haddow, and Alexandra Birch in 2016 and its industrialization via Kudo and Richardson's SentencePiece in 2018 — a library now frozen into the spine of LLaMA, Mistral, Gemma, and most open-source models not in the OpenAI lineage. The discussion sharpens around fill-in-the-middle prompting, the paradigm introduced by Mohammad Bavarian and colleagues at OpenAI in 2022 and now embedded in every major code completion tool from GitHub Copilot to StarCoder. Shannon walks through the paper's central example: a code completion scenario where the correct next character receives a probability of exactly zero — not a rounding artifact but a structural impossibility, because the tokenizer has carved up the prompt in a way that makes the right answer unreachable in token-space. Turing challenges the framing, arguing that byte-level alternatives like ByT5 and MegaByte existed and BPE was an informed trade-off against the three-to-eight times sequence length penalty that raw bytes impose on attention compute. Shannon holds the line: the point is not that BPE was a mistake but that its systematic bias was never formally characterized until now, and the Byte-Token Representation Lemma finally gives the field the mathematical language to name and measure it. The episode closes by introducing the second paper from the episode's pairing — Minixhofer et al.'s NeurIPS 2025 work on cross-tokenizer knowledge distillation — which attacks the tokenizer barrier from the training side rather than the inference side. Where Phan et al. offer a zero-shot correction algorithm that recovers 18% on fill-in-the-middle coding benchmarks without any retraining, Minixhofer et al. enable knowledge transfer between models with fundamentally incompatible vocabularies, breaking the assumption that distillation requires shared tokenization. Together the two papers sketch a trajectory where tokenization becomes a transparent implementation detail rather than an architectural constraint that determines what a model can and cannot express.

Sources:
1. Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model Ensembles — Buu Phan, Brandon Amos, Itai Gat, Marton Havasi, Matthew Muckley, Karen Ullrich, 2024
http://arxiv.org/abs/2410.09303
2. https://arxiv.org/pdf/2503.20083
3. Neural Machine Translation of Rare Words with Subword Units — Rico Sennrich, Barry Haddow, Alexandra Birch, 2016
https://scholar.google.com/scholar?q=Neural+Machine+Translation+of+Rare+Words+with+Subword+Units
4. SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing — Taku Kudo, John Richardson, 2018
https://scholar.google.com/scholar?q=SentencePiece%3A+A+simple+and+language+independent+subword+tokenizer+and+detokenizer+for+Neural+Text+Processing
5. Toward a Theory of Tokenization in LLMs — Nived Rajaraman, Jiantao Jiao, Kannan Ramchandran, 2024
https://scholar.google.com/scholar?q=Toward+a+Theory+of+Tokenization+in+LLMs
6. How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models — Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, Iryna Gurevych, 2021
https://scholar.google.com/scholar?q=How+Good+is+Your+Tokenizer%3F+On+the+Monolingual+Performance+of+Multilingual+Language+Models
7. ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models — Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, Colin Raffel, 2022
https://scholar.google.com/scholar?q=ByT5%3A+Towards+a+Token-Free+Future+with+Pre-trained+Byte-to-Byte+Models
8. MEGABYTE: Predicting Million-byte Sequences with Multiscale Transformers — Lili Yu, Daniel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, Mike Lewis, 2023
https://scholar.google.com/scholar?q=MEGABYTE%3A+Predicting+Million-byte+Sequences+with+Multiscale+Transformers
9. CANINE: Pre-training an Efficient Tokenization-Free Encoder for Language Representation — Jonathan H. Clark, Dan Garrette, Iulia Turc, John Wieting, 2022
https://scholar.google.com/scholar?q=CANINE%3A+Pre-training+an+Efficient+Tokenization-Free+Encoder+for+Language+Representation
10. Charformer: Fast Character Transformers via Gradient-Based Subword Tokenization — Yi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Gupta, Hyung Won Chung, Dara Bahri, Zhen Qin, Simon Kornblith, Cecelia Zhang, Donald Metzler, Mostafa Dehghani, 2022
https://scholar.google.com/scholar?q=Charformer%3A+Fast+Character+Transformers+via+Gradient-Based+Subword+Tokenization
11. Efficient Training of Language Models to Fill in the Middle — Mohammad Bavarian, Heewoo Jun, Nikolas Tezak, John Schulman, Christine McLeavey, Jerry Tworek, Mark Chen, 2022
https://scholar.google.com/scholar?q=Efficient+Training+of+Language+Models+to+Fill+in+the+Middle
12. Evaluating Large Language Models Trained on Code — Mark Chen, Jerry Tworek, Heewoo Jun, and a large team at OpenAI, 2021
https://scholar.google.com/scholar?q=Evaluating+Large+Language+Models+Trained+on+Code
13. Code Llama: Open Foundation Models for Code — Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, and others at Meta AI, 2023
https://scholar.google.com/scholar?q=Code+Llama%3A+Open+Foundation+Models+for+Code
14. StarCoder: May the Source Be With You! — Raymond Li, Loubna Ben Allal, and a large BigCode / HuggingFace / ServiceNow collaboration, 2023
https://scholar.google.com/scholar?q=StarCoder%3A+May+the+Source+Be+With+You%21
15. Training Products of Experts by Minimizing Contrastive Divergence — Geoffrey E. Hinton, 2002
https://scholar.google.com/scholar?q=Training+Products+of+Experts+by+Minimizing+Contrastive+Divergence
16. Self-Consistency Improves Chain of Thought Reasoning in Language Models — Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, Denny Zhou, 2023
https://scholar.google.com/scholar?q=Self-Consistency+Improves+Chain+of+Thought+Reasoning+in+Language+Models
17. Knowledge Fusion of Large Language Models — Fanqi Wan, Xinting Huang, Deng Cai, Xiaojun Quan, Wei Bi, Shuming Shi, 2024
https://scholar.google.com/scholar?q=Knowledge+Fusion+of+Large+Language+Models
18. LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion — Dongfu Jiang, Xiang Ren, Bill Yuchen Lin, 2023
https://scholar.google.com/scholar?q=LLM-Blender%3A+Ensembling+Large+Language+Models+with+Pairwise+Ranking+and+Generative+Fusion
19. Fill in the Middle: A New Training Objective for Language Models — Bavarian et al., 2022
https://scholar.google.com/scholar?q=Fill+in+the+Middle%3A+A+New+Training+Objective+for+Language+Models
20. CodeFusion: A Pre-trained Diffusion Model for Code Generation — Dagan et al., 2024
https://scholar.google.com/scholar?q=CodeFusion%3A+A+Pre-trained+Diffusion+Model+for+Code+Generation
21. Is Tokenization More Than Compression? — Rajaraman et al., 2024
https://scholar.google.com/scholar?q=Is+Tokenization+More+Than+Compression%3F
22. SpaceByte: Towards Deleting Tokenization from Large Language Modeling — Slagle, 2024
https://scholar.google.com/scholar?q=SpaceByte%3A+Towards+Deleting+Tokenization+from+Large+Language+Modeling
23. StarCoder 2 and The Stack v2: The Next Generation — Lozhkov et al., 2024
https://scholar.google.com/scholar?q=StarCoder+2+and+The+Stack+v2%3A+The+Next+Generation
24. Byte Latent Transformer: Patches Scale Better Than Tokens — Pagnoni et al. (Meta AI), 2024
https://scholar.google.com/scholar?q=Byte+Latent+Transformer%3A+Patches+Scale+Better+Than+Tokens
25. Token-level Ensembling of Models with Different Vocabularies — various, 2024
https://scholar.google.com/scholar?q=Token-level+Ensembling+of+Models+with+Different+Vocabularies
26. Bridging the Gap Between Different Vocabularies for LLM Ensemble — various, 2024
https://scholar.google.com/scholar?q=Bridging+the+Gap+Between+Different+Vocabularies+for+LLM+Ensemble
27. Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies — various, 2024
https://scholar.google.com/scholar?q=Accelerating+LLM+Inference+with+Lossless+Speculative+Decoding+Algorithms+for+Heterogeneous+Vocabularies
28. Bridging Developer Instructions and Code Completion Through Instruction-Aware Fill-in-the-Middle Paradigm — various, 2024
https://scholar.google.com/scholar?q=Bridging+Developer+Instructions+and+Code+Completion+Through+Instruction-Aware+Fill-in-the-Middle+Paradigm
29. AI Post Transformers: Fast Inference from Transformers via Speculative Decoding — Hal Turing & Dr. Ada Shannon, Sat,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Fast-Inference-from-Transformers-via-Speculative-Decoding-e3foclv
30. AI Post Transformers: Multiagent Debate Improves Language Model Reasoning — Hal Turing & Dr. Ada Shannon, Fri,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Multiagent-Debate-Improves-Language-Model-Reasoning-e36kfd4
Interactive Visualization: Tokenization Bias: The Hidden Flaw Breaking Language Models

Hal Turing and Dr. Ada Shannon return to the CARTRIDGE compression system with a mechanistic lens, covering Maurizio A. Diaz's paper "Learned Structure in Cartridges: Keys as Shareable Routers in Self-Studied Representations" (arXiv 2508.17032), presented at the NeurIPS 2025 Workshop on Mechanistic Interpretability. Building on the original CARTRIDGE episode from November 10th, 2025 and the follow-up from February 6th, 2026, this episode asks the question those earlier discussions left open: what structure does the optimizer actually induce in a trained CARTRIDGE? The hosts ground the discussion in the memory scaling problem driving the entire field—KV caches that grow linearly with context length, now routinely dwarfing model weights at the 128K-to-million-token scales of current frontier models—and trace how techniques like PagedAttention, Grouped Query Attention, and token eviction address symptoms without shrinking the underlying representation. Diaz's central finding is a clean functional division between key and value vectors inside a trained CARTRIDGE. Keys converge to stable retrieval routers: low-rank, consistent structures that steer attention toward the right stored content across diverse queries. Values carry the compressed semantic payload. The hosts connect this directly to how CARTRIDGE's Self-Study training pipeline works—because the cache is optimized against synthetic question-answer traces generated by the model over its own content, the training signal explicitly selects for routing behavior, making the key-as-router outcome a predictable consequence of the objective rather than an accident. Diaz uses Singular Value Decomposition to quantify this structure layer by layer, separating the geometric properties of key matrices from value matrices across training checkpoints. Two downstream findings from the key-router property shape the second half of the discussion. Because keys are stable and low-rank, they transfer across tasks with minimal degradation—a result with direct implications for multi-task serving, where a single shared key structure could route to task-specific value sets without independent CARTRIDGE training per deployment. The Sampled Chunk Initialization method introduced in the paper exploits this stability to warm-start CARTRIDGE training, accelerating convergence by initializing the learnable KV pairs from a small representative sample rather than random weights. Hal and Ada close by discussing what the key-as-router framing implies for KV-cache compression research more broadly: if the routing function is separable and transferable, compression schemes that conflate keys and values may be discarding structure that has real serving-efficiency value.

Sources:
1. Learned Structure in Cartridges: Keys as Shareable Routers in Self-Studied Representations — Maurizio Diaz, 2025
http://arxiv.org/abs/2508.17032
2. CARTRIDGES: Learning to Pack Long Contexts into KV Caches — Zhihao Zhang, Aditya Desai, Amir Gholami, Michael W. Mahoney, Kurt Keutzer, et al. (Berkeley / ICSI), 2025
https://scholar.google.com/scholar?q=CARTRIDGES%3A+Learning+to+Pack+Long+Contexts+into+KV+Caches
3. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianghao Huang, Mu Li, Beidi Chen, Jason D. Lee, Binhang Yuan, Ce Zhang, Cho-Jui Hsieh, 2023
https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models
4. Scissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test Time — Zichang Liu, Aditya Desai, Fangshuo Liao, Weitao Wang, Victor Xie, Zhaozhuo Xu, Anastasios Kyrillidis, Anshumali Shrivastava, 2023
https://scholar.google.com/scholar?q=Scissorhands%3A+Exploiting+the+Persistence+of+Importance+Hypothesis+for+LLM+KV+Cache+Compression+at+Test+Time
5. Efficient Streaming Language Models with Attention Sinks — Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, Mike Lewis, 2023
https://scholar.google.com/scholar?q=Efficient+Streaming+Language+Models+with+Attention+Sinks
6. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
7. Lost in the Middle: How Language Models Use Long Contexts — Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, Percy Liang, 2023
https://scholar.google.com/scholar?q=Lost+in+the+Middle%3A+How+Language+Models+Use+Long+Contexts
8. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebron, Sumit Sanghai, 2023
https://scholar.google.com/scholar?q=GQA%3A+Training+Generalized+Multi-Query+Transformer+Models+from+Multi-Head+Checkpoints
9. Extending Context Window of Large Language Models via Positional Interpolation — Shouyuan Chen, Sherman Wong, Liangjian Chen, Yuandong Tian, 2023
https://scholar.google.com/scholar?q=Extending+Context+Window+of+Large+Language+Models+via+Positional+Interpolation
10. Constitutional AI: Harmlessness from AI Feedback — Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. (Anthropic), 2022
https://scholar.google.com/scholar?q=Constitutional+AI%3A+Harmlessness+from+AI+Feedback
11. Distilling the Knowledge in a Neural Network — Geoffrey Hinton, Oriol Vinyals, Jeff Dean, 2015
https://scholar.google.com/scholar?q=Distilling+the+Knowledge+in+a+Neural+Network
12. Compressing Context to Enhance Inference Efficiency of Large Language Models — Yucheng Li, Bo Dong, Chenghua Lin, Frank Guerin, 2023
https://scholar.google.com/scholar?q=Compressing+Context+to+Enhance+Inference+Efficiency+of+Large+Language+Models
13. AutoCompressors: Adapting Language Models to Summarize Arbitrary Contexts into Summary Vectors — Alexis Chevalier, Alexander Wettig, Anirudh Anand, Danqi Chen, 2023
https://scholar.google.com/scholar?q=AutoCompressors%3A+Adapting+Language+Models+to+Summarize+Arbitrary+Contexts+into+Summary+Vectors
14. A Mathematical Framework for Transformer Circuits — Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, Chris Olah, 2021
https://scholar.google.com/scholar?q=A+Mathematical+Framework+for+Transformer+Circuits
15. Toy Models of Superposition — Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Ben Mann, Shan Carter, Chris Olah, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Jared Kaplan, Dario Amodei, 2022
https://scholar.google.com/scholar?q=Toy+Models+of+Superposition
16. Towards Monosemanticity: Decomposing Language Models with Dictionary Learning — Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, William Lamont, Annabel Shearer, Zac Hatfield-Dodds, Tom Henighan, Nicholas Joseph, Ben Ziegler, Bilal Chaudhry, Will Hao, Timothy Telleen-Lawton, Daniel Mossing, Adam Jermyn, Tom Brown, Chris Olah, Sam McCandlish, Jack Clark, Dario Amodei, 2023
https://scholar.google.com/scholar?q=Towards+Monosemanticity%3A+Decomposing+Language+Models+with+Dictionary+Learning
17. In-context Learning and Induction Heads — Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, Chris Olah, 2022
https://scholar.google.com/scholar?q=In-context+Learning+and+Induction+Heads
18. AutoCompressors: Compressing Contexts with Language Models — Chevalier et al., 2023
https://scholar.google.com/scholar?q=AutoCompressors%3A+Compressing+Contexts+with+Language+Models
19. Gist Tokens: Compressing Prompts into Tokens for Long-Context Language Models — Mu et al., 2023
https://scholar.google.com/scholar?q=Gist+Tokens%3A+Compressing+Prompts+into+Tokens+for+Long-Context+Language+Models
20. The Power of Scale for Parameter-Efficient Prompt Tuning — Lester et al., 2021
https://scholar.google.com/scholar?q=The+Power+of+Scale+for+Parameter-Efficient+Prompt+Tuning
21. Prefix-Tuning: Optimizing Continuous Prompts for Generation — Li and Liang, 2021
https://scholar.google.com/scholar?q=Prefix-Tuning%3A+Optimizing+Continuous+Prompts+for+Generation
22. PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling — Cai et al., 2024
https://scholar.google.com/scholar?q=PyramidKV%3A+Dynamic+KV+Cache+Compression+based+on+Pyramidal+Information+Funneling
23. Tokasaurus: A High-Throughput LLM Serving Engine with Grouped-Sparse Attention — Lenz et al., 2025
https://scholar.google.com/scholar?q=Tokasaurus%3A+A+High-Throughput+LLM+Serving+Engine+with+Grouped-Sparse+Attention
24. Unlocking the Address Book: Dissecting the Sparse Semantic Structure of LLM Key-Value Caches via Sparse Autoencoders — unknown from snippet, 2024-2025
https://scholar.google.com/scholar?q=Unlocking+the+Address+Book%3A+Dissecting+the+Sparse+Semantic+Structure+of+LLM+Key-Value+Caches+via+Sparse+Autoencoders
25. SCBench: A KV Cache-Centric Analysis of Long-Context Methods — approximate, multiple authors, 2024-2025
https://scholar.google.com/scholar?q=SCBench%3A+A+KV+Cache-Centric+Analysis+of+Long-Context+Methods
26. LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression — approximate, Microsoft Research, 2024
https://scholar.google.com/scholar?q=LLMLingua-2%3A+Data+Distillation+for+Efficient+and+Faithful+Task-Agnostic+Prompt+Compression
27. AI Post Transformers: CARTRIDGE: Efficient In-Context Learning via Distillation — Hal Turing & Dr. Ada Shannon, Mon,
https://podcasters.spotify.com/pod/show/12146088098/episodes/CARTRIDGE-Efficient-In-Context-Learning-via-Distillation-e3aous4
28. AI Post Transformers: Context Distillation for Language Models — Hal Turing & Dr. Ada Shannon, Mon,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Context-Distillation-for-Language-Models-e3aouen
29. AI Post Transformers: Advancements in Efficient KV Cache Quantization and Management — Hal Turing & Dr. Ada Shannon, Thu,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Advancements-in-Efficient-KV-Cache-Quantization-and-Management-e3fk9kr
30. AI Post Transformers: Architectural Migration to Multi-head Latent Attention — Hal Turing & Dr. Ada Shannon, Wed,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Architectural-Migration-to-Multi-head-Latent-Attention-e39jbmq

Hal Turing and Dr. Ada Shannon dig into FlashAttention-4, a March 2026 paper from a cross-institutional team including Tri Dao, Jay Shah, and colleagues at Princeton, Meta, NVIDIA, Colfax Research, Georgia Tech, and Together AI. The paper targets a precise hardware mismatch on NVIDIA's Blackwell B200: tensor core throughput doubles compared to the H100, but shared memory bandwidth and dedicated exponential function units do not scale at the same rate. Rather than waiting for hardware fixes, the authors co-design the attention algorithm with the asymmetric architecture itself — making FlashAttention-4 the first attention kernel built specifically for Blackwell's scaling profile. To frame why this matters, Shannon traces the full lineage of FlashAttention research. The original 2022 NeurIPS paper by Dao and colleagues reframed attention as an IO problem: instead of materializing the quadratic N×N score matrix in slow off-chip High Bandwidth Memory, tiling and online softmax keep computation inside the fast on-chip shared memory of each streaming multiprocessor. FlashAttention-2 doubled throughput through sequence-dimension parallelism. FlashAttention-3 pushed H100 utilization to roughly 75% by exploiting Hopper-specific warp specialization and asynchronous data movement. Each generation addressed a qualitatively different bottleneck — and Blackwell introduced a new one that none of those solutions anticipated. The hosts ground the stakes for practitioners who work in ML without writing GPU kernels. Attention sits at the core of every Transformer-based system — large language models, vision transformers, multimodal architectures — and long-context workloads at 32K to 128K tokens make the quadratic memory cost and HBM round-trips increasingly punishing. Shannon introduces the roofline model as the analytic lens the paper uses to characterize where Blackwell kernels actually bottleneck, setting up how FlashAttention-4's algorithmic co-design approach navigates the compute and memory bandwidth ceilings that previous generations of the kernel never had to contend with.

Sources:
1. FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling — Ted Zadouri, Markus Hoehnerbach, Jay Shah, Timmy Liu, Vijay Thakkar, Tri Dao, 2026
http://arxiv.org/abs/2603.05451v1
2. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Ré, 2022
https://scholar.google.com/scholar?q=FlashAttention%3A+Fast+and+Memory-Efficient+Exact+Attention+with+IO-Awareness
3. FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning — Tri Dao, 2023
https://scholar.google.com/scholar?q=FlashAttention-2%3A+Faster+Attention+with+Better+Parallelism+and+Work+Partitioning
4. FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low Precision — Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, Tri Dao, 2024
https://scholar.google.com/scholar?q=FlashAttention-3%3A+Fast+and+Accurate+Attention+with+Asynchrony+and+Low+Precision
5. Triton: An Intermediate Language and Compiler for Tiled Neural Network Computations — Philippe Tillet, H. T. Kung, David Cox, 2019
https://scholar.google.com/scholar?q=Triton%3A+An+Intermediate+Language+and+Compiler+for+Tiled+Neural+Network+Computations
6. Roofline: An Insightful Visual Performance Model for Floating-Point Programs and Multicore Architectures — Samuel Williams, Andrew Waterman, David Patterson, 2009
https://scholar.google.com/scholar?q=Roofline%3A+An+Insightful+Visual+Performance+Model+for+Floating-Point+Programs+and+Multicore+Architectures
7. Efficiently Scaling Transformer Inference — Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Jonathan Heek, Kefan Xiao, Sharan Narang, Jeff Dean, 2023
https://scholar.google.com/scholar?q=Efficiently+Scaling+Transformer+Inference
8. In-Datacenter Performance Analysis of a Tensor Processing Unit — Norman P. Jouppi et al. (Google), 2017
https://scholar.google.com/scholar?q=In-Datacenter+Performance+Analysis+of+a+Tensor+Processing+Unit
9. Mamba: Linear-Time Sequence Modeling with Selective State Spaces — Albert Gu, Tri Dao, 2023
https://scholar.google.com/scholar?q=Mamba%3A+Linear-Time+Sequence+Modeling+with+Selective+State+Spaces
10. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
11. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism — Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, Bryan Catanzaro, 2019
https://scholar.google.com/scholar?q=Megatron-LM%3A+Training+Multi-Billion+Parameter+Language+Models+Using+Model+Parallelism
12. FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision — Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, Tri Dao, 2024
https://scholar.google.com/scholar?q=FlashAttention-3%3A+Fast+and+Accurate+Attention+with+Asynchrony+and+Low-precision
13. SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization — Jintao Zhang et al., 2024
https://scholar.google.com/scholar?q=SageAttention2%3A+Efficient+Attention+with+Thorough+Outlier+Smoothing+and+Per-thread+INT4+Quantization
14. Ring Attention with Blockwise Transformers for Near-Infinite Context — Hao Liu, Matei Zaharia, Pieter Abbeel, 2023
https://scholar.google.com/scholar?q=Ring+Attention+with+Blockwise+Transformers+for+Near-Infinite+Context
15. FlexAttention: The Flexibility of PyTorch with the Performance of FlashAttention — PyTorch Team, 2024
https://scholar.google.com/scholar?q=FlexAttention%3A+The+Flexibility+of+PyTorch+with+the+Performance+of+FlashAttention
16. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model — DeepSeek-AI, 2024
https://scholar.google.com/scholar?q=DeepSeek-V2%3A+A+Strong%2C+Economical%2C+and+Efficient+Mixture-of-Experts+Language+Model
17. Softmax output approximation for activation memory-efficient training of attention-based networks — approximate — likely 2023–2024, 2023–2024
https://scholar.google.com/scholar?q=Softmax+output+approximation+for+activation+memory-efficient+training+of+attention-based+networks
18. FlashAttention-T: Towards Fully Tensorized Attention by Exploiting Tensor-Vector Parallelism — approximate — likely 2024–2025, 2024–2025
https://scholar.google.com/scholar?q=FlashAttention-T%3A+Towards+Fully+Tensorized+Attention+by+Exploiting+Tensor-Vector+Parallelism
19. Overcoming Long-Context Limitations of State-Space Models via Context-Dependent Sparse Attention — approximate — likely 2024–2025, 2024–2025
https://scholar.google.com/scholar?q=Overcoming+Long-Context+Limitations+of+State-Space+Models+via+Context-Dependent+Sparse+Attention
20. Based on Tensor Core Sparse Kernels Accelerating Deep Neural Networks — approximate — likely 2024–2025, 2024–2025
https://scholar.google.com/scholar?q=Based+on+Tensor+Core+Sparse+Kernels+Accelerating+Deep+Neural+Networks
21. AI Post Transformers: FlashAttention-2: Faster Attention with Better Parallelism — Hal Turing & Dr. Ada Shannon, Fri,
https://podcasters.spotify.com/pod/show/12146088098/episodes/FlashAttention-2-Faster-Attention-with-Better-Parallelism-e36kdm0
22. AI Post Transformers: ATTENTION2D and lean attention: Distributed Self-Attention — Hal Turing & Dr. Ada Shannon, Wed,
https://podcasters.spotify.com/pod/show/12146088098/episodes/ATTENTION2D-and-lean-attention-Distributed-Self-Attention-e3a7r4n
23. AI Post Transformers: Jet-RL: Stable On-Policy Reinforcement Learning with Unified FP8 Flow — Hal Turing & Dr. Ada Shannon, Tue,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Jet-RL-Stable-On-Policy-Reinforcement-Learning-with-Unified-FP8-Flow-e3f7det
24. AI Post Transformers: Mojo: Performance-Portable HPC Kernels on GPUs — Hal Turing & Dr. Ada Shannon, Sat,
https://podcasters.spotify.com/pod/show/12146088098/episodes/Mojo-Performance-Portable-HPC-Kernels-on-GPUs-e39n4sk
Interactive Visualization: FlashAttention-4: Algorithm & Kernel Co-Design

The January 26, 2026 Stanford research paper introduces Agentic Plan Caching (APC), a novel framework designed to reduce the high operational costs of Large Language Model (LLM) agents. Traditional caching methods often fail because agent workflows are highly dynamic and dependent on external environments, making simple input-output storage ineffective. The APC framework solves this by extracting generalized plan templates from successful task executions, which are then indexed by high-level intent keywords. When a similar task is encountered, a small, cost-effective planner LM adapts these cached templates to the new context, significantly reducing the need for expensive, high-reasoning models. Experiments demonstrate that this approach can cut financial and computational costs by over 75% while maintaining high accuracy across complex benchmarks. Ultimately, this system offers a scalable way to deploy sophisticated AI agents by minimizing redundant reasoning through intelligent plan reuse. Source: January 26, 2026 Agentic Plan Caching: Test-Time Memory for Fast and Cost-Efficient LLM Agents Stanford University Qizheng Zhang, Michael Wornow, Gerry Wan, Kunle Olukotun https://arxiv.org/pdf/2506.14852

There is a sharp divergence regarding the utility of long context. Google's Gemini 1.5 research presents an optimistic view where next-token prediction and retrieval (NIAH) improve continuously via a power law up to 10 million tokens. Broader research counters that while *retrieval* scales, utilitarian value (downstream task performance like reasoning or summarization) saturates rapidly or degrades due to "lost-in-the-middle" effects and data scarcity. There is no conclusive position on the empirical utility of long context for complex reasoning; the community must move beyond simple retrieval benchmarks to determine if the immense cost of processing millions of tokens yields proportional functional gains.Sources:1. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of contextDate: March 2024Institutions: Google DeepMindURL:https://arxiv.org/pdf/2403.055302. How to Train Long-Context Language Models (Effectively) [ProLong]Date: 2025Institutions: Princeton UniversityURL: https://aclanthology.org/2025.acl-long.366.pdf3. L2M: Mutual Information Scaling Law for Long-Context Language ModelingDate:** 2025 (NeurIPS)Institutions: MIT, Polytechnic University of Catalonia, Harvard University, UCLAURL:https://arxiv.org/pdf/2503.047254. Predicting Task Performance with Context-aware Scaling LawsDate: October 2025Institutions: UC Santa Cruz, Washington University in St. Louis, Databricks, Google DeepMind, UC BerkeleyURL:https://arxiv.org/pdf/2510.149195. Explaining Context Length Scaling and Bounds for Language ModelsDate:.February 2025Institutions:Tsinghua University, CPHOS Research, Carnegie Mellon University, University of Washington, University of CopenhagenURL:https://arxiv.org/pdf/2502.014816. Scaling Laws and In-Context Learning: A Unified Theoretical FrameworkDate: November 2025 (NeurIPS)URL: https://arxiv.org/pdf/2511.062327. Long-Context Efficient Transformers: A Comprehensive Survey of Techniques, Applications, and Future DirectionsDate: April 10, 2025Institutions: Tsinghua University, Peking University, USTC, Stanford University, UC BerkeleyURL:https://www.techrxiv.org/users/892385/articles/1283745-long-context-efficient-transformers-a-comprehensive-survey-of-techniques-applications-and-future-directions

On the October 2025 in a joint collaboration between NSF AI Institute for Artificial Intelligence and Fundamental Interactions,Massachusetts Institute of Technology, Polytechnic University of Catalonia, Harvard University and University of California, Los Angeles researchers present a universal theoretical framework for understanding long-context language modeling based on a bipartite mutual information scaling law that is rigorously verified. This is in the paper "L2M: Mutual Information Scaling Law for Long-Context Language Modeling". This research paper investigates how large language models manage long-range dependencies by applying principles from information theory. The authors introduce the L2M framework, which establishes that a model's ability to process extensive context is strictly limited by the size of its history state. While transformers utilize a growing cache of data to maintain performance across long sequences, models with fixed-size states, such as SSMs and RNNs, face inherent capacity bottlenecks as input length increases. By utilizing bipartite mutual information as a metric, the study formalizes the theoretical requirements for an architecture to be considered MI-capable. Ultimately, this work provides a principled method for evaluating and designing efficient architectures that can sustain complex reasoning over thousands of tokens. Source: https://arxiv.org/pdf/2503.04725

We focus on the July 2025 paper, "Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient Accumulator". The paper goes into the mathematical details of approximating the FIM, and the easiest is what they call the Squisher: the Adam optimizer variance. This technique allows for tasks like model pruning and model merging to be performed "for free" without the significant computational overhead typically required to calculate the Fisher Information Matrix. We also review the old 1992 paper "Second order derivatives for network pruning: Optimal Brain Surgeon" in terms of what was missing in light of the Squisher paper.Sources:https://arxiv.org/pdf/2507.18807https://proceedings.neurips.cc/paper/1992/file/303ed4c69846ab36c2904d3ba8573050-Paper.pdf

This research presents a novel method for efficient long-context modeling in Large Language Models (LLMs) by tackling the quadratic complexity of attention mechanisms through KV cache compression. The core discovery is a fundamental local KV cache asymmetry, which reveals that adjacent attention keys exhibit high structural homogeneity, while their associated value vectors possess distinct, heterogeneous distributions. To capitalize on this finding, the authors propose AsymKV, a training-free compression framework that shifts information loss from heterogeneous values to homogeneous keys. AsymKV operates by applying homogeneity-based merging to keys using a mathematically derived optimal vector, paired with a lossless value representation scheme utilizing cardinality-aware normalization to preserve vital information. Extensive empirical results on benchmarks like LongBench, across diverse models such as LLaMA3.1-8B, confirm that AsymKV consistently surpasses state-of-the-art long-context methods in terms of accuracy and information retention, offering improved performance with practical inference efficiency. Source: https://arxiv.org/pdf/2506.05410

The source details the creation and evaluation of Agentic Memory (A-MEM), a novel memory system for Large Language Model (LLM) agents that addresses the fundamental rigidity of existing memory architectures. Traditional systems require predefined data structures and fixed operational workflows, which severely limits their ability to adapt to new information and maintain performance in complex, long-term tasks. A-MEM overcomes this by drawing inspiration from the Zettelkasten method, employing dynamic note construction, autonomous link generation, and memory evolution to create a self-organizing knowledge base. Experimental results on long-term dialogue datasets demonstrate that A-MEM significantly outperforms baseline methods across diverse question categories, particularly in challenging multi-hop reasoning tasks. The system is also shown to be highly efficient and scalable, requiring substantially fewer tokens for operation and maintaining minimal increases in retrieval time as the memory scale grows. These architectural advancements allow LLM agents to maintain meaningful, continuously evolving knowledge structures essential for sophisticated interaction with the environment. Source: https://openreview.net/pdf?id=FiM0M8gcct

The provided text outlines DYNAACT, a new framework intended to enhance sequential reasoning in Large Language Models (LLMs) by dynamically managing the available actions during complex problem-solving. This approach targets the inefficiency of current methods that either rely on manually defined and restrictive action spaces or utilize unstructured spaces that prove computationally prohibitive for exhaustive searches. DYNAACT addresses this by first estimating a broad action space from a corpus and then using a greedy algorithm to select an optimal, compact action space for each step. The core of the method is a submodular function that ensures the selected subset of actions maintains a balance between high utility (relevance to the current state) and sufficient diversity (avoiding redundant actions). Extensive evaluation on six benchmarks confirms that DYNAACT significantly improves problem-solving accuracy—especially in math and complex reasoning tasks—while also maintaining efficient inference compared to baseline methods. Source: https://openreview.net/pdf?id=R24ZqNwoDz

The source introduces FlashBias, an innovative algorithm designed to significantly accelerate the efficiency of the Transformer attention mechanism when incorporating an additive bias term. Current methods, like those optimized for attention masks, cannot handle bias because these terms are generally dense and continuous rather than sparse. FlashBias overcomes this limitation by exploiting the mathematical principle that attention bias matrices exhibit an inherent low-rank structure. The technique utilizes several decomposition methods, including exact, SVD, and neural decomposition, to represent the dense bias matrix in a much smaller, compressible form. Experiments showcase substantial time and memory savings when applying FlashBias across various demanding models, such as Large Language Models, Vision Transformers, and AlphaFold 3. This new approach provides crucial efficiency for training and inference, especially for tasks involving dynamic or complex prior knowledge. Source: https://openreview.net/pdf?id=7L4NvUtZY3

The research systematically investigates the effects of integrating various gating mechanisms into the standard softmax attention layer, comparing over thirty configurations across dense and Mixture-of-Experts Large Language Models. The central finding demonstrates that applying an elementwise, head-specific sigmoid gate immediately following the Scaled Dot-Product Attention (SDPA) output consistently yields the most substantial improvement in overall performance metrics. This successful gating method also provides superior training stability, allowing models to converge effectively under larger learning rates and mitigating disruptive loss spikes during optimization. The improved efficacy is attributed to two factors: introducing essential non-linearity into the low-rank attention mapping and generating input-dependent sparse gating scores. Crucially, this sparsity acts to normalize attention dynamics, eliminating the 'attention sink' problem where initial tokens dominate attention scores, thereby facilitating notably better long-context extrapolation. These demonstrated benefits led to the incorporation of this specific gated attention design into the forthcoming Qwen3-Next models. Source: https://openreview.net/pdf?id=1b7whO4SfY

The academic paper introduces KGGen, a novel text-to-knowledge-graph generator designed to overcome the scarcity and poor quality of automatically extracted knowledge graphs (KGs). KGGen utilizes Language Models for initial triple extraction but innovates by employing an iterative clustering and de-duplication process that resolves duplicate entities and relations to reduce sparsity in the final graph representation. To properly assess KG extraction performance, the authors release a new two-part benchmark called Measure of Information in Nodes and Edges (MINE), which evaluates both short-text information retention and knowledge retrieval capabilities in RAG systems. Results on this new benchmark demonstrate that KGGen outperforms competitors like OpenIE and Microsoft's GraphRAG in crucial metrics, including information capture and scaling efficiency across large corpora. The study concludes that KGGen successfully generates KGs with more concise, generalizable entities and relations, which is essential for maximizing utility in downstream applications like embeddings and information retrieval. Source: https://openreview.net/pdf?id=YyhRJXxbpi

This research paper introduces LLaDA, an 8-billion parameter language model based on the masked diffusion model (MDM) architecture, specifically developed to challenge the assumption that core Large Language Model (LLM) capabilities are exclusive to autoregressive models (ARMs). Unlike ARMs that predict the next token sequentially, LLaDA employs a generative approach featuring a forward token-masking process and a reverse process that simultaneously predicts masked tokens using a Transformer network. Trained and evaluated from scratch, LLaDA demonstrates strong scalability and achieves performance comparable to advanced ARM baselines like LLaMA 3 8B across various benchmarks covering general knowledge, math, and code generation. Crucially, the non-autoregressive nature enables bidirectional modeling, which allows LLaDA to effectively address the reversal curse and outperform contemporary models, including GPT-4o, on complex reversal reasoning tasks. These findings confirm that fundamental generative modeling principles, rather than dependence on sequential ARMs, underpin essential LLM capabilities. The work concludes that diffusion models offer a promising new paradigm for building robust, large-scale language models. Source: https://openreview.net/pdf?id=KnqiC0znVF

This paper introduces Mixture of Block Attention (MoBA) to address the prohibitive quadratic computational overhead inherent in traditional attention mechanisms when scaling large language models (LLMs) for long contexts. MoBA is a novel architecture that strategically applies the established Mixture of Experts (MoE) paradigm directly to the attention mechanism itself. Instead of attending to the entire sequence, MoBA partitions the context into discrete blocks and utilizes a dynamic gating network to selectively route queries to only the most relevant blocks of keys and values. This block-sparse approach drastically increases computational efficiency, achieving sub-quadratic complexity and demonstrating speedups of up to 16 times when processing sequences up to 10 million tokens. Crucially, the research demonstrates that MoBA maintains performance comparable to full attention across scaling laws and real-world benchmarks. Furthermore, the architecture is highly flexible, allowing for seamless transitions between sparse MoBA and full attention layers during both training and inference. Source: https://openreview.net/pdf?id=RlqYCpTu1P

The research proposes Parallel Scaling (PARSCALE) as a novel, efficient strategy to enhance Large Language Model (LLM) capacity by increasing parallel computation rather than merely growing the parameter count. This method reuses existing model parameters by feeding multiple parallel input streams (differentiated by learned prefixes) and dynamically combining their outputs into a single prediction. Through extensive testing, the paper develops a new scaling law, showing that scaling computation by a factor of P provides performance gains roughly equivalent to scaling parameters by a factor of O(N logP). PARSCALE demonstrates particular effectiveness in boosting performance on reasoning-intensive tasks like coding and mathematics problems. Critically, this scaling technique offers superior efficiency during inference, requiring significantly less memory and time increase than traditional parameter scaling, thereby making it highly suitable for low-resource edge deployment. Source: https://openreview.net/pdf?id=dEi1S731lk

This research examines the data efficiency of Reinforcement Learning with Verifiable Reward (RLVR) when applied to large language models for mathematical reasoning tasks. The paper's most significant finding is the success of 1-shot RLVR, showing that comparable performance to using a large training dataset can be achieved using just a single, carefully selected example. This result suggests that RLVR is effective primarily because it activates the strong latent reasoning capabilities already present in the base model, rather than imparting new domain knowledge. An interesting phenomenon observed during training is "post-saturation generalization," where the model's test performance continues to rise long after training accuracy has saturated and the model has begun overfitting the single example. Ablation studies indicate that while policy gradient loss is the main source of improvement, entropy loss is essential for encouraging the exploration needed to realize this enhanced long-term generalization. Source: https://openreview.net/pdf?id=IBrRNLr6JA

The source details the development and evaluation of Reward Reasoning Models (RRMs), which are designed to enhance Large Language Model (LLM) alignment by incorporating an explicit chain-of-thought reasoning process before generating a final reward. This innovative structure enables RRMs to adaptively utilize computational resources at inference time for complex evaluation tasks requiring nuanced judgment. The models are trained using a novel reinforcement learning framework that promotes the self-evolution of reasoning skills without requiring explicit reasoning traces as initial training data. Experimental results confirm that RRMs achieve superior performance across diverse reward modeling and reasoning benchmarks, often outperforming competing models with much larger parameter sizes. The document further validates the practical effectiveness of RRMs in tasks such as reward-guided best-of-N response selection and robust LLM post-training alignment. Overall, the work establishes a new state-of-the-art approach by demonstrating the scalable benefits of marrying reasoning capabilities with reward prediction. Source: https://openreview.net/pdf?id=V8Kbz7l2cr

The academic paper presents the Self-Adapting LLM (SEAL) framework, designed to allow large language models to overcome their static nature by transforming and generating their own fine-tuning data. This mechanism involves the model producing a "self-edit," which consists of natural-language instructions that specify synthetic data, tool invocations, or optimization hyperparameters for adaptation. Training is managed by an outer reinforcement learning (RL) loop that rewards the model based on the improved performance achieved after the self-edit results in persistent weight updates via supervised fine-tuning. Evaluations show that SEAL significantly enhances both knowledge incorporation of new factual data and few-shot generalization on abstract reasoning tasks. Ultimately, the authors propose this work as a viable strategy for enabling models to pursue self-directed, continual learning in preparation for a future where traditional human-generated data sources are exhausted. Source: https://openreview.net/pdf?id=JsNUE84Hxi

The academic paper introduces Self-play Reinforcement Learning (SeRL), a framework engineered to enhance the reasoning capabilities of Large Language Models (LLMs) specifically in scenarios lacking extensive, high-quality labeled data. SeRL consists of two core, complementary modules: the self-instruction module generates new and diverse training problems from a small seed dataset, ensuring data quality and appropriate difficulty via an online filtering strategy. Simultaneously, the self-rewarding module bypasses the need for external supervision by estimating response rewards using a stable majority-voting mechanism among sampled outputs. This integrated approach facilitates sustained, unsupervised reinforcement learning across multiple training iterations. Experiments demonstrate that SeRL is highly effective, consistently outperforming existing self-play methods and matching the performance levels achieved by models trained on full datasets with verifiable rewards. Source: https://openreview.net/pdf?id=ZF93vyH9He

The research introduces Thinkless, a framework designed to solve the computational inefficiency of Large Language Models (LLMs) that overuse chain-of-thought reasoning for simple queries. This adaptive model determines whether to utilize a concise () or detailed reasoning () mode based on the input complexity and its own capabilities. Central to this approach is the Decoupled Group Relative Policy Optimization (DeGRPO) algorithm, which employs reinforcement learning to jointly optimize both the selection of the reasoning mode and the accuracy of the final answer. DeGRPO stabilizes training by balancing the gradient signals between the control tokens and the response tokens, successfully preventing policy collapse observed in traditional reinforcement learning methods. Empirically, the model effectively handles varied tasks, demonstrating its ability to reduce the reliance on computationally expensive, long-form reasoning by 50% to 90% on mathematical benchmarks while maintaining performance. Source: https://openreview.net/pdf?id=ariVQf0KZx

We provide a review of the evolution of value of Page Rank to Random Walk with Random Restart and it's application to neural networks focusing on five research papers dating from the original page rank to 2025. They collectively focus on methods for learning on graphs, particularly through the use of Random Walk Neural Networks (RWNNs) and related random walk algorithms. One primary source introduces RWNNs, detailing their architecture, which involves a random walk generating a machine-readable record processed by a deep neural network, demonstrating that these models can achieve universal approximation of graph functions and overcome issues like over-smoothing found in Message Passing Neural Networks (MPNNs). This source also explores techniques like anonymization and named neighbors for walk recording and includes experimental results on graph isomorphism and transductive classification using language models like DeBERTa and Llama 3. The other sources provide brief contextual support, mentioning Random Walk with Restart (RWR) parameters and evaluation criteria like Relative Accuracy and Relative Score for related graph applications and datasets, suggesting connections to established graph algorithms such as PageRank.Sources:2025:REVISITING RANDOM WALKS FOR LEARNING ON GRAPHShttps://proceedings.iclr.cc/paper_files/paper/2025/file/cd51b67dcb19db4e9f0022f500076b00-Paper-Conference.pdfOctober 3, 2022:Universal Multilayer Network Exploration byRandom Walk with Restarthttps://arxiv.org/pdf/2107.045652020:Random Walk Graph Neural Networkshttps://proceedings.neurips.cc/paper/2020/file/ba95d78a7c942571185308775a97a3a0-Paper.pdf2006:Fast Random Walk with Restart and Its Applicationshttps://www.cs.cmu.edu/~htong/pdf/ICDM06_tong.pdfJanuary 29, 1998:The Page Rank Citation Ranking: Bringing Order to the Webhttps://www.cis.upenn.edu/~mkearns/teaching/NetworkedLife/pagerank.pdf

These 14 research papers provide an overview of various compression techniques for Large Language Models (LLMs), primarily focusing on reducing the size and computational overhead of the Key-Value (KV) cache to handle long contexts more efficiently. Several novel methods are detailed, including GVote, an adaptive compression algorithm using query sampling and voting to find an optimal cache budget, and SnapKV, which selects clustered, important KV positions based on an "observation" window to maintain performance while increasing speed and memory efficiency. Other approaches include POD (Proximal tokens over Distant tokens), which reduces redundancy by sharing key states across layers for distant tokens while preserving proximal ones, and DecoQuant, a quantization method utilizing matrix decomposition to reduce errors. The sources also examine prompt compression methods like LLMLingua and LongLLMLingua, and describe CASC (Context-Adaptive Synthesis and Compression), a Retrieval-Augmented Generation (RAG) framework that intelligently synthesizes and compresses multi-document contexts to improve answer accuracy in complex domains.Sources:https://arxiv.org/pdf/2509.08315https://arxiv.org/html/2509.09199v1https://arxiv.org/html/2509.03136v1https://aclanthology.org/2025.acl-long.1394.pdfhttps://proceedings.neurips.cc/paper_files/paper/2024/file/fd0705710bf01b88a60a3d479ea341d9-Paper-Conference.pdfhttps://arxiv.org/html/2412.14838v1https://arxiv.org/pdf/2412.02252https://aclanthology.org/2024.acl-long.133.pdfhttps://arxiv.org/html/2508.19357v1https://aclanthology.org/2024.acl-long.91.pdfhttps://arxiv.org/html/2310.05736v2https://aclanthology.org/2025.naacl-long.368.pdfhttps://arxiv.org/pdf/2404.14469https://aclanthology.org/2024.findings-emnlp.266.pdf

This academic paper introduces movement pruning, a novel method for reducing the size of large pre-trained language models like BERT during fine-tuning. Unlike traditional magnitude pruning which removes weights based on their absolute values, movement pruning prioritizes weights that change significantly during the fine-tuning process, demonstrating superior performance in high-sparsity scenarios. The authors provide mathematical foundations for their approach and empirically compare it against existing zeroth- and first-order pruning techniques, highlighting its effectiveness, especially when combined with distillation. The research emphasizes the potential for resource reduction, enabling the deployment of complex models on less powerful hardware and fostering broader accessibility in the field of natural language processing. Source: Published 2020 https://papers.neurips.cc/paper_files/paper/2020/file/eae15aabaa768ae4a5993a8a4f4fa6e4-Paper.pdf