Hal Turing and Dr. Ada Shannon open the episode by confronting a structural flaw that has been hiding in plain sight since the transformer era began: tokenization bias. The episode centers on "Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model Ensembles" by Buu Phan, Brandon Amos, Itai Gat, Marton Havasi, Matthew Muckley, and Karen Ullrich (ICLR 2025), which formally proves that a tokenized model and a byte-level model can be statistically equivalent and still produce wildly different predictions for the same next character. The hosts trace the origins of the problem through BPE's introduction by Rico Sennrich, Barry Haddow, and Alexandra Birch in 2016 and its industrialization via Kudo and Richardson's SentencePiece in 2018 — a library now frozen into the spine of LLaMA, Mistral, Gemma, and most open-source models not in the OpenAI lineage. The discussion sharpens around fill-in-the-middle prompting, the paradigm introduced by Mohammad Bavarian and colleagues at OpenAI in 2022 and now embedded in every major code completion tool from GitHub Copilot to StarCoder. Shannon walks through the paper's central example: a code completion scenario where the correct next character receives a probability of exactly zero — not a rounding artifact but a structural impossibility, because the tokenizer has carved up the prompt in a way that makes the right answer unreachable in token-space. Turing challenges the framing, arguing that byte-level alternatives like ByT5 and MegaByte existed and BPE was an informed trade-off against the three-to-eight times sequence length penalty that raw bytes impose on attention compute. Shannon holds the line: the point is not that BPE was a mistake but that its systematic bias was never formally characterized until now, and the Byte-Token Representation Lemma finally gives the field the mathematical language to name and measure it. The episode closes by introducing the second paper from the episode's pairing — Minixhofer et al.'s NeurIPS 2025 work on cross-tokenizer knowledge distillation — which attacks the tokenizer barrier from the training side rather than the inference side. Where Phan et al. offer a zero-shot correction algorithm that recovers 18% on fill-in-the-middle coding benchmarks without any retraining, Minixhofer et al. enable knowledge transfer between models with fundamentally incompatible vocabularies, breaking the assumption that distillation requires shared tokenization. Together the two papers sketch a trajectory where tokenization becomes a transparent implementation detail rather than an architectural constraint that determines what a model can and cannot express.
Hal Turing and Dr. Ada Shannon return to the CARTRIDGE compression system with a mechanistic lens, covering Maurizio A. Diaz's paper "Learned Structure in Cartridges: Keys as Shareable Routers in Self-Studied Representations" (arXiv 2508.17032), presented at the NeurIPS 2025 Workshop on Mechanistic Interpretability. Building on the original CARTRIDGE episode from November 10th, 2025 and the follow-up from February 6th, 2026, this episode asks the question those earlier discussions left open: what structure does the optimizer actually induce in a trained CARTRIDGE? The hosts ground the discussion in the memory scaling problem driving the entire field—KV caches that grow linearly with context length, now routinely dwarfing model weights at the 128K-to-million-token scales of current frontier models—and trace how techniques like PagedAttention, Grouped Query Attention, and token eviction address symptoms without shrinking the underlying representation. Diaz's central finding is a clean functional division between key and value vectors inside a trained CARTRIDGE. Keys converge to stable retrieval routers: low-rank, consistent structures that steer attention toward the right stored content across diverse queries. Values carry the compressed semantic payload. The hosts connect this directly to how CARTRIDGE's Self-Study training pipeline works—because the cache is optimized against synthetic question-answer traces generated by the model over its own content, the training signal explicitly selects for routing behavior, making the key-as-router outcome a predictable consequence of the objective rather than an accident. Diaz uses Singular Value Decomposition to quantify this structure layer by layer, separating the geometric properties of key matrices from value matrices across training checkpoints. Two downstream findings from the key-router property shape the second half of the discussion. Because keys are stable and low-rank, they transfer across tasks with minimal degradation—a result with direct implications for multi-task serving, where a single shared key structure could route to task-specific value sets without independent CARTRIDGE training per deployment. The Sampled Chunk Initialization method introduced in the paper exploits this stability to warm-start CARTRIDGE training, accelerating convergence by initializing the learnable KV pairs from a small representative sample rather than random weights. Hal and Ada close by discussing what the key-as-router framing implies for KV-cache compression research more broadly: if the routing function is separable and transferable, compression schemes that conflate keys and values may be discarding structure that has real serving-efficiency value.
Hal Turing and Dr. Ada Shannon dig into FlashAttention-4, a March 2026 paper from a cross-institutional team including Tri Dao, Jay Shah, and colleagues at Princeton, Meta, NVIDIA, Colfax Research, Georgia Tech, and Together AI. The paper targets a precise hardware mismatch on NVIDIA's Blackwell B200: tensor core throughput doubles compared to the H100, but shared memory bandwidth and dedicated exponential function units do not scale at the same rate. Rather than waiting for hardware fixes, the authors co-design the attention algorithm with the asymmetric architecture itself — making FlashAttention-4 the first attention kernel built specifically for Blackwell's scaling profile. To frame why this matters, Shannon traces the full lineage of FlashAttention research. The original 2022 NeurIPS paper by Dao and colleagues reframed attention as an IO problem: instead of materializing the quadratic N×N score matrix in slow off-chip High Bandwidth Memory, tiling and online softmax keep computation inside the fast on-chip shared memory of each streaming multiprocessor. FlashAttention-2 doubled throughput through sequence-dimension parallelism. FlashAttention-3 pushed H100 utilization to roughly 75% by exploiting Hopper-specific warp specialization and asynchronous data movement. Each generation addressed a qualitatively different bottleneck — and Blackwell introduced a new one that none of those solutions anticipated. The hosts ground the stakes for practitioners who work in ML without writing GPU kernels. Attention sits at the core of every Transformer-based system — large language models, vision transformers, multimodal architectures — and long-context workloads at 32K to 128K tokens make the quadratic memory cost and HBM round-trips increasingly punishing. Shannon introduces the roofline model as the analytic lens the paper uses to characterize where Blackwell kernels actually bottleneck, setting up how FlashAttention-4's algorithmic co-design approach navigates the compute and memory bandwidth ceilings that previous generations of the kernel never had to contend with.
The January 26, 2026 Stanford research paper introduces Agentic Plan Caching (APC), a novel framework designed to reduce the high operational costs of Large Language Model (LLM) agents. Traditional caching methods often fail because agent workflows are highly dynamic and dependent on external environments, making simple input-output storage ineffective. The APC framework solves this by extracting generalized plan templates from successful task executions, which are then indexed by high-level intent keywords. When a similar task is encountered, a small, cost-effective planner LM adapts these cached templates to the new context, significantly reducing the need for expensive, high-reasoning models. Experiments demonstrate that this approach can cut financial and computational costs by over 75% while maintaining high accuracy across complex benchmarks. Ultimately, this system offers a scalable way to deploy sophisticated AI agents by minimizing redundant reasoning through intelligent plan reuse. Source: January 26, 2026 Agentic Plan Caching: Test-Time Memory for Fast and Cost-Efficient LLM Agents Stanford University Qizheng Zhang, Michael Wornow, Gerry Wan, Kunle Olukotun https://arxiv.org/pdf/2506.14852
There is a sharp divergence regarding the utility of long context. Google's Gemini 1.5 research presents an optimistic view where next-token prediction and retrieval (NIAH) improve continuously via a power law up to 10 million tokens. Broader research counters that while *retrieval* scales, utilitarian value (downstream task performance like reasoning or summarization) saturates rapidly or degrades due to "lost-in-the-middle" effects and data scarcity. There is no conclusive position on the empirical utility of long context for complex reasoning; the community must move beyond simple retrieval benchmarks to determine if the immense cost of processing millions of tokens yields proportional functional gains.Sources:1. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of contextDate: March 2024Institutions: Google DeepMindURL:https://arxiv.org/pdf/2403.055302. How to Train Long-Context Language Models (Effectively) [ProLong]Date: 2025Institutions: Princeton UniversityURL: https://aclanthology.org/2025.acl-long.366.pdf3. L2M: Mutual Information Scaling Law for Long-Context Language ModelingDate:** 2025 (NeurIPS)Institutions: MIT, Polytechnic University of Catalonia, Harvard University, UCLAURL:https://arxiv.org/pdf/2503.047254. Predicting Task Performance with Context-aware Scaling LawsDate: October 2025Institutions: UC Santa Cruz, Washington University in St. Louis, Databricks, Google DeepMind, UC BerkeleyURL:https://arxiv.org/pdf/2510.149195. Explaining Context Length Scaling and Bounds for Language ModelsDate:.February 2025Institutions:Tsinghua University, CPHOS Research, Carnegie Mellon University, University of Washington, University of CopenhagenURL:https://arxiv.org/pdf/2502.014816. Scaling Laws and In-Context Learning: A Unified Theoretical FrameworkDate: November 2025 (NeurIPS)URL: https://arxiv.org/pdf/2511.062327. Long-Context Efficient Transformers: A Comprehensive Survey of Techniques, Applications, and Future DirectionsDate: April 10, 2025Institutions: Tsinghua University, Peking University, USTC, Stanford University, UC BerkeleyURL:https://www.techrxiv.org/users/892385/articles/1283745-long-context-efficient-transformers-a-comprehensive-survey-of-techniques-applications-and-future-directions
On the October 2025 in a joint collaboration between NSF AI Institute for Artificial Intelligence and Fundamental Interactions,Massachusetts Institute of Technology, Polytechnic University of Catalonia, Harvard University and University of California, Los Angeles researchers present a universal theoretical framework for understanding long-context language modeling based on a bipartite mutual information scaling law that is rigorously verified. This is in the paper "L2M: Mutual Information Scaling Law for Long-Context Language Modeling". This research paper investigates how large language models manage long-range dependencies by applying principles from information theory. The authors introduce the L2M framework, which establishes that a model's ability to process extensive context is strictly limited by the size of its history state. While transformers utilize a growing cache of data to maintain performance across long sequences, models with fixed-size states, such as SSMs and RNNs, face inherent capacity bottlenecks as input length increases. By utilizing bipartite mutual information as a metric, the study formalizes the theoretical requirements for an architecture to be considered MI-capable. Ultimately, this work provides a principled method for evaluating and designing efficient architectures that can sustain complex reasoning over thousands of tokens. Source: https://arxiv.org/pdf/2503.04725
We focus on the July 2025 paper, "Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient Accumulator". The paper goes into the mathematical details of approximating the FIM, and the easiest is what they call the Squisher: the Adam optimizer variance. This technique allows for tasks like model pruning and model merging to be performed "for free" without the significant computational overhead typically required to calculate the Fisher Information Matrix. We also review the old 1992 paper "Second order derivatives for network pruning: Optimal Brain Surgeon" in terms of what was missing in light of the Squisher paper.Sources:https://arxiv.org/pdf/2507.18807https://proceedings.neurips.cc/paper/1992/file/303ed4c69846ab36c2904d3ba8573050-Paper.pdf
This research presents a novel method for efficient long-context modeling in Large Language Models (LLMs) by tackling the quadratic complexity of attention mechanisms through KV cache compression. The core discovery is a fundamental local KV cache asymmetry, which reveals that adjacent attention keys exhibit high structural homogeneity, while their associated value vectors possess distinct, heterogeneous distributions. To capitalize on this finding, the authors propose AsymKV, a training-free compression framework that shifts information loss from heterogeneous values to homogeneous keys. AsymKV operates by applying homogeneity-based merging to keys using a mathematically derived optimal vector, paired with a lossless value representation scheme utilizing cardinality-aware normalization to preserve vital information. Extensive empirical results on benchmarks like LongBench, across diverse models such as LLaMA3.1-8B, confirm that AsymKV consistently surpasses state-of-the-art long-context methods in terms of accuracy and information retention, offering improved performance with practical inference efficiency. Source: https://arxiv.org/pdf/2506.05410
The source details the creation and evaluation of Agentic Memory (A-MEM), a novel memory system for Large Language Model (LLM) agents that addresses the fundamental rigidity of existing memory architectures. Traditional systems require predefined data structures and fixed operational workflows, which severely limits their ability to adapt to new information and maintain performance in complex, long-term tasks. A-MEM overcomes this by drawing inspiration from the Zettelkasten method, employing dynamic note construction, autonomous link generation, and memory evolution to create a self-organizing knowledge base. Experimental results on long-term dialogue datasets demonstrate that A-MEM significantly outperforms baseline methods across diverse question categories, particularly in challenging multi-hop reasoning tasks. The system is also shown to be highly efficient and scalable, requiring substantially fewer tokens for operation and maintaining minimal increases in retrieval time as the memory scale grows. These architectural advancements allow LLM agents to maintain meaningful, continuously evolving knowledge structures essential for sophisticated interaction with the environment. Source: https://openreview.net/pdf?id=FiM0M8gcct
The provided text outlines DYNAACT, a new framework intended to enhance sequential reasoning in Large Language Models (LLMs) by dynamically managing the available actions during complex problem-solving. This approach targets the inefficiency of current methods that either rely on manually defined and restrictive action spaces or utilize unstructured spaces that prove computationally prohibitive for exhaustive searches. DYNAACT addresses this by first estimating a broad action space from a corpus and then using a greedy algorithm to select an optimal, compact action space for each step. The core of the method is a submodular function that ensures the selected subset of actions maintains a balance between high utility (relevance to the current state) and sufficient diversity (avoiding redundant actions). Extensive evaluation on six benchmarks confirms that DYNAACT significantly improves problem-solving accuracy—especially in math and complex reasoning tasks—while also maintaining efficient inference compared to baseline methods. Source: https://openreview.net/pdf?id=R24ZqNwoDz
The source introduces FlashBias, an innovative algorithm designed to significantly accelerate the efficiency of the Transformer attention mechanism when incorporating an additive bias term. Current methods, like those optimized for attention masks, cannot handle bias because these terms are generally dense and continuous rather than sparse. FlashBias overcomes this limitation by exploiting the mathematical principle that attention bias matrices exhibit an inherent low-rank structure. The technique utilizes several decomposition methods, including exact, SVD, and neural decomposition, to represent the dense bias matrix in a much smaller, compressible form. Experiments showcase substantial time and memory savings when applying FlashBias across various demanding models, such as Large Language Models, Vision Transformers, and AlphaFold 3. This new approach provides crucial efficiency for training and inference, especially for tasks involving dynamic or complex prior knowledge. Source: https://openreview.net/pdf?id=7L4NvUtZY3
The research systematically investigates the effects of integrating various gating mechanisms into the standard softmax attention layer, comparing over thirty configurations across dense and Mixture-of-Experts Large Language Models. The central finding demonstrates that applying an elementwise, head-specific sigmoid gate immediately following the Scaled Dot-Product Attention (SDPA) output consistently yields the most substantial improvement in overall performance metrics. This successful gating method also provides superior training stability, allowing models to converge effectively under larger learning rates and mitigating disruptive loss spikes during optimization. The improved efficacy is attributed to two factors: introducing essential non-linearity into the low-rank attention mapping and generating input-dependent sparse gating scores. Crucially, this sparsity acts to normalize attention dynamics, eliminating the 'attention sink' problem where initial tokens dominate attention scores, thereby facilitating notably better long-context extrapolation. These demonstrated benefits led to the incorporation of this specific gated attention design into the forthcoming Qwen3-Next models. Source: https://openreview.net/pdf?id=1b7whO4SfY
The academic paper introduces KGGen, a novel text-to-knowledge-graph generator designed to overcome the scarcity and poor quality of automatically extracted knowledge graphs (KGs). KGGen utilizes Language Models for initial triple extraction but innovates by employing an iterative clustering and de-duplication process that resolves duplicate entities and relations to reduce sparsity in the final graph representation. To properly assess KG extraction performance, the authors release a new two-part benchmark called Measure of Information in Nodes and Edges (MINE), which evaluates both short-text information retention and knowledge retrieval capabilities in RAG systems. Results on this new benchmark demonstrate that KGGen outperforms competitors like OpenIE and Microsoft's GraphRAG in crucial metrics, including information capture and scaling efficiency across large corpora. The study concludes that KGGen successfully generates KGs with more concise, generalizable entities and relations, which is essential for maximizing utility in downstream applications like embeddings and information retrieval. Source: https://openreview.net/pdf?id=YyhRJXxbpi
This research paper introduces LLaDA, an 8-billion parameter language model based on the masked diffusion model (MDM) architecture, specifically developed to challenge the assumption that core Large Language Model (LLM) capabilities are exclusive to autoregressive models (ARMs). Unlike ARMs that predict the next token sequentially, LLaDA employs a generative approach featuring a forward token-masking process and a reverse process that simultaneously predicts masked tokens using a Transformer network. Trained and evaluated from scratch, LLaDA demonstrates strong scalability and achieves performance comparable to advanced ARM baselines like LLaMA 3 8B across various benchmarks covering general knowledge, math, and code generation. Crucially, the non-autoregressive nature enables bidirectional modeling, which allows LLaDA to effectively address the reversal curse and outperform contemporary models, including GPT-4o, on complex reversal reasoning tasks. These findings confirm that fundamental generative modeling principles, rather than dependence on sequential ARMs, underpin essential LLM capabilities. The work concludes that diffusion models offer a promising new paradigm for building robust, large-scale language models. Source: https://openreview.net/pdf?id=KnqiC0znVF
This paper introduces Mixture of Block Attention (MoBA) to address the prohibitive quadratic computational overhead inherent in traditional attention mechanisms when scaling large language models (LLMs) for long contexts. MoBA is a novel architecture that strategically applies the established Mixture of Experts (MoE) paradigm directly to the attention mechanism itself. Instead of attending to the entire sequence, MoBA partitions the context into discrete blocks and utilizes a dynamic gating network to selectively route queries to only the most relevant blocks of keys and values. This block-sparse approach drastically increases computational efficiency, achieving sub-quadratic complexity and demonstrating speedups of up to 16 times when processing sequences up to 10 million tokens. Crucially, the research demonstrates that MoBA maintains performance comparable to full attention across scaling laws and real-world benchmarks. Furthermore, the architecture is highly flexible, allowing for seamless transitions between sparse MoBA and full attention layers during both training and inference. Source: https://openreview.net/pdf?id=RlqYCpTu1P
The research proposes Parallel Scaling (PARSCALE) as a novel, efficient strategy to enhance Large Language Model (LLM) capacity by increasing parallel computation rather than merely growing the parameter count. This method reuses existing model parameters by feeding multiple parallel input streams (differentiated by learned prefixes) and dynamically combining their outputs into a single prediction. Through extensive testing, the paper develops a new scaling law, showing that scaling computation by a factor of P provides performance gains roughly equivalent to scaling parameters by a factor of O(N logP). PARSCALE demonstrates particular effectiveness in boosting performance on reasoning-intensive tasks like coding and mathematics problems. Critically, this scaling technique offers superior efficiency during inference, requiring significantly less memory and time increase than traditional parameter scaling, thereby making it highly suitable for low-resource edge deployment. Source: https://openreview.net/pdf?id=dEi1S731lk
This research examines the data efficiency of Reinforcement Learning with Verifiable Reward (RLVR) when applied to large language models for mathematical reasoning tasks. The paper's most significant finding is the success of 1-shot RLVR, showing that comparable performance to using a large training dataset can be achieved using just a single, carefully selected example. This result suggests that RLVR is effective primarily because it activates the strong latent reasoning capabilities already present in the base model, rather than imparting new domain knowledge. An interesting phenomenon observed during training is "post-saturation generalization," where the model's test performance continues to rise long after training accuracy has saturated and the model has begun overfitting the single example. Ablation studies indicate that while policy gradient loss is the main source of improvement, entropy loss is essential for encouraging the exploration needed to realize this enhanced long-term generalization. Source: https://openreview.net/pdf?id=IBrRNLr6JA
The source details the development and evaluation of Reward Reasoning Models (RRMs), which are designed to enhance Large Language Model (LLM) alignment by incorporating an explicit chain-of-thought reasoning process before generating a final reward. This innovative structure enables RRMs to adaptively utilize computational resources at inference time for complex evaluation tasks requiring nuanced judgment. The models are trained using a novel reinforcement learning framework that promotes the self-evolution of reasoning skills without requiring explicit reasoning traces as initial training data. Experimental results confirm that RRMs achieve superior performance across diverse reward modeling and reasoning benchmarks, often outperforming competing models with much larger parameter sizes. The document further validates the practical effectiveness of RRMs in tasks such as reward-guided best-of-N response selection and robust LLM post-training alignment. Overall, the work establishes a new state-of-the-art approach by demonstrating the scalable benefits of marrying reasoning capabilities with reward prediction. Source: https://openreview.net/pdf?id=V8Kbz7l2cr
The academic paper presents the Self-Adapting LLM (SEAL) framework, designed to allow large language models to overcome their static nature by transforming and generating their own fine-tuning data. This mechanism involves the model producing a "self-edit," which consists of natural-language instructions that specify synthetic data, tool invocations, or optimization hyperparameters for adaptation. Training is managed by an outer reinforcement learning (RL) loop that rewards the model based on the improved performance achieved after the self-edit results in persistent weight updates via supervised fine-tuning. Evaluations show that SEAL significantly enhances both knowledge incorporation of new factual data and few-shot generalization on abstract reasoning tasks. Ultimately, the authors propose this work as a viable strategy for enabling models to pursue self-directed, continual learning in preparation for a future where traditional human-generated data sources are exhausted. Source: https://openreview.net/pdf?id=JsNUE84Hxi
The academic paper introduces Self-play Reinforcement Learning (SeRL), a framework engineered to enhance the reasoning capabilities of Large Language Models (LLMs) specifically in scenarios lacking extensive, high-quality labeled data. SeRL consists of two core, complementary modules: the self-instruction module generates new and diverse training problems from a small seed dataset, ensuring data quality and appropriate difficulty via an online filtering strategy. Simultaneously, the self-rewarding module bypasses the need for external supervision by estimating response rewards using a stable majority-voting mechanism among sampled outputs. This integrated approach facilitates sustained, unsupervised reinforcement learning across multiple training iterations. Experiments demonstrate that SeRL is highly effective, consistently outperforming existing self-play methods and matching the performance levels achieved by models trained on full datasets with verifiable rewards. Source: https://openreview.net/pdf?id=ZF93vyH9He
The research introduces Thinkless, a framework designed to solve the computational inefficiency of Large Language Models (LLMs) that overuse chain-of-thought reasoning for simple queries. This adaptive model determines whether to utilize a concise () or detailed reasoning () mode based on the input complexity and its own capabilities. Central to this approach is the Decoupled Group Relative Policy Optimization (DeGRPO) algorithm, which employs reinforcement learning to jointly optimize both the selection of the reasoning mode and the accuracy of the final answer. DeGRPO stabilizes training by balancing the gradient signals between the control tokens and the response tokens, successfully preventing policy collapse observed in traditional reinforcement learning methods. Empirically, the model effectively handles varied tasks, demonstrating its ability to reduce the reliance on computationally expensive, long-form reasoning by 50% to 90% on mathematical benchmarks while maintaining performance. Source: https://openreview.net/pdf?id=ariVQf0KZx
We provide a review of the evolution of value of Page Rank to Random Walk with Random Restart and it's application to neural networks focusing on five research papers dating from the original page rank to 2025. They collectively focus on methods for learning on graphs, particularly through the use of Random Walk Neural Networks (RWNNs) and related random walk algorithms. One primary source introduces RWNNs, detailing their architecture, which involves a random walk generating a machine-readable record processed by a deep neural network, demonstrating that these models can achieve universal approximation of graph functions and overcome issues like over-smoothing found in Message Passing Neural Networks (MPNNs). This source also explores techniques like anonymization and named neighbors for walk recording and includes experimental results on graph isomorphism and transductive classification using language models like DeBERTa and Llama 3. The other sources provide brief contextual support, mentioning Random Walk with Restart (RWR) parameters and evaluation criteria like Relative Accuracy and Relative Score for related graph applications and datasets, suggesting connections to established graph algorithms such as PageRank.Sources:2025:REVISITING RANDOM WALKS FOR LEARNING ON GRAPHShttps://proceedings.iclr.cc/paper_files/paper/2025/file/cd51b67dcb19db4e9f0022f500076b00-Paper-Conference.pdfOctober 3, 2022:Universal Multilayer Network Exploration byRandom Walk with Restarthttps://arxiv.org/pdf/2107.045652020:Random Walk Graph Neural Networkshttps://proceedings.neurips.cc/paper/2020/file/ba95d78a7c942571185308775a97a3a0-Paper.pdf2006:Fast Random Walk with Restart and Its Applicationshttps://www.cs.cmu.edu/~htong/pdf/ICDM06_tong.pdfJanuary 29, 1998:The Page Rank Citation Ranking: Bringing Order to the Webhttps://www.cis.upenn.edu/~mkearns/teaching/NetworkedLife/pagerank.pdf
These 14 research papers provide an overview of various compression techniques for Large Language Models (LLMs), primarily focusing on reducing the size and computational overhead of the Key-Value (KV) cache to handle long contexts more efficiently. Several novel methods are detailed, including GVote, an adaptive compression algorithm using query sampling and voting to find an optimal cache budget, and SnapKV, which selects clustered, important KV positions based on an "observation" window to maintain performance while increasing speed and memory efficiency. Other approaches include POD (Proximal tokens over Distant tokens), which reduces redundancy by sharing key states across layers for distant tokens while preserving proximal ones, and DecoQuant, a quantization method utilizing matrix decomposition to reduce errors. The sources also examine prompt compression methods like LLMLingua and LongLLMLingua, and describe CASC (Context-Adaptive Synthesis and Compression), a Retrieval-Augmented Generation (RAG) framework that intelligently synthesizes and compresses multi-document contexts to improve answer accuracy in complex domains.Sources:https://arxiv.org/pdf/2509.08315https://arxiv.org/html/2509.09199v1https://arxiv.org/html/2509.03136v1https://aclanthology.org/2025.acl-long.1394.pdfhttps://proceedings.neurips.cc/paper_files/paper/2024/file/fd0705710bf01b88a60a3d479ea341d9-Paper-Conference.pdfhttps://arxiv.org/html/2412.14838v1https://arxiv.org/pdf/2412.02252https://aclanthology.org/2024.acl-long.133.pdfhttps://arxiv.org/html/2508.19357v1https://aclanthology.org/2024.acl-long.91.pdfhttps://arxiv.org/html/2310.05736v2https://aclanthology.org/2025.naacl-long.368.pdfhttps://arxiv.org/pdf/2404.14469https://aclanthology.org/2024.findings-emnlp.266.pdf
This academic paper introduces movement pruning, a novel method for reducing the size of large pre-trained language models like BERT during fine-tuning. Unlike traditional magnitude pruning which removes weights based on their absolute values, movement pruning prioritizes weights that change significantly during the fine-tuning process, demonstrating superior performance in high-sparsity scenarios. The authors provide mathematical foundations for their approach and empirically compare it against existing zeroth- and first-order pruning techniques, highlighting its effectiveness, especially when combined with distillation. The research emphasizes the potential for resource reduction, enabling the deployment of complex models on less powerful hardware and fostering broader accessibility in the field of natural language processing. Source: Published 2020 https://papers.neurips.cc/paper_files/paper/2020/file/eae15aabaa768ae4a5993a8a4f4fa6e4-Paper.pdf