This research presents a novel method for efficient long-context modeling in Large Language Models (LLMs) by tackling the quadratic complexity of attention mechanisms through KV cache compression. The core discovery is a fundamental local KV cache asymmetry, which reveals that adjacent attention keys exhibit high structural homogeneity, while their associated value vectors possess distinct, heterogeneous distributions. To capitalize on this finding, the authors propose AsymKV, a training-free compression framework that shifts information loss from heterogeneous values to homogeneous keys. AsymKV operates by applying homogeneity-based merging to keys using a mathematically derived optimal vector, paired with a lossless value representation scheme utilizing cardinality-aware normalization to preserve vital information. Extensive empirical results on benchmarks like LongBench, across diverse models such as LLaMA3.1-8B, confirm that AsymKV consistently surpasses state-of-the-art long-context methods in terms of accuracy and information retention, offering improved performance with practical inference efficiency. Source: https://arxiv.org/pdf/2506.05410

The research paper presents PageANN, a novel framework engineered to overcome the severe latency and scalability limitations facing existing disk-based Approximate Nearest Neighbor Search (ANNS) methods used in vector databases. Current systems suffer from inefficient search paths and a crucial misalignment between logical graph node size and the physical I/O granularity of Solid-State Drives (SSDs). PageANN introduces a core innovation: a page-node graph structure that directly maps logical graph nodes to physical SSD pages, significantly shortening I/O traversal paths and maximizing data utility during retrieval. This is supported by a co-designed disk data layout that embeds compressed neighbor vectors within each page and a dynamic memory management strategy utilizing lightweight indexing for fast query routing. According to experimental results, PageANN consistently outperforms state-of-the-art techniques, achieving substantial gains in throughput and latency across diverse datasets and memory constraints while maintaining comparable recall accuracy. Source: https://arxiv.org/pdf/2509.25487

Browse by month

Sep 2026 Aug 2026 Jul 2026 Jun 2026 May 2026 Apr 2026 Mar 2026 Feb 2026 Jan 2026 Dec 2025 Nov 2025 Oct 2025 Sep 2025 Aug 2025