This September 2024 paper introduces FusionANNS, a novel system designed to improve Approximate Nearest Neighbor Search (ANNS) for extremely large datasets. It addresses challenges in existing ANNS systems, such as performance bottlenecks, high operational costs, and accuracy limitations, particularly when dealing with billion-scale vector data in modern AI infrastructure like Large Language Models (LLMs). FusionANNS achieves this through a cooperative CPU/GPU architecture that employs multi-tiered indexing, heuristic re-ranking, and redundancy-aware I/O deduplication. The system is shown to significantly outperform state-of-the-art SSD-based and GPU-accelerated in-memory ANNS solutions in terms of throughput (QPS), cost efficiency, and memory efficiency, while maintaining low latency and high accuracy. Source: https://arxiv.org/pdf/2409.16576
This August 2025 paper introduces rStar2-Agent, a 14B math reasoning model developed by Microsoft Research that achieves state-of-the-art performance comparable to much larger models by employing agentic reinforcement learning. The model is trained to "think smarter" through three key innovations: an efficient RL infrastructure that manages high-throughput code execution, a novel GRPO-RoC algorithm for effective reasoning in a noisy code environment by filtering high-quality trajectories, and an efficient training recipe that minimizes computational cost. Demonstrating superior accuracy on challenging math benchmarks like AIME24/25, rStar2-Agent-14B also exhibits strong generalization to other domains like scientific reasoning and tool-use tasks, all while producing shorter, more concise responses than its counterparts. The paper further explores how high-entropy tokens in the model's reasoning traces indicate advanced cognitive behaviors such as self-reflection and adaptive exploration in response to tool feedback and errors. Source: https://arxiv.org/pdf/2508.20722
This August 2025 paper offers an extensive overview of the evolution and application of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) within scientific research, primarily focusing on the period from 2018 to 2025. It details how these AI models have progressed through various paradigm shifts, from initial transfer learning to sophisticated scientific agents capable of autonomous research. The document thoroughly examines the diverse data modalities—including visual spectra, microscopy images, molecular encodings, and time-series data—across six key scientific domains: Chemistry, Materials Science, Physics, Life Sciences, Astronomy, and Earth Science. Furthermore, it addresses critical issues surrounding data quality, traceability, timeliness, privacy, and bias within scientific datasets, while also highlighting the importance of robust evaluation benchmarks and tool integration for advancing scientific AI. Source: https://arxiv.org/pdf/2508.21148
This paper introduces Apertus, a large language model developed by the Swiss AI Initiative, a partnership between ETH Zurich and EPFL. The GitHub repository appears to host technical documentation or code related to Apertus, while the Hugging Face page provides a comprehensive overview of the Apertus-8B-Instruct-2509 model. This model is highlighted for being fully open, massively multilingual (supporting over 1800 languages), and compliant with data privacy regulations, even incorporating mechanisms for data protection and copyright requests. The Hugging Face page also outlines the model's technical specifications, how to use it, its evaluation, training details, and limitations.Sources:https://github.com/swiss-ai/apertus-tech-report/blob/main/Apertus_Tech_Report.pdfhttps://huggingface.co/swiss-ai/Apertus-8B-Instruct-2509
This May 2025 paper introduces FastVLM, an innovative approach designed to enhance the efficiency of Vision Language Models (VLMs). The authors explain that while increasing image resolution is crucial for VLM performance, traditional visual encoders become inefficient. FastVLM addresses this by incorporating FastViTHD, a novel hybrid vision encoder that reduces both the number of visual tokens and encoding time for high-resolution images. This optimization, achieved solely through input image scaling, leads to a significant 3.2x improvement in time-to-first-token (TTFT) while maintaining strong performance on VLM benchmarks, making it a more efficient solution compared to prior methods. The paper, submitted to CVPR 2025, also provides access to the code and models for its research.Sources:https://arxiv.org/abs/2412.13303https://machinelearning.apple.com/research/fast-vision-language-models
This September 2025 paper article from Nature, authored by Kevin M. Cherry and Lulu Qian, introduces a novel DNA-based neural network capable of supervised learning in vitro. The authors demonstrate how DNA molecules can be programmed to autonomously classify patterns from molecular examples. This system integrates training data directly into molecular memories and uses these memories for subsequent classification, moving beyond previous systems that relied on in silico learning. The work highlights the potential of molecular circuits to perform complex information processing, opening doors for adaptive decision-making in various physical systems, from biomedicine to soft materials. The article meticulously details the design, characterization, and scalability of their DNA neural network, showcasing its robustness and outlining future challenges like unsupervised learning and increased complexity through spatial organization. Source: https://www.nature.com/articles/s41586-025-09479-w
This June 2024 paper examines the current state and future potential of Physical Neural Networks (PNNs), which are AI systems implemented directly in physical hardware rather than purely digital software. It explores various training methodologies for PNNs, including in-silico (digital simulation), in-situ (real-world hardware training), and hybrid approaches like physics-aware training, each with its own advantages and limitations regarding accuracy, speed, cost, and complexity. The text also discusses alternative training paradigms such as Feedback Alignment, Local Learning, and gradient-free methods that aim to overcome challenges associated with traditional backpropagation in physical systems. Furthermore, it highlights the promise of PNNs for advanced applications like continual learning and the development of energy-efficient large AI models, identifying emerging technologies such as quantum and photonic hardware as key to their future scalability and performance benefits over conventional digital systems. Source: https://arxiv.org/pdf/2406.03372
This September 2025 paper introduces DeepResearch Arena, a novel benchmark designed to evaluate the research capabilities of large language models (LLMs) by mirroring real-world academic inquiry. This benchmark addresses limitations of existing evaluation methods, which often suffer from data leakage or lack authenticity, by grounding its tasks in academic seminars and expert discourse. A Multi-Agent Hierarchical Task Generation (MAHTG) system is utilized to automatically generate over 10,000 diverse research tasks across multiple disciplines, covering phases from synthesis to evaluation. The paper also proposes a hybrid evaluation framework that combines Keypoint-Aligned Evaluation (KAE) for factual correctness and Adaptively-generated Checklist Evaluation (ACE) for nuanced, open-ended reasoning. Experimental results demonstrate the challenges DeepResearch Arena poses to current state-of-the-art LLMs, revealing varying strengths and limitations across models. Source: https://arxiv.org/pdf/2509.01396
This document announces EmbeddingGemma, a new open embedding model from Google, specifically designed for on-device artificial intelligence (AI). It highlights the model's efficiency, compact size, and best-in-class performance for its category, particularly in multilingual text embedding. The source explains how EmbeddingGemma enables mobile-first Retrieval Augmented Generation (RAG) pipelines and semantic search by generating high-quality text embeddings directly on user hardware, ensuring privacy and offline functionality. It also details the model's compatibility with popular development tools and its ability to offer flexible output dimensions while maintaining a small memory footprint. Finally, it contrasts EmbeddingGemma's strengths for on-device applications with other Google models suited for large-scale server-side use. Source: https://developers.googleblog.com/en/introducing-embeddinggemma/
This September 2025 paper introduces Inverse IFEval, a novel benchmark designed to evaluate Large Language Models (LLMs) for their Counter-intuitive Ability. This refers to an LLM's capacity to override its ingrained training patterns and comply with instructions that conflict with conventional norms or standardized formats. The benchmark includes eight distinct categories of such challenging instructions, like "Code without Comments" or "Deliberately Incorrect Answers," to expose the cognitive inertia and overfitting that current LLMs exhibit. The study underscores the need for future LLM development to prioritize adaptability in unconventional contexts beyond mere fluency and factual accuracy. Findings demonstrate that while some models perform well on traditional instruction-following tasks, their performance significantly declines when faced with these inverse instructions. Source: https://arxiv.org/pdf/2509.04292
These academic papers introduce and detail the Massive Multilingual Text Embedding Benchmark (MMTEB), a comprehensive evaluation framework for text embedding models. The MMTEB expands upon existing benchmarks by offering over 500 tasks across 250+ languages and various domains, significantly increasing the diversity and scale of evaluation. It incorporates optimizations like downsampling and caching to reduce computational costs, making the benchmark more accessible, especially for low-resource languages. The papers also evaluate various models, including large language models (LLMs) and smaller, multilingual models, revealing that instruction-tuned models often perform better, and smaller models can surprisingly outperform larger LLMs in highly multilingual or low-resource settings. Ultimately, the MMTEB aims to provide a robust and extensive platform for assessing and advancing text embedding capabilities across a wide spectrum of linguistic and thematic challenges.Sources:June 2025: MMTEB: MASSIVE MULTILINGUAL TEXT EMBEDDING BENCHMARKhttps://arxiv.org/pdf/2502.13595March 2023: MTEB: Massive Text Embedding Benchmarkhttps://arxiv.org/pdf/2210.07316https://github.com/embeddings-benchmark/mteb
This August 2025 paper from Google DeepMind, titled "On the Theoretical Limitations of Embedding-Based Retrieval," explores the fundamental constraints of vector embedding models in information retrieval. The authors demonstrate that the number of relevant document combinations an embedding can represent is inherently limited by its dimension. Through empirical "free embedding" experiments and the introduction of a new dataset called LIMIT, they show that even state-of-the-art models struggle with simple queries designed to stress these theoretical boundaries. The research concludes that for complex, instruction-following queries, alternative retrieval approaches like cross-encoders or multi-vector models may be necessary to overcome these inherent limitations. Source: https://arxiv.org/pdf/2508.21038
This September 2025 paper describe SAIR, the Structurally Augmented IC50 Repository, a groundbreaking open-source dataset developed by SandboxAQ in collaboration with NVIDIA. SAIR is the largest publicly available collection of over 5 million AI-generated 3D protein-ligand structures, each linked with experimentally measured drug potency data (IC₅₀ values). This dataset aims to bridge a critical data gap in AI-powered drug discovery by providing comprehensive structural intelligence, thereby enabling researchers to accelerate R&D, explore novel drug targets, and improve the accuracy of AI models for predicting drug properties. The creation of SAIR involved extensive high-performance computing, taking over 130,000 GPU hours, and its structures were rigorously validated with industry-standard tools, achieving a 97% pass rate. By offering this resource for free commercial and non-commercial use on platforms like Hugging Face, SAIR seeks to revolutionize how pharmaceutical, biotech, and tech-bio leaders approach drug design and optimization.Sources:https://go.sandboxaq.com/rs/175-UKR-711/images/sair_paper.pdfhttps://huggingface.co/datasets/SandboxAQ/SAIRhttps://huggingface.co/blog/SandboxAQ/sair-data-accelerating-drug-discovery-with-ai
This blog post series from Chris Lattner extensively examines CUDA's pervasive dominance in AI compute, detailing its evolution from a graphics processor to a layered software platform integral to NVIDIA's success, while also highlighting the challenges and complexities it presents to developers and alternative hardware vendors. The articles critically assess various attempts to democratize AI compute, including OpenCL, TVM, XLA, and MLIR, explaining why these alternatives largely failed to dislodge CUDA due to fragmentation, misaligned incentives, and a lack of unified vision. Ultimately, the texts introduce Modular's approach to addressing these issues through its Mojo language, MAX framework, and Mammoth cluster management system, aiming to provide a portable, performant, and programmable solution for the rapidly evolving Generative AI landscape. Source: https://www.modular.com/blog/democratizing-compute-part-1-deepseeks-impact-on-ai
These sources collectively explore various approaches to evaluating and improving Large Language Models (LLMs). Several papers introduce new benchmark datasets designed to test LLMs on complex reasoning tasks, such as the "BIG-Bench Hard (BBH)" suite, the graduate-level "GPQA" questions in science, and "MuSR" for multistep soft reasoning in natural language narratives. A key technique discussed across these sources is Chain-of-Thought (CoT) prompting, which encourages LLMs to show their step-by-step reasoning, leading to improved performance, often surpassing human-rater averages on challenging tasks. Additionally, the "Instruction-Following Eval (IFEval)" introduces a reproducible benchmark for verifiable instructions, allowing for objective assessment of an LLM's ability to follow explicit directives. The "MMLU-Pro Benchmark" further contributes a large-scale dataset across diverse disciplines to rigorously assess model capabilities, emphasizing the need for robust evaluation metrics and challenging data to push the boundaries of AI reasoning.Sources:https://github.com/EleutherAI/lm-evaluation-harnesshttps://github.com/EleutherAI/lm-evaluation-harness/blob/main/lm_eval/tasks/leaderboard/README.mdhttps://arxiv.org/pdf/2103.03874 - Measuring Mathematical Problem Solving With theMATH Datasethttps://arxiv.org/pdf/2210.09261 - Challenging BIG-Bench tasks andwhether chain-of-thought can solve themhttps://arxiv.org/pdf/2310.16049 - MUSR: TESTING THE LIMITS OF CHAIN-OF-THOUGHTWITH MULTISTEP SOFT REASONINGhttps://arxiv.org/pdf/2311.07911 - Instruction-Following Evaluation for Large LanguageModelshttps://arxiv.org/pdf/2311.12022 - GPQA: A Graduate-Level Google-ProofQ&A Benchmarkhttps://arxiv.org/pdf/2406.01574 - MMLU-Pro: A More Robust and ChallengingMulti-Task Language Understanding Benchmark
This July 2021 paper documents the development and evaluation of OpenAI's Codex models, which are large language models specialized in code generation, particularly Python functions from docstrings. They introduce HumanEval, a hand-written dataset designed to assess the functional correctness of generated code through unit tests, a more robust metric than traditional match-based scores like BLEU. The papers compare the performance of various Codex iterations, including supervised fine-tuned versions (Codex-S), against other models like GPT-3, demonstrating significant improvements in pass rates with increased model size and sample generation. Furthermore, the texts explore the limitations, broader impacts, and potential hazards of these models, discussing issues such as over-reliance, misalignment, economic implications for the labor market, and security concerns related to generating vulnerable or biased code. Finally, the sources touch upon Codex-D, a model for generating docstrings from code, and emphasize the need for continued research into safe and responsible AI deployment.Sources:https://arxiv.org/pdf/2107.03374https://github.com/openai/human-eval
These September 2025 posts describe HuggingFaceM4/FineVision, a large dataset designed for image and text modalities. It features a substantial size, ranging from 10M to 100M, and is available in the parquet format. This dataset includes various ratings, such as relevance, visual dependency, image correspondence, and formatting, indicating its use in evaluating the quality and relationship between visual and textual content. The examples provided demonstrate that FineVision contains question-and-answer pairs related to diverse charts and diagrams, covering topics like population trends, genetic diseases, software update frequencies, and demographic distributions, suggesting its application in training models for visual question answering and chart comprehension.Sources:https://huggingface.co/spaces/HuggingFaceM4/FineVisionhttps://huggingface.co/datasets/HuggingFaceM4/FineVision
Thus describes EleutherAI's GPT-NeoX library, a robust open-source framework for training large-scale autoregressive language models on GPUs, building upon the Megatron and DeepSpeed libraries. It highlights the library's advanced features like distributed training, support for various hardware and systems, and cutting-edge architectural innovations. The text also provides practical guidance on setup, configuration, data preparation, training, inference, and evaluation, alongside details on pretrained models like GPT-NeoX-20B and Pythia. Furthermore, it details how to export models to Hugging Face and monitor experiments, underscoring its widespread adoption in research and industry. Source: https://github.com/EleutherAI/gpt-neox
The provided May 2024 sources center around CoreNet, an Apple-developed library for training deep neural networks, and OpenELM, an efficient language model family built using CoreNet. CoreNet is a versatile toolkit supporting various tasks, including foundation models like large language models (LLMs), object classification, and semantic segmentation, with its development evolving from the earlier CVNets. A key innovation highlighted is OpenELM's layer-wise scaling strategy, which optimizes parameter allocation within transformer models to achieve superior accuracy with fewer pre-training tokens compared to other open LLMs. The resources emphasize reproducibility and transparency by providing comprehensive frameworks for OpenELM's training and evaluation, including code for inference and fine-tuning on Apple devices using the MLX library, and detailed benchmarks on both NVIDIA CUDA and Apple Silicon hardware.Sources:https://arxiv.org/pdf/2404.14619https://machinelearning.apple.com/research/openelmhttps://github.com/apple/corenethttps://github.com/apple/corenet/tree/main/projects/kv-prediction
This June 2024 paper introduces SGLang, a framework designed to enhance the efficiency of Large Language Model (LLM) and Vision Language Model (VLM) serving. It achieves this through a co-design of a flexible frontend language and a fast backend runtime. The frontend simplifies programming with primitives for generation and parallelism, while the backend utilizes novel optimizations like RadixAttention for KV cache reuse and compressed finite state machines for faster structured output decoding. These innovations allow SGLang to significantly improve throughput and reduce latency compared to existing systems across various LLM applications and hardware platforms. The framework is open-source, boasts extensive model support, and has seen wide industry adoption due to its performance benefits in complex LM programs.Sources:https://arxiv.org/pdf/2312.07104https://docs.sglang.ai/https://github.com/sgl-project/sglang
This March 2025 paper examines the input/output (I/O) characteristics of offloading large language model (LLM) components to NVMe SSDs during inference, a critical solution for overcoming GPU memory limitations with ever-growing LLMs. Researchers analyzed block-layer I/O traces from two prominent LLM frameworks, DeepSpeed and FlexGen, to understand how model weights and key-value (KV) caches are handled. The findings indicate that asynchronous I/O using libaio significantly outperforms POSIX for tensor transfers, although neither method fully saturates the NVMe SSD's theoretical bandwidth. For model offloading, I/O is predominantly characterized by 128KiB reads, primarily occurring at the beginning of the inference process, while KV cache offloading involves both reads and writes of similar size, with read bandwidth being substantially higher. Ultimately, the research suggests that modern NVMe SSDs are capable of supporting current LLM inference workloads but highlights opportunities for further optimization in SSD design and KV cache management. Source: https://dl.acm.org/doi/10.1145/3719330.3721230
This September 2025 paper introduces "Behavioral Fingerprinting," a novel framework designed to evaluate Large Language Models (LLMs) beyond traditional performance scores like MMLU. It aims to understand how models "think," creating a multi-faceted profile of their intrinsic cognitive and interactive styles. The methodology employs a diagnostic prompt suite and an automated evaluation pipeline where a powerful LLM acts as a judge, analyzing eighteen different models across four key dimensions: internal world model, reasoning abilities, biases and personality (including sycophancy), and semantic robustness. Findings indicate a convergence in core reasoning abilities among top models but a significant divergence in alignment-related behaviors such as sycophancy and robustness, which are influenced by specific developer strategies. The framework also identifies default personality clustering (ISTJ/ESTJ types), reflecting common training paradigms that reward logical, structured, and decisive responses. Source: https://arxiv.org/pdf/2509.04504
This September 2025 paper investigates the reliability and robustness of Large Language Models (LLMs) when evaluated using traditional benchmarks. The authors systematically paraphrased questions across six common benchmarks and observed how 34 different LLMs performed. Their findings indicate that while LLM rankings remain relatively consistent, their absolute effectiveness scores significantly decline when faced with reworded questions, suggesting a lack of robustness to linguistic variability. The study highlights that current benchmark evaluations may overstate LLM generalization abilities and advocates for more robustness-aware evaluation methodologies that better reflect real-world language use. Source: https://arxiv.org/pdf/2509.04013
This September 2025 paper introduces TraceRL, a novel reinforcement learning framework designed to enhance diffusion language models (DLMs) across various architectural types. The core idea behind TraceRL is to align the training process with the preferred inference trajectories of the model, which demonstrably improves performance on complex reasoning tasks like mathematics and coding. The authors also propose a diffusion-based value model to boost training stability. Through experiments, the paper showcases the effectiveness of TraceRL, yielding state-of-the-art DLMs called TraDo that outperform larger autoregressive models. Furthermore, the source provides an open-source framework to facilitate the development, training, and deployment of these advanced DLMs, including accelerated inference techniques and diverse post-training methods. Source: https://arxiv.org/pdf/2509.06949
The May - June 2025 sources introduce AlphaEvolve, a novel AI coding agent developed by Google DeepMind in collaboration with mathematicians like Javier Gómez Serrano and Terence Tao. This Gemini-powered tool utilizes an evolutionary process, similar to natural selection, to generate and iteratively refine code solutions for complex problems. AlphaEvolve has demonstrated its capability in scientific and algorithmic discovery, successfully tackling open mathematical challenges such as improving bounds for matrix multiplication and the kissing number problem in 11 dimensions. Beyond theoretical advancements, it has also been applied to optimize critical components within Google's computing infrastructure, including data center scheduling, Gemini kernel engineering, and hardware circuit design. The sources highlight AlphaEvolve's potential to accelerate research by combining human intuition with machine exploration, marking a significant shift in how complex problems are approached.Sources:https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/AlphaEvolve.pdfhttps://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/https://www.crm.cat/javier-gomez-serrano-collaborates-with-terence-tao-and-deepmind-on-an-ai-project-to-solve-open-mathematical-problems/
This July 2002 paper introduced BLEU (Bilingual Evaluation Understudy), an automatic and inexpensive method for evaluating machine translation (MT) quality. It highlights the limitations of human evaluation, such as its high cost and time consumption, and proposes BLEU as a quick, language-independent alternative that correlates strongly with human judgment. The core concept of BLEU involves measuring the "closeness" of a machine translation to one or more human reference translations through a modified n-gram precision metric and a brevity penalty. The paper details the mathematical formulation of the BLEU score, its evaluation against human and machine translations, and its proven correlation with human assessment across various languages. Source: https://dl.acm.org/doi/10.3115/1073083.1073135
This September 2025 paper introduces Diversity-Aware Reinforcement Learning (Darling), a novel framework designed to enhance both the quality and semantic diversity of large language model (LLM) generations. Recognizing that traditional post-training methods often sacrifice diversity for accuracy, Darling integrates a learned partition function to measure semantic diversity beyond simple lexical variations. This diversity signal is then multiplied with a quality reward during online reinforcement learning, which encourages LLMs to produce responses that are not only high-quality but also distinct and novel. Experiments on both non-verifiable tasks, such as creative writing, and verifiable tasks, like competition math, demonstrate that Darling consistently outperforms quality-only baselines, achieving improved scores in both quality metrics and diversity measures (e.g., pass@1 and pass@k). A key finding is that explicitly optimizing for diversity can catalyze exploration in online RL, leading to a simultaneous improvement in the overall quality of the generated responses. Source: https://arxiv.org/pdf/2509.02534
This February 2025 paper introduces INF2, a novel framework designed to enhance the generative inference throughput of large language models (LLMs) by utilizing computational storage devices (CSDs). The core innovation, attention-near storage (ANS), offloads memory-intensive self-attention operations directly to accelerators within these storage devices, significantly reducing data transfer bottlenecks over the system interconnect. To further boost performance, INF2 incorporates delayed KV cache writeback which minimizes storage write latency by batching updates to the KV cache, and cooperative X-cache, which optimizes host memory usage by storing input activations instead of key-value caches for cooperative processing between the GPU and CSDs. Through these methods, INF2 demonstrates substantial throughput improvements, achieving up to 3.46 times faster performance compared to existing state-of-the-art baselines in real-world evaluations. Source: https://arxiv.org/html/2502.09921v1
The September 9 2025 press release and paper announce and detail K2 Think, an advanced open-source AI reasoning system developed by the Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) and G42 in the UAE. K2 Think stands out for its parameter efficiency, achieving performance comparable to much larger models, particularly in mathematical reasoning, with only 32 billion parameters. This breakthrough is attributed to a six-pillar approach, including supervised fine-tuning, reinforcement learning with verifiable rewards, agentic planning, test-time scaling, and optimization for Cerebras Wafer-Scale Engine hardware for rapid inference. The system also undergoes red-teaming for safety evaluation, demonstrating a solid baseline in content refusal and conversational robustness, while identifying areas for improvement in cybersecurity and jailbreak resistance.Sources:https://k2think-about.pages.dev/assets/tech-report/K2-Think_Tech-Report.pdfhttps://k2think-about.pages.dev/assets/announcement/K2_Think_-_Sep_9_English_Release.pdfhttps://www.k2think.ai/k2think
This September 2025 paper analyzes the theoretical benefits and limitations of Masked Diffusion Models (MDMs) for text generation, contrasting them with auto-regressive models. While MDMs can sample multiple tokens in parallel, offering a potential for efficiency, the research demonstrates that their actual performance depends heavily on the evaluation metric. Specifically, MDMs can achieve near-optimal fluency (low Token Error Rate) with a constant number of sampling steps, regardless of sequence length. However, when assessed for correctness (low Sequence Error Rate), particularly for tasks requiring logical reasoning, MDMs necessitate a number of sampling steps that scales linearly with sequence length, effectively negating their efficiency advantage. Empirical results using formal languages and large open-sourced MDMs support these theoretical findings, indicating that MDMs are better suited for fluent text generation but less so for accuracy-critical reasoning tasks. Source: https://arxiv.org/pdf/2502.09622
This September 2025 paper introduces Mini-o3, a Vision-Language Model (VLM) designed to overcome the limitations of existing VLMs in handling complex visual search tasks that require multi-turn reasoning and trial-and-error exploration. The researchers developed a three-component training recipe, including the creation of the Visual Probe Dataset with challenging, high-resolution images, a pipeline for synthesizing diverse multi-turn trajectories for supervised finetuning, and an over-turn masking technique in reinforcement learning. This masking prevents penalization of long, incomplete reasoning paths, encouraging deeper exploration without increasing training time. Mini-o3 demonstrates state-of-the-art performance on various visual search benchmarks, showcasing its enhanced ability for complex, adaptive visual understanding through iterative observation, thought, and action. Source: https://arxiv.org/pdf/2509.07969
This July 2025 paper introduces AdLlama, a new large language model (LLM) for generating Facebook ad text, trained using Reinforcement Learning with Performance Feedback (RLPF). Unlike previous models that relied on supervised fine-tuning to imitate curated ads, AdLlama utilizes historical ad performance data, specifically click-through rates (CTR), as a reward signal to optimize its text generation. A large-scale A/B test on Facebook, involving nearly 35,000 advertisers, demonstrated that AdLlama significantly improved advertiser-level CTR by 6.7% and increased the number of ad variations advertisers created by 18.5%. The findings highlight RLPF as a promising, generalizable approach for metric-driven LLM post-training, moving beyond subjective human preferences to achieve tangible business outcomes in real-world settings like online advertising. Source: https://arxiv.org/pdf/2507.21983
This July 2024 paper introduces ByteCheckpoint, a novel PyTorch-native system designed for Large Language Model (LLM) development. This system addresses critical challenges in LLM training, particularly the high I/O costs associated with saving and loading checkpoints, and the complexities of checkpoint resharding across different parallel configurations and training frameworks. ByteCheckpoint achieves this through a data/metadata disaggregated storage architecture and asynchronous tensor merging, enabling automatic online resharding and multi-framework support. The paper highlights ByteCheckpoint's significant performance improvements in reducing checkpoint saving and loading times compared to existing methods, making LLM development more efficient and robust. The authors detail various I/O performance optimization techniques. Source: https://arxiv.org/html/2407.20143v1
This April 2025 paper introduces SODA, a novel framework designed to enhance digital advertising strategies by making opaque AI systems more understandable for marketers. The authors highlight the current challenges faced by advertisers due to the lack of transparency in major ad platforms like Meta, which often results in wasted ad spend and reliance on intuition. To address this, SODA integrates Large Language Models (LLMs) with explainable AI techniques to provide clear, actionable insights into ad performance. The framework initially employs an improved Click-Through Rate (CTR) prediction model, SoWide-v2, which also offers visual explanations through attention maps. Furthermore, SODA leverages LLMs to automate comprehensive ad analysis, including identifying target audiences, brand personas, and key messaging, providing marketers with summarized and comparative data that was previously difficult to obtain without extensive manual effort. A case study with marketing professionals validated the framework's practical value and potential to streamline decision-making in fast-paced advertising environments. Source: https://arxiv.org/pdf/2504.20064
This April 2025 paper introduces HyperController, a novel and computationally efficient algorithm designed to optimize hyperparameters during the training of reinforcement learning neural networks. Hyperparameter optimization is crucial for improving machine learning models, but traditional methods can be slow and computationally intensive. HyperController addresses these challenges by modeling the hyperparameter optimization problem as an unknown Linear Gaussian Dynamical System and leveraging the Kalman filter for efficient prediction. The algorithm is validated through experiments on various OpenAI Gymnasium environments, where it demonstrates faster training times and superior or comparable performance compared to existing methods, achieving the highest median reward in four out of five environments. The research highlights HyperController's potential for stable and efficient training of reinforcement learning neural networks, offering advantages like quicker deployment and easier fine-tuning. Source: https://arxiv.org/pdf/2504.19382
This August 2025 paper introduces NOVELTYBENCH, a new benchmark designed to evaluate how well large language models (LLMs) generate diverse and high-quality outputs, addressing the problem of "mode collapse" where models produce repetitive responses. The research found that current state-of-the-art LLMs consistently generate less diversity than human writers, with larger models often exhibiting even lower diversity than their smaller counterparts. The benchmark uses a unique approach to measure functional equivalence between generations, ensuring that diversity is meaningful to users. While certain prompting strategies, like in-context regeneration, can enhance diversity, the study suggests that this capability is not inherent in the models themselves, highlighting the need for new training and evaluation paradigms that prioritize both diversity and quality in LLM development. Source: https://arxiv.org/pdf/2504.05228
This September 10, 2025 technical report from Tencent AI Lab introduces Parallel-R1, a novel reinforcement learning (RL) framework designed to enhance large language models (LLMs) with parallel thinking capabilities for complex mathematical reasoning tasks. Unlike previous methods relying on supervised fine-tuning (SFT) over synthetic data, Parallel-R1 utilizes a progressive curriculum to address the cold-start problem in RL, initially using SFT on simpler tasks to instill the basic format of parallel thinking before transitioning to RL for exploration and generalization on more challenging problems. The research highlights that parallel thinking evolves from an exploratory strategy to a multi-perspective verification tool during training, and can also serve as a mid-training exploration scaffold to unlock higher performance ceilings. The framework demonstrated significant accuracy improvements across various math benchmarks, and the authors plan to open-source their model, data, and code. Source: https://arxiv.org/pdf/2509.07980
This September 2025 published research investigates how large language models (LLMs) perform mental math, particularly focusing on the flow of information and computational processes within their transformer architecture. The authors introduce two novel techniques, Context-Aware Mean Ablation (CAMA) and Attention-Based Peeking (ABP), to identify a minimal computational subgraph called All-for-One (AF1). This subgraph reveals that for mental math tasks, input-specific computation is largely deferred to later layers and primarily handled by the final token, which receives necessary information from other tokens during a few specific intermediate layers. The study demonstrates that this sparse AF1 subgraph is sufficient and necessary for high performance across various arithmetic expressions and models, offering significant insights into the mechanistic interpretability of LLM arithmetic reasoning. Source: https://www.arxiv.org/pdf/2509.09650
This October 2024 paper introduces synthetic continued pretraining (synthetic CPT), a novel method designed to enhance language model knowledge acquisition from small, specialized text collections. Current large language models often struggle with data efficiency and learning niche facts from limited sources. The core of this approach is EntiGraph, a synthetic data augmentation algorithm that extracts entities and their relationships from a small corpus to generate a much larger, more diverse synthetic dataset. Experiments using the QuALITY dataset demonstrate that EntiGraph CPT significantly improves a model's ability to answer questions and summarize content about these specialized domains, outperforming direct training on raw data or simple rephrasing techniques. The authors also explore the scaling properties of EntiGraph, revealing distinct phases of knowledge acquisition. Source: https://arxiv.org/pdf/2409.07431
These September 2025 papers present a technical report on SpikingBrain, a novel family of large language models (LLMs) that draw inspiration from brain mechanisms to address the efficiency challenges of traditional Transformer architectures. The research focuses on efficient long-context training and inference by developing hybrid linear attention architectures and an adaptive threshold spiking neuron scheme. A significant aspect of this work is the successful training and deployment of these models on non-NVIDIA GPU clusters, specifically the MetaX platform, demonstrating the feasibility of large-scale LLM development on alternative hardware. The authors highlight substantial speedups in inference for long sequences and significant reductions in energy consumption through the sparse, event-driven spiking design, with performance comparable to established Transformer baselines while using considerably less training data. This work ultimately aims to advance the design of energy-efficient and scalable brain-inspired LLMs for next-generation computing systems, including neuromorphic hardware.Sources:https://arxiv.org/pdf/2509.05276https://arxiv.org/html/2509.05276v1
This September 2025 paper explores the critical role of statistical methods in enhancing the reliability and functionality of Generative AI (GenAI), which inherently lacks guarantees regarding correctness or safety. It discusses various statistical applications, including improving and altering model behavior through techniques like output trimming and abstention based on risk scores, often utilizing conformal prediction for provable guarantees. The text also covers diagnostics and uncertainty quantification (UQ), differentiating between epistemic and aleatoric uncertainty and addressing challenges like semantic multiplicity and the need for calibration in GenAI outputs. Furthermore, it highlights the importance of statistical inference in evaluating GenAI models, particularly with limited data, and examines interventions and experiment design, such as using "steering vectors" and "causal mediation analysis" to understand and mitigate biases within these complex systems. Source: https://arxiv.org/pdf/2509.07054
This September 2025 paper provides a comprehensive overview of Reinforcement Learning (RL) as applied to Large Reasoning Models (LRMs). It breaks down the field into foundational components such as reward design and policy optimization, explaining various algorithms like PPO and GRPO. The document also discusses training resources, distinguishing between static corpora and dynamic environments, and highlights diverse applications of RL in LRMs, including coding, agentic tasks, and multimodal understanding, with a focus on models from 2025. Ultimately, the paper aims to identify future directions for scaling RL in LRMs towards achieving Artificial Superintelligence (ASI). Source: https://arxiv.org/pdf/2509.08827
This May 2025 paper explores the structural patterns of knowledge within Large Language Models (LLMs) by adopting a graph-based perspective. The authors quantify LLM knowledge at both the triplet and entity levels, analyzing its relationship with graph properties like node degree. Key findings include the discovery of knowledge homophily, where closely connected entities exhibit similar knowledgeability, and a positive correlation between an entity's degree and its knowledge. These insights further motivate the development of graph machine learning models to predict entity knowledge, which can then be used to strategically select less-known triplets for fine-tuning LLMs, leading to improved performance. The study evaluates several prominent LLMs across diverse knowledge graphs, highlighting domain-specific variations in knowledge distribution and the consistent presence of structural patterns. Source: https://arxiv.org/html/2505.19286v2
On this November 2025 paper the Meta Llama Team's paper introduces Llama 3, a new family of large language models featuring 8B, 70B, and 405B parameters, designed with native multilingual support, coding, reasoning, and tool usage capabilities. The development emphasizes data quality and diversity, employing extensive filtering, de-duplication, and heuristic cleaning processes for both English and multilingual data, alongside scaling laws to optimize model size and training budgets. The models utilize a standard dense Transformer architecture with minor adaptations like grouped query attention and an attention mask for multi-document sequences, demonstrating comparable performance to leading models such as GPT-4 across various benchmarks. Furthermore, the research explores integrating multimodal capabilities—image, video, and speech—through compositional approaches involving specialized encoders and adapters, which are trained through multi-stage pre-training and fine-tuning. A significant focus is also placed on safety and responsible development, incorporating comprehensive data cleaning, iterative safety finetuning with reward models and DPO, and robust red teaming efforts to address risks like insecure coding and prompt injection, while publicly releasing Llama Guard 3 as a system-level safety classifier. Source: https://arxiv.org/pdf/2407.21783
This July 2024 paper introduces Activation-aware Weight Quantization (AWQ), a novel method for compressing Large Language Models (LLMs) by quantizing weights to low-bit integers for efficient deployment on edge devices. It highlights that AWQ identifies and protects crucial "salient" weights by observing activation distributions, which significantly reduces quantization error without requiring computationally intensive training or overfitting to specific datasets. Complementing AWQ, the paper also presents TinyChat, an inference framework specifically designed to optimize and accelerate these 4-bit quantized LLMs on various hardware, including mobile GPUs and even resource-constrained devices like the Raspberry Pi, achieving substantial speedups compared to traditional implementations. The combination of AWQ and TinyChat aims to make powerful LLMs accessible for on-device applications, addressing challenges like memory limitations and power consumption. Source: https://arxiv.org/pdf/2306.00978
This June 2023 paper introduces FlexGen, a novel high-throughput generation engine designed to overcome the substantial computational and memory demands of large language model (LLM) inference on limited hardware, specifically a single commodity GPU. It details FlexGen's ability to aggregate memory and computation across the GPU, CPU, and disk, employing an optimized scheduling approach and a linear programming-based policy search to store and access tensors efficiently. Furthermore, FlexGen incorporates 4-bit compression for model weights and attention caches, which significantly reduces memory footprint with minimal accuracy loss. The research demonstrates FlexGen's superior performance, achieving substantially higher throughput compared to existing offloading systems, even enabling the operation of models as large as OPT-175B on a single 16GB GPU. Source: https://arxiv.org/pdf/2303.06865
This September 2018 paper introduces GraphSAGE, a novel inductive framework designed to generate node embeddings for large, evolving graphs, addressing limitations of prior transductive methods that struggle with unseen data. Instead of learning a specific embedding for each node, GraphSAGE learns a function that generates these embeddings by sampling and aggregating features from a node's local neighborhood. The authors evaluate various aggregator architectures, including mean, LSTM, and pooling functions, demonstrating that GraphSAGE significantly outperforms strong baselines on node classification tasks across diverse datasets, such as citation networks, Reddit posts, and protein-protein interaction graphs. The research also highlights GraphSAGE's computational efficiency and provides a theoretical analysis of its capability to learn local graph structural information, like clustering coefficients. Source: https://arxiv.org/pdf/1706.02216
This January 2025 paper introduces HybridServe, an LLM inference system designed to enhance throughput and cost-effectiveness for large language models by optimizing memory usage and host-GPU communication. It tackles the challenges of host memory offloading, where model parameters and KV cache are stored on slower host memory to reduce costs but can lead to GPU underutilization due to limited transfer bandwidth. HybridServe proposes a novel activation checkpointing technique with a KV-Activation hybrid caching scheme that stores intermediate activations, allowing for faster recomputation of the KV cache while model parameters are transferred. This system dynamically balances communication overhead and recomputation time to maximize throughput, demonstrating significant improvements over existing state-of-the-art methods like FlexGen. Source: https://arxiv.org/pdf/2501.01792
This August 2025 paper explores the critical area of fact-checking and factuality evaluation in Large Language Models (LLMs). It systematically analyzes the challenges of misinformation generation, particularly hallucinations, which are factually incorrect but fluent outputs from LLMs. The paper investigates various mitigation strategies, including fine-tuning, instruction tuning, and Retrieval-Augmented Generation (RAG), which grounds LLM outputs in external knowledge. It further examines evaluation metrics, datasets, and prompting strategies used to assess and enhance the factual accuracy of these models, highlighting the need for more robust, explainable, and domain-specific fact-checking frameworks. The review concludes by identifying open issues and future research agendas to foster more trustworthy and context-aware LLMs. Source: https://arxiv.org/pdf/2508.03860
This September 2025 paper presents MetaGraph, a novel methodology for constructing knowledge graphs from scientific literature, specifically applied to Financial Natural Language Processing (NLP) research between 2022 and 2025. The authors utilized Large Language Models (LLMs) to extract key information from 681 papers, including tasks, datasets, models, motivations, and limitations, and organized it into a structured, queryable format. The analysis highlights three phases in Financial NLP's evolution: initial LLM adoption and task/dataset innovation, subsequent critical reflection on LLM limitations, and a current trend toward integrating peripheral techniques into modular systems. The research reveals a shift toward Financial Question Answering (QA), increased use of synthetic data, and a growing emphasis on open-source models and system-level solutions like Retrieval-Augmented Generation (RAG). Ultimately, MetaGraph offers a reusable framework for quantitatively mapping scientific progress and understanding changing priorities in a rapidly evolving field. Source: https://www.arxiv.org/pdf/2509.09544
This September 2023 paper introduces PyTorch Fully Sharded Data Parallel (FSDP), an advanced solution designed to scale the training of exceptionally large machine learning models. It addresses limitations of previous methods like Distributed Data Parallel (DDP) by sharding model parameters, gradients, and optimizer states across multiple GPUs, thereby drastically reducing individual GPU memory consumption. FSDP employs various techniques, including deferred initialization, flexible sharding strategies, and optimizations for communication overlap and prefetching, to ensure high efficiency and a user-friendly experience. The research demonstrates FSDP's effectiveness in training models with billions of parameters, achieving near-linear scalability and enabling broader access to large model development. The authors also discuss interoperability with other parallelization paradigms and known limitations. Source: https://arxiv.org/pdf/2304.11277
This September 2025 paper explores the concept of long-horizon execution in Large Language Models (LLMs), arguing that marginal gains in single-step accuracy can lead to exponential improvements in the length of tasks LLMs can complete. The authors introduce a novel framework to isolate execution capabilities by providing models with necessary knowledge and plans, revealing that larger models can execute significantly more steps, even when smaller models achieve perfect single-turn accuracy. A key finding is the "self-conditioning effect," where LLMs become more prone to errors when their past mistakes are present in the context, a challenge not fully mitigated by increasing model size. However, the paper concludes that "thinking" models, which employ sequential test-time compute, effectively address this self-conditioning and can execute substantially longer tasks in a single turn. Source: https://arxiv.org/pdf/2509.09677
This May 2025 paper introduces a resource-aware algorithm designed to optimize the performance of Large Language Models (LLMs) for low-latency inference on edge computing devices. The core innovation lies in its fine-grained partitioning of the Transformer architecture, specifically at the attention head-level, rather than coarser layer-level divisions. This approach allows for dynamic reassignment and migration of these individual attention heads and their associated Key/Value (K/V) caches across heterogeneous edge devices. By managing the expanding memory footprint of K/V caches and exploiting parallel execution of attention heads, the proposed method significantly reduces inference latency and memory usage compared to existing static or layer-based partitioning strategies. Source: https://arxiv.org/pdf/2505.02533
This February 2025 paper introduce CodeI/O, a novel training method for Large Language Models (LLMs) that enhances general reasoning abilities by transforming code into an input-output prediction task. Instead of focusing on generating code, CodeI/O trains models to predict inputs or outputs of a given code in natural language Chain-of-Thought (CoT) rationales. This approach allows LLMs to learn universal reasoning primitives embedded in code, such as logic flow and decision-making, while decoupling them from specific programming syntax. An improved version, CodeI/O++, further refines training data through multi-turn revision based on execution feedback. Experimental results demonstrate that both CodeI/O and CodeI/O++ lead to consistent and balanced performance improvements across a wide range of symbolic, scientific, logical, and mathematical reasoning tasks in LLMs. Source: https://arxiv.org/html/2502.07316v1
This December 2024 paper introduces a collaborative inference framework designed for large-scale models in 5G smart city edge computing environments, addressing the challenge of limited memory and computing capacity on individual edge nodes. The framework partitions large models into sub-models deployed across multiple edge nodes and incorporates an early exit mechanism to accelerate inference. To manage the complexities of heterogeneous systems and dynamic environments, the authors propose a distributed algorithm called DTO-EE, which jointly optimizes task offloading strategies and confidence thresholds for early exits. Experimental results demonstrate that DTO-EE significantly reduces response delay and improves inference accuracy compared to existing methods. Source: https://arxiv.org/pdf/2412.08284
This August 2025 paper examines the evolving landscape of Federated Large Language Models (FedLLM), focusing on how large language models are post-trained while preserving user data privacy. The authors introduce a novel taxonomy that categorizes FedLLM approaches based on model accessibility (white-box, gray-box, and black-box) and parameter efficiency. It highlights various techniques within these categories, such as adapter-based tuning and prompt tuning, which reduce computational and communication overhead. The paper also discusses the growing importance of inference-only black-box settings for future FedLLM development and identifies open challenges like federated value alignment and enhanced security in constrained environments. Source: https://arxiv.org/html/2508.16261v1
This April 2025 paper introduces Self-Principled Critique Tuning (SPCT), a novel method designed to enhance the inference-time scalability of Generative Reward Models (GRMs) for various domains. It details how SPCT, through a combination of rejective fine-tuning and rule-based online reinforcement learning, facilitates the adaptive generation of principles and critiques, thereby improving the quality and inference-time scalability of GRMs. The paper compares different reward generation paradigms (scalar, semi-scalar, generative) and scoring patterns (pointwise, pairwise), demonstrating that the proposed DeepSeek-GRM models, particularly when guided by a meta Reward Model, consistently outperform existing methods across multiple benchmarks without significant domain biases. The research highlights the potential for GRMs to serve as a versatile interface for generalist reward systems, advancing Large Language Model (LLM) post-training and inference, while also acknowledging ethical considerations regarding bias and the importance of human-in-the-loop frameworks. Source: https://arxiv.org/pdf/2504.02495
This August 2025 paper introduces the Hierarchical Reasoning Model (HRM), a novel AI architecture inspired by the human brain's hierarchical and multi-timescale processing. This model aims to overcome the limitations of current large language models (LLMs) and Chain-of-Thought (CoT) techniques in complex reasoning tasks, which often suffer from computational inefficiencies and extensive data requirements. HRM utilizes two interdependent recurrent modules: a high-level module for abstract planning and a low-level module for detailed computations, enabling it to achieve significant computational depth. Notably, HRM demonstrates exceptional performance on challenging reasoning benchmarks like Sudoku and maze navigation with minimal training data (around 1000 samples) and without pre-training or CoT data. The research further explores the brain-like hierarchical dimensionality organization within HRM, where the high-level module operates in a higher-dimensional space, mirroring principles observed in the mammalian cortex. Source: https://arxiv.org/pdf/2506.21734
This January 2025 paper introduces Janus-Pro, an enhanced artificial intelligence model for multimodal understanding and generation. It builds upon its predecessor, Janus, through optimized training strategies, expanded data, and increased model size. The authors demonstrate that Janus-Pro achieves significant improvements in both multimodal understanding benchmarks and text-to-image generation capabilities, producing more stable and aesthetically pleasing outputs. This work highlights the benefits of decoupling visual encoding for understanding and generation tasks within a unified autoregressive transformer architecture. Source: https://arxiv.org/pdf/2501.17811
This February 2025 paper introduces Native Sparse Attention (NSA), a novel approach to address the computational demands of long-context modeling in large language models. NSA combines algorithmic innovations like a dynamic hierarchical sparse strategy with hardware-aligned optimizations to significantly improve efficiency. The paper highlights NSA's ability to maintain or even surpass the performance of traditional "Full Attention" models across various benchmarks, including general language, long-context tasks, and instruction-based reasoning, while achieving substantial speedups in decoding, forward, and backward propagation. It critically analyzes the shortcomings of existing sparse attention methods, particularly their failure to achieve practical speedups and support end-to-end training, thus motivating NSA's natively trainable and hardware-efficient design. NSA's architecture incorporates token compression, blockwise token selection, and a sliding window mechanism, underpinned by a specialized kernel designed for optimal GPU utilization. Source: https://arxiv.org/pdf/2502.11089
This June 2025 paper introduces Non-Penetrative Tensor Partitioning (NPTP), a novel method designed to improve the speed of collaborative inference for Deep Neural Networks (DNNs) on Internet of Things (IoT) devices. It addresses the common challenge of limited resources and strict latency requirements by minimizing the communication overhead that typically arises when large images are divided and processed across multiple devices. Unlike existing methods that utilize penetrative partitioning, which leads to substantial data sharing between devices, NPTP employs a non-penetrative approach and a Multilevel Partitioning Algorithm (MPA) to reduce this inter-device communication. Experimental results demonstrate that NPTP significantly outperforms state-of-the-art collaborative inference algorithms like CoEdge, achieving notable inference speedups, particularly for larger DNN models and image sizes, while maintaining device memory efficiency. The paper details the computational and communication overhead formulations, along with the algorithm design for optimal tensor partitioning. Source: https://arxiv.org/pdf/2501.04489
This March 2023 paper introduces PETALS, a novel system designed to facilitate the collaborative inference and fine-tuning of large language models (LLMs) by pooling resources from multiple participants. It addresses the significant computational and memory demands of LLMs, which typically restrict access for many researchers. PETALS proposes an alternative to traditional methods like slow RAM offloading or inflexible inference APIs by allowing distributed processing across a network of consumer GPUs, enhancing speed and flexibility. The system incorporates optimizations like 8-bit quantization and dynamic load balancing to improve performance and reliability. Ultimately, PETALS aims to democratize access to powerful LLMs, enabling broader research and application development that was previously cost-prohibitive. Source: https://arxiv.org/pdf/2209.01188
This August 2025 paper introduces UQ, a novel evaluation framework designed to challenge large language models (LLMs) with complex, unsolved questions sourced from platforms like Stack Exchange, where no definitive ground truth answers currently exist. The framework consists of three main components: UQ-Dataset, a collection of 500 hand-filtered, difficult, and unsolved questions; UQ-Validators, a set of LLM-based validation strategies that assess candidate solutions by leveraging the observation that models are often better at verifying answers than generating them; and UQ-Platform, which facilitates community engagement and human verification. The paper highlights the generator-validator gap, demonstrating that LLMs show improved performance in validating answers as their capabilities increase, and emphasizes that UQ aims to accelerate research in domains lacking clear ground-truth verification. Ultimately, UQ offers a new paradigm for evaluating advanced AI by focusing on realistic, open-ended problems that push the boundaries of current model capabilities. Source: https://arxiv.org/pdf/2508.17580
This July 2025 paper introduces Hierarchical Networks (H-Nets), a novel architecture designed to move beyond traditional tokenization in large language models by implementing dynamic chunking. This mechanism allows the model to automatically learn content- and context-dependent segmentation strategies directly from raw data, eliminating the need for predefined pre-processing steps like byte-pair encoding (BPE). H-Nets utilize a recursive, multi-stage structure that processes data at varying levels of abstraction, from bytes to more complex semantic units. Experiments demonstrate that H-Nets, particularly multi-stage configurations, outperform tokenized Transformers in perplexity, downstream tasks, and robustness to textual perturbations, especially in languages and modalities with weak or absent tokenization cues, such as Chinese, code, and DNA sequences. The authors highlight that this end-to-end learning of data chunking represents a significant step towards more generalized and efficient foundation models. Source: https://arxiv.org/html/2507.07955
This April 2025 paper introduces Infini-gram, a novel engine designed to scale n-gram language models to an unprecedented 5 trillion tokens and support unbounded n (∞-gram LMs). Unlike traditional methods that rely on pre-computed count tables, Infini-gram leverages suffix arrays for efficient, low-latency calculation of n-gram and ∞-gram probabilities, even for extremely long contexts. The authors demonstrate that this modernized approach significantly improves the perplexity of neural Large Language Models (LLMs), by up to 73%, by offering complementary insights into human-written and machine-generated text. Beyond enhancing LLMs, the Infini-gram engine also enables various applications such as corpus analysis, data curation, document retrieval, and detection of data contamination. A public web interface, API endpoint, and source code are provided to encourage further exploration and development within the community. Source: https://arxiv.org/pdf/2401.17377
This September 2025 paper introduces LoFT, a novel framework designed to improve Long-Tailed Semi-Supervised Learning (LTSSL) by leveraging parameter-efficient fine-tuning of pre-trained foundation models. The core idea is to enhance confidence calibration and generate more reliable pseudo-labels, which are crucial for addressing the imbalance inherent in long-tailed datasets. Furthermore, the paper extends this approach to open-world scenarios with LoFT-OW, specifically incorporating mechanisms to detect and filter out-of-distribution (OOD) samples from unlabeled data. The authors demonstrate that these fine-tuned models achieve superior performance on various benchmarks, even when utilizing significantly less unlabeled data compared to previous methods. Source: https://arxiv.org/pdf/2509.09926
This July 2025 paper discusses advanced memory optimization techniques for Large Language Models (LLMs), particularly focusing on KV cache management in multi-tenant serving environments. The primary subject, MIRAGE, introduces parameter remapping, a novel method that dynamically repurposes GPU memory allocated for model parameters to expand KV cache capacity, outperforming traditional CPU-offloading and KV cache swapping by reducing latency and increasing throughput. Complementary research highlights challenges in on-device LLM deployment and proposes solutions like quantization (AWQ) for model compression and two-level scheduling (FineServe, Nexus) for efficient GPU sharing to mitigate memory fragmentation and improve performance. Overall, the papers underscore the critical need for innovative memory management to address the growing memory demands of LLMs and enhance their inference serving efficiency across diverse hardware configurations. Source: https://www.researchgate.net/publication/393724496_MIRAGE_KV_Cache_Optimization_through_Parameter_Remapping_for_Multi-tenant_LLM_Serving
This September 2025 paper describes QuantAgent, a novel multi-agent large language model (LLM) framework designed for high-frequency quantitative trading based solely on price-derived market signals. The system decomposes trading decisions into four specialized agents—IndicatorAgent, PatternAgent, TrendAgent, and DecisionAgent—which analyze market dynamics from complementary perspectives and communicate through structured prompts. QuantAgent consistently outperforms baseline models across diverse assets, including commodities, equities, and cryptocurrencies, demonstrating robust generalization and achieving high directional accuracy in predicting price movements. A key feature is its ability to produce traceable, language-native explanations for its trading decisions, enhancing transparency and interpretability in a field traditionally marked by opaque algorithms. The framework also includes a local, browser-based interface that allows users to visualize market data and interact with LLM-generated analyses, emphasizing user control and understanding. Source: https://arxiv.org/pdf/2509.09995
This April 2025 paper introduces ShadowKV, an innovative inference system for long-context Large Language Models (LLMs) designed to significantly enhance throughput and support larger batch sizes without compromising accuracy. It achieves this by strategically managing the Key-Value (KV) cache: specifically, it compresses the low-rank pre-Rotary Position Embedding (RoPE) key cache on the GPU and offloads the value cache to the CPU. ShadowKV further optimizes performance through an accurate KV selection strategy that reconstructs minimal sparse KV pairs on-the-fly, thus minimizing decoding latency. Empirical evaluations demonstrate that ShadowKV can support up to 6x larger batch sizes and boost throughput by up to 3.04x on an A100 GPU across various LLMs and benchmarks, even outperforming theoretical infinite memory scenarios. Source: https://arxiv.org/pdf/2410.21465
This May 2025 paper introduces TailorKV, a novel hybrid framework designed to optimize Key-Value (KV) cache management in large language models (LLMs) for long-context inference. It addresses challenges like high GPU memory consumption and inference latency that arise from the linear growth of KV cache size with sequence length. TailorKV categorizes Transformer layers into quantization-friendly and sparsity-friendly based on their attention patterns, applying 1-bit quantization to the former and dynamic retrieval of Top-K tokens from CPU memory for the latter. This tailored approach significantly reduces memory usage and decoding latency while maintaining model accuracy, enabling LLMs to operate efficiently on resource-limited hardware. Source: https://arxiv.org/pdf/2505.19586
This September 2025 paper introduces WebSailor-V2, an open-source deep research agent developed by Alibaba Group's Tongyi Lab. The paper details a post-training pipeline that uses a novel synthetic data construction scheme, SailorFog-QA-V2, and a dual-environment reinforcement learning framework. WebSailor-V2, built on the Qwen3-30B-A3B model, demonstrates state-of-the-art performance among open-source agents and is competitive with leading proprietary systems on various web-agent benchmarks, including BrowseComp and Humanity's Last Exam. The authors emphasize that high-quality data and a stable training environment are more crucial than the specific RL algorithm for developing robust AI agents. Source: https://arxiv.org/pdf/2509.13305
This paper published on Nature on September 17 2025, "DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning," details the development of DeepSeek-R1-Zero and DeepSeek-R1, two large language models (LLMs) engineered to enhance reasoning capabilities. The authors explain how reinforcement learning (RL) is used to enable emergent advanced reasoning patterns like self-reflection and dynamic strategy adaptation, moving beyond reliance on human-annotated data. The paper discusses a multistage training pipeline for DeepSeek-R1, integrating rejection sampling, RL, and supervised fine-tuning to improve both reasoning and general language tasks while addressing issues like language mixing. Furthermore, the researchers highlight the release of these models and their distilled, smaller versions to the public to contribute to ongoing AI research. Ultimately, the source concludes by acknowledging the ethical considerations and limitations of their pure RL methodology, such as reward hacking and token efficiency. Source: https://www.nature.com/articles/s41586-025-09422-z
How can pre-computing and reusing Key-Value (KV) caches accelerate inference for Retrieval-Augmented Generation and other long-context LLM tasks? The provided sources identify the same core problem—high latency in Large Language Model (LLM) inference due to processing long, repetitive contexts—and converge on a unified solution: leveraging pre-computed Key-Value (KV) caches. Each source then contributes a unique perspective on *how* to implement this solution effectively, addressing specific challenges that arise from this approach. The unified answer proposed by all sources is to avoid redundant computation by pre-computing, storing, and reusing the KV caches of recurring text segments (referred to as chunks, documents, or prompt modules).Sources:https://arxiv.org/html/2502.15734v1https://arxiv.org/html/2412.15605v1https://arxiv.org/html/2502.16002v1https://arxiv.org/html/2310.07240v6https://arxiv.org/pdf/2404.12457https://openreview.net/pdf?id=x7NbaU8RSUhttps://proceedings.mlsys.org/paper_files/paper/2024/file/a66caa1703fe34705a4368c3014c1966-Paper-Conference.pdfhttps://www.cs.princeton.edu/~ravian/COS597_F24/papers/cacheblend.pdf
This September 2025 academic paper, titled "REFRAG: Rethinking RAG based Decoding," appears on the alphaXiv pre-print server. It focuses on Reframing Retrieval-Augmented Generation (RAG) within the context of decoding processes. The research likely explores new approaches or improvements to how RAG mechanisms are integrated and utilized during the generation of text. By "rethinking" this integration, the authors probably aim to optimize the performance or efficiency of language models that leverage external knowledge through RAG. The source, specifically from alphaXiv, indicates it is a scholarly work shared for peer review and discussion within the scientific community. Source: https://www.alphaxiv.org/abs/2509.01092v1
This September 2025 paper source is a research paper from Tencent AI Lab and academic collaborators that introduces EVOL-RL, an Evolution-Oriented and Label-free Reinforcement Learning framework for Large Language Models (LLMs). The paper addresses a critical flaw, termed entropy collapse, in existing label-free self-improvement methods like Test-Time Reinforcement Learning (TTRL), where reliance solely on a majority vote leads to a shrink in solution diversity and poor generalization. EVOL-RL overcomes this by incorporating a novel reward system that explicitly balances majority-based selection (for stability) with a novelty-aware reward (for variation), preventing the model from converging to repetitive, low-entropy solutions. Experimental results on mathematical reasoning benchmarks demonstrate that EVOL-RL significantly improves accuracy and generalization by sustaining diverse, longer chains of thought compared to the TTRL baseline. Source: https://arxiv.org/pdf/2509.15194
This September 2025 paper introduces FlowRL, a novel reinforcement learning (RL) algorithm for large language models (LLMs) that shifts the optimization objective from reward maximization to reward distribution matching via flow balancing. Traditional RL methods like PPO and GRPO tend to over-optimize high-reward paths, leading to limited solution diversity and mode collapse, particularly in complex tasks like long Chain-of-Thought (CoT) reasoning. FlowRL addresses this by minimizing the reverse KL divergence between the policy and a reward-weighted target distribution, which is shown to be equivalent to the trajectory balance loss from GFlowNets, thereby jointly promoting reward and entropy maximization. Through experiments on math and code reasoning benchmarks, FlowRL demonstrates significant performance gains—an average of 10.0% over GRPO and 5.1% over PPO on math tasks—by generating substantially more diverse and generalizable reasoning trajectories. Source: https://arxiv.org/pdf/2509.15207
This September 2025 paper introduces SearchInstruct, a novel framework designed to enhance Supervised Fine-Tuning (SFT) of large language models (LLMs) by constructing high-quality, domain-specific instruction datasets. This approach overcomes challenges like data scarcity and outdated model knowledge by dynamically retrieving external, up-to-date documents to generate accurate, context-grounded answers for augmented questions. SearchInstruct operates via a four-stage pipeline: starting with a small set of human-generated questions, expanding them using an LLM, retrieving relevant documents (via RAG or web search), and synthesizing context-aware responses. Experimental results confirm that the method significantly boosts LLM performance in specialized domains like Iranian culture and efficiently facilitates model editing for factual updates, creating a measurable improvement over baseline models. Source: https://arxiv.org/pdf/2509.10708
This September 2025 paper introduces Single-stream Policy Optimization (SPO), a new reinforcement learning algorithm for training Large Language Models (LLMs) developed by Tencent researchers. SPO challenges the prevailing group-based optimization methods like Group Relative Policy Optimization (GRPO), which suffer from high computational waste due to "degenerate groups" and synchronization bottlenecks, particularly in complex agentic tasks. The core of SPO involves returning to a single-stream paradigm, using a persistent, KL-adaptive value tracker as a stable baseline, and applying global advantage normalization to ensure efficient and stable learning. Empirical results on challenging math benchmarks, using the Qwen3-8B model, demonstrate that SPO consistently outperforms GRPO in terms of accuracy and achieves a significant 4.35x speedup in simulated high-variance agentic training environments, validating its superior scalability and efficiency. Source: https://arxiv.org/pdf/2509.13232
The "Anthropic Economic Index report" documents the rapid and uneven adoption of Artificial Intelligence (AI), specifically using data from the company's Claude.ai consumer platform and its enterprise API. The report highlights that AI adoption is geographically concentrated in high-income regions and that the speed of adoption surpasses previous technologies like the internet. Analyzing usage patterns, the authors find that coding tasks dominate usage across both consumer and enterprise users, but enterprise use via API is highly automated (77%) compared to the more collaborative (augmentation) consumer use. The study introduces the Anthropic AI Usage Index (AUI) to compare per-capita AI use across regions and suggests that for businesses, model capability and economic value matter more than cost, while the availability of contextual information may be a crucial bottleneck for sophisticated deployment. Source: https://assets.anthropic.com/m/218c82b858610fac/original/Economic-Index.pdf
This September 2025 paper describes THOR (Tool-Integrated Hierarchical Optimization via RL), a novel approach designed to enhance the mathematical reasoning and code generation capabilities of Large Language Models (LLMs) by integrating external code-execution tools. The methodology introduces TIRGen, a pipeline for creating high-quality Tool-Integrated Reasoning (TIR) data, which is crucial for training the model using a hierarchical reinforcement learning (RL) strategy. This RL framework incorporates both trajectory-level optimization for overall problem-solving ability and step-level optimization to correct code generation errors, addressing the sparse reward problem common in long reasoning tasks. Experimental results demonstrate that THOR achieves state-of-the-art (SOTA) performance across various mathematical and code benchmarks for both reasoning and non-reasoning models. Finally, the system leverages code execution feedback for a self-correction inference enhancement mechanism, which further improves performance, especially on more challenging problems. Source: https://arxiv.org/pdf/2509.13761
These 14 research papers provide an overview of various compression techniques for Large Language Models (LLMs), primarily focusing on reducing the size and computational overhead of the Key-Value (KV) cache to handle long contexts more efficiently. Several novel methods are detailed, including GVote, an adaptive compression algorithm using query sampling and voting to find an optimal cache budget, and SnapKV, which selects clustered, important KV positions based on an "observation" window to maintain performance while increasing speed and memory efficiency. Other approaches include POD (Proximal tokens over Distant tokens), which reduces redundancy by sharing key states across layers for distant tokens while preserving proximal ones, and DecoQuant, a quantization method utilizing matrix decomposition to reduce errors. The sources also examine prompt compression methods like LLMLingua and LongLLMLingua, and describe CASC (Context-Adaptive Synthesis and Compression), a Retrieval-Augmented Generation (RAG) framework that intelligently synthesizes and compresses multi-document contexts to improve answer accuracy in complex domains.Sources:https://arxiv.org/pdf/2509.08315https://arxiv.org/html/2509.09199v1https://arxiv.org/html/2509.03136v1https://aclanthology.org/2025.acl-long.1394.pdfhttps://proceedings.neurips.cc/paper_files/paper/2024/file/fd0705710bf01b88a60a3d479ea341d9-Paper-Conference.pdfhttps://arxiv.org/html/2412.14838v1https://arxiv.org/pdf/2412.02252https://aclanthology.org/2024.acl-long.133.pdfhttps://arxiv.org/html/2508.19357v1https://aclanthology.org/2024.acl-long.91.pdfhttps://arxiv.org/html/2310.05736v2https://aclanthology.org/2025.naacl-long.368.pdfhttps://arxiv.org/pdf/2404.14469https://aclanthology.org/2024.findings-emnlp.266.pdf
The September 2025 academic paper introduces LLM-Interleaved (LLM-I), a novel, flexible framework for interleaved image-text generation that reframes the task as a tool-use problem to overcome the "one-tool" limitation of unified models. Authored by researchers from Zhejiang University and ByteDance, BandAI, the system uses a central Large Language Model (LLM) or Multimodal LLM (MLLM) agent to orchestrate a diverse toolkit of specialized visual tools, including online image search, diffusion generation, code execution, and image editing. The agent is trained using a Reinforcement Learning (RL) framework featuring a hybrid reward system that combines rule-based logic with LLM and MLLM evaluators. The research demonstrates that LLM-I achieves state-of-the-art performance across four benchmarks by moving from an "omniscient solver" to a "proficient tool-user" paradigm, allowing for factually grounded and programmatically precise visual outputs. Source: https://arxiv.org/pdf/2509.13642
The September 25 2035 paper introduces a novel reinforcement learning (RL) algorithm, Controlling Entropy via Gradient-Preserving Policy Optimization (CE-GPPO), designed to fine-tune large language models (LLMs) for complex reasoning tasks. The authors analyze how policy entropy, which represents the balance between exploration and exploitation, becomes unstable in existing methods like Proximal Policy Optimization (PPO) due to the clipping of low-probability tokens. CE-GPPO addresses this by reintroducing gradients from these clipped tokens—specifically Positive-advantage Low-Probability (PA&LP) and Negative-advantage Low-Probability (NA&LP) tokens—in a bounded and controlled manner. The goal is to regulate entropy dynamics and prevent both entropy collapse and entropy explosion. Empirical results on mathematical reasoning benchmarks show that CE-GPPO consistently outperforms strong baselines by maintaining more stable and optimal entropy throughout training. Source: https://arxiv.org/pdf/2509.20712
The September 24 2025 paper introduces EmbeddingGemma, a novel, lightweight text embedding model developed by Google DeepMind, built upon the Gemma 3 language model family. The paper details the innovative training methodology, which involves encoder-decoder initialization and geometric embedding distillation from larger models like Gemini Embedding, alongside a "spread-out" regularizer and model souping for improved expressiveness and generalizability. Through extensive evaluation on the Massive Text Embedding Benchmark (MTEB), the 308M-parameter model is shown to achieve state-of-the-art performance among models under 500M parameters across multilingual, English, and code tasks, often rivaling models double its size, thus offering an exceptional performance-to-cost ratio suitable for low-latency, on-device applications. Ablation studies support the design choices, concluding that the encoder-decoder initialization and mean pooling provide the strongest foundation for high-quality embeddings. Source: https://arxiv.org/pdf/2509.20354
The September 25 2025 dated sources introduce GDPval, a novel benchmark created by OpenAI to evaluate the performance of AI models on economically valuable, real-world tasks. This evaluation spans 44 knowledge work occupations across the top nine sectors contributing to the U.S. GDP, using tasks meticulously crafted by experienced industry professionals. Results indicate that the best frontier models are approaching human expert quality on these tasks, with models like Claude Opus 4.1 and GPT-5 demonstrating strengths in different areas, such as aesthetics and accuracy, respectively. Furthermore, the analysis suggests that integrating AI can potentially lead to significant speed and cost improvements in expert workflows, while noting that model performance is still limited by the real-world complexity of multi-draft and ambiguous tasks. Finally, OpenAI is open-sourcing a subset of tasks and an automated grader to facilitate further research in tracking AI capabilities.Sources:https://openai.com/index/gdpval/https://cdn.openai.com/pdf/d5eb7428-c4e9-4a33-bd86-86dd4bcf12ce/GDPval.pdf
The September 24 2025 paper is a technical report from ByteDance Seed detailing the Seedream 4.0 system, an advanced multimodal image generation model. This single framework efficiently unifies text-to-image synthesis, image editing, and multi-image composition. The core innovation is an efficient diffusion transformer and powerful VAE that facilitates fast generation of high-resolution images (up to 4K), outperforming competitors like GPT-4o and Gemini 2.5 in human and automatic evaluations. The system uses sophisticated training techniques, including multi-modal post-training and adversarial acceleration methods, to achieve state-of-the-art results and ultra-fast inference speed. Seedream 4.0 is designed for both creative and professional applications, featuring capabilities like in-context reasoning and advanced text rendering. Source: https://arxiv.org/pdf/2509.20427
The September 25 2025 paper introduces Tree-based Group Relative Policy Optimization (Tree-GRPO), a new reinforcement learning (RL) method designed to enhance the agentic capabilities of large language models (LLMs) in multi-turn tasks where supervision is typically sparse. Tree-GRPO addresses the challenges of sparse rewards and heavy rollout costs associated with existing chain-based RL by employing a tree-search sampling strategy where each node represents a complete agent interaction step, allowing for prefix sharing and reduced budget use. This tree structure inherently creates finer-grained process supervision signals from outcome rewards, a mechanism shown to be structurally equivalent to step-level direct preference learning. Empirical results across multiple datasets demonstrate that the tree-based approach consistently achieves higher performance with less rollout budget compared to chain-based methods. Source: https://arxiv.org/pdf/2509.21240
This Meta September 24 2025 paper provides an extensive overview of Code World Model (CWM), a 32-billion-parameter dense decoder-only Transformer designed for coding and reasoning tasks, highlighting its architecture and multi-stage training process. Training involves pre-training, mid-training, and post-training stages which include supervised fine-tuning (SFT) and joint reinforcement learning (RL) across environments like software engineering (SWE) tasks, coding problems, and mathematics. A core feature is CWM's ability to process and predict Python execution traces and perform agentic interactions using a minimal set of tools within containerized environments. The document details the model's competitive performance against other large language models on benchmarks like SWE-bench Verified and discusses infrastructure choices, such as asynchronous RL and fp8 matrix multiplication, used to achieve training efficiency.Sources:https://ai.meta.com/temp/research/publications/cwm-an-open-weights-llm-for-research-on-code-generation-with-world-models/https://scontent-lax3-2.xx.fbcdn.net/v/t39.2365-6/553592426_661450129912484_4072750821656455102_n.pdf?_nc_cat=103&ccb=1-7&_nc_sid=3c67a6&_nc_ohc=0-g1m0kIX7cQ7kNvwFJRCs6&_nc_oc=AdnQYyoahmWWXKZWyQuh4F09IuhBGd08Uwox14N8BdY_tMilZ5_Tl5u7P82HLIJ9RSc8nDy188xiuxmmByXhkJ1S&_nc_zt=14&_nc_ht=scontent-lax3-2.xx&_nc_gid=RcYYIy-y9eIernCj-naXRQ&oh=00_AfYC3ol0MNn_PokD4H3hoOcpxg6OoOsgWfD4pDyThszIpA&oe=68DD28B5
This September 20 2025 paper introduce a novel, efficient architecture for training retrieval models used in retrieval-augmented generation (RAG) systems. This architecture addresses the inefficiency of fine-tuning large models by combining adapters for soft embeddings with a Classifier-as-Retriever (CaR) approach. The soft embeddings, created by lightweight layers in a frozen small language model (SLM), efficiently adapt the model to new corpora, while the CaR replaces static maximum inner product search (MIPS) with a trainable classifier for significantly higher accuracy (up to 99%). Furthermore, the methods integrate naturally with federated learning (FL) to achieve distributed training speedups (up to 2.6x faster) and utilize differential privacy (DP) techniques to safeguard client data during edge device training. This combined approach results in a lighter, faster, and privacy-preserving solution for domain-specific RAG.Sources:https://www.webai.com/blog/federated-learning-with-soft-embeddings-a-new-efficient-way-to-train-retrieval-modelshttps://arxiv.org/pdf/2509.16508
This September 18 2025 paper introduces a research project that applies Schoenfeld’s Episode Theory, a classic cognitive framework for analyzing human mathematical problem-solving, to understand the reasoning processes of Large Reasoning Models (LRMs). The authors created a novel, publicly available benchmark by annotating thousands of sentences and paragraphs from model-generated solutions to math problems, using seven cognitive labels such as Plan, Implement, and Verify. This approach offers a theoretically grounded methodology for interpreting LRM cognition, demonstrating that machine reasoning exhibits structured, episodic patterns similar to human behavior. The resulting annotated corpus and analytical protocol aim to enable the development of more transparent and controllable reasoning systems. Source: https://arxiv.org/pdf/2509.14662
This September 26 2025 paper is an excerpt from a research paper introducing a variational reasoning framework designed to enhance the reasoning capabilities of language models (LLMs). This framework conceptualizes thinking traces as latent variables and uses variational inference to optimize them, building upon the Evidence Lower Bound (ELBO) and extending it with a tighter, multi-trace IWAE-style bound. Crucially, the paper proposes a forward-KL objective for stabilizing the training of the variational posterior, which samples high-quality thinking paths. The research also interprets existing methods like Rejection Sampling Finetuning (RFT) and binary-reward Reinforcement Learning (RL) as local forward-KL objectives, highlighting a previously unrecognized bias toward easier questions in these traditional approaches. Empirical validation on the Qwen 2.5 and Qwen 3 models across diverse benchmarks confirms that this principled probabilistic perspective leads to consistent performance improvements and greater training stability compared to strong baselines. Source: https://arxiv.org/pdf/2509.22637
These May and September 2025 technical reports introduce and evaluate two distinct but related large language models: the Qwen3 family and the Qwen3-Omni multimodal system. The first source focuses on the text-based Qwen3 models, highlighting their development process, which includes a sophisticated multilingual data annotation system and a multi-stage training pipeline incorporating Strong-to-Weak Distillation and "Thinking Mode" for complex tasks like coding and mathematics. The second, more comprehensive source describes Qwen3-Omni, a single model designed to excel across text, image, audio, and video modalities without performance degradation, utilizing a novel Thinker–Talker Mixture-of-Experts (MoE) architecture for real-time speech and reasoning. Crucially, Qwen3-Omni achieves state-of-the-art performance in audio tasks and boasts extremely low first-packet latency for interactive applications, supporting a wide range of multilingual capabilities in both speech and text.Sources:https://arxiv.org/pdf/2505.09388https://arxiv.org/pdf/2509.17765