This episode explores a speech recognition paper that replaces slow left-to-right transcript generation with a faster draft-and-edit approach, where a speech model produces an initial hypothesis and a bidirectional LLM corrects it in parallel. It explains the tradeoff between CTC-based systems, which are fast but weaker at using linguistic context, and autoregressive decoders, which are more expressive but too slow for low-latency use cases like captioning and meetings. The discussion highlights the paper’s key ideas, including transcript editing with insertion slots, latent alignment inspired by CTC, and the use of LoRA to adapt pretrained language models efficiently. Listeners would find it interesting because it shows a concrete path to pushing ASR onto a better speed-accuracy frontier, with reported gains such as a 27x speedup over an autoregressive baseline while staying competitive on word error rate.
Sources:
1. NLE: Non-autoregressive LLM-based ASR by Transcript Editing — Avihu Dekel, Samuel Thomas, Takashi Fukada, George Saon, 2026
http://arxiv.org/abs/2603.083972. Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks — Alex Graves, Santiago Fernandez, Faustino Gomez, Jurgen Schmidhuber, 2006
https://scholar.google.com/scholar?q=Connectionist+Temporal+Classification%3A+Labelling+Unsegmented+Sequence+Data+with+Recurrent+Neural+Networks3. Listen, Attend and Spell — William Chan, Navdeep Jaitly, Quoc V. Le, Oriol Vinyals, 2015
https://scholar.google.com/scholar?q=Listen%2C+Attend+and+Spell4. Exploring Architectures, Data and Units for Streaming End-to-End Speech Recognition with RNN-Transducer — Hasim Sak, Kanishka Rao, Rohit Prabhavalkar, 2017
https://scholar.google.com/scholar?q=Exploring+Architectures%2C+Data+and+Units+for+Streaming+End-to-End+Speech+Recognition+with+RNN-Transducer5. Robust Speech Recognition via Large-Scale Weak Supervision — Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, Ilya Sutskever, 2022
https://scholar.google.com/scholar?q=Robust+Speech+Recognition+via+Large-Scale+Weak+Supervision6. Mask CTC: Non-Autoregressive End-to-End ASR with CTC and Mask Predict — Yosuke Higuchi, Shinji Watanabe, Nanxin Chen, Tetsuji Ogawa, Tetsunori Kobayashi, 2020
https://scholar.google.com/scholar?q=Mask+CTC%3A+Non-Autoregressive+End-to-End+ASR+with+CTC+and+Mask+Predict7. Align-Refine: Non-autoregressive Speech Recognition via Iterative Realignment — Ethan A. Chi, Julian Salazar, Katrin Kirchhoff, 2021
https://scholar.google.com/scholar?q=Align-Refine%3A+Non-autoregressive+Speech+Recognition+via+Iterative+Realignment8. A Survey on Non-Autoregressive Generation for Neural Machine Translation and Beyond — Yisheng Xiao, Lijun Wu, Junliang Guo, Juntao Li, Min Zhang, Tao Qin, Tie-Yan Liu, 2022
https://scholar.google.com/scholar?q=A+Survey+on+Non-Autoregressive+Generation+for+Neural+Machine+Translation+and+Beyond9. Encode, Tag, Realize: High-Precision Text Editing — Eric Malmi, Sebastian Krause, Sascha Rothe, Daniil Mirylenka, Aliaksei Severyn, 2019
https://scholar.google.com/scholar?q=Encode%2C+Tag%2C+Realize%3A+High-Precision+Text+Editing10. Levenshtein Transformer — Jiatao Gu, Changhan Wang, Junbo Zhao, 2019
https://scholar.google.com/scholar?q=Levenshtein+Transformer11. GECToR - Grammatical Error Correction: Tag, Not Rewrite — Kostiantyn Omelianchuk, Vitaliy Atrasevych, Artem Chernodub, Oleksandr Skurzhanskyi, 2020
https://scholar.google.com/scholar?q=GECToR+-+Grammatical+Error+Correction%3A+Tag%2C+Not+Rewrite12. FELIX: Flexible Text Editing Through Tagging and Insertion — Jonathan Mallinson, Aliaksei Severyn, Eric Malmi, Guillermo Garrido, 2020
https://scholar.google.com/scholar?q=FELIX%3A+Flexible+Text+Editing+Through+Tagging+and+Insertion13. LoRA: Low-Rank Adaptation of Large Language Models — Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, 2021
https://scholar.google.com/scholar?q=LoRA%3A+Low-Rank+Adaptation+of+Large+Language+Models14. Token-and-Duration Transducer — Dmitriy Genzel, Yanzhang He, et al., 2023
https://scholar.google.com/scholar?q=Token-and-Duration+Transducer15. Mask-Predict: Parallel Decoding of Conditional Masked Language Models — Marjan Ghazvininejad, Omer Levy, Yinhan Liu, Luke Zettlemoyer, 2019
https://scholar.google.com/scholar?q=Mask-Predict%3A+Parallel+Decoding+of+Conditional+Masked+Language+Models16. QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache — Rishabh Tiwari, Haocheng Xi, Aditya Tomar, et al., 2025
https://scholar.google.com/scholar?q=QuantSpec%3A+Self-Speculative+Decoding+with+Hierarchical+Quantized+KV+Cache17. SpeCache: Speculative Key-Value Caching for Efficient Generation of LLMs — Shibo Jie, Yehui Tang, Kai Han, Zhi-Hong Deng, Jing Han, 2025
https://scholar.google.com/scholar?q=SpeCache%3A+Speculative+Key-Value+Caching+for+Efficient+Generation+of+LLMs18. SoftCorrect: Error Correction with Soft Detection for Automatic Speech Recognition — Yichong Leng, Xu Tan, Wenjie Liu, Kaitao Song, Rui Wang, Xiang-Yang Li, Tao Qin, Edward Lin, Tie-Yan Liu, 2022
https://scholar.google.com/scholar?q=SoftCorrect%3A+Error+Correction+with+Soft+Detection+for+Automatic+Speech+Recognition19. Semantic Segmentation with Bidirectional Language Models Improves Long-form ASR — W. Ronny Huang, Hao Zhang, Shankar Kumar, Shuo-Yiin Chang, Tara N. Sainath, 2023
https://scholar.google.com/scholar?q=Semantic+Segmentation+with+Bidirectional+Language+Models+Improves+Long-form+ASR20. Bi-directional Context-Enhanced Speech Large Language Models for Multilingual Conversational ASR — Yizhou Peng, Hexin Liu, Eng Siong Chng, 2025
https://scholar.google.com/scholar?q=Bi-directional+Context-Enhanced+Speech+Large+Language+Models+for+Multilingual+Conversational+ASR21. AI Post Transformers: Qwen3.5-Omni Thinker-Talker for Omnimodal Streaming — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-21-qwen35-omni-thinker-talker-for-omnimodal-36b26c.mp322. AI Post Transformers: Batch-Aware Expert Routing for Faster MoE Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-batch-aware-expert-routing-for-faster-mo-683ab6.mp3Interactive Visualization: Fast Speech Recognition by Transcript Editing