This episode explores ELF: Embedded Language Flows, a continuous-time diffusion language model that stays in embedding space until the final decoding step instead of repeatedly snapping back to discrete tokens during generation. It explains how that design lets the model borrow flow-matching and guidance techniques from image diffusion, while arguing that earlier continuous text models may have underperformed because of token-level constraints rather than any fundamental weakness. The discussion highlights reported results on OpenWebText, where a 105M-parameter ELF model achieves better generative perplexity than 170M baselines with far fewer training tokens and fewer sampling steps, while also extending to translation and summarization. It also digs into the main caveat: whether the gains really come from late discretization and continuous-time modeling, or from a bundle of confounded training and inference choices, making the episode interesting both as a technical walkthrough and as a skeptical evaluation of a bold research claim.
Sources:
1. ELF and Continuous Language Diffusion
https://arxiv.org/pdf/2605.109382. Structured Denoising Diffusion Models in Discrete State-Spaces — Jacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow, Rianne van den Berg, 2021
https://scholar.google.com/scholar?q=Structured+Denoising+Diffusion+Models+in+Discrete+State-Spaces3. Diffusion-LM Improves Controllable Text Generation — Xiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang, Tatsunori B. Hashimoto, 2022
https://scholar.google.com/scholar?q=Diffusion-LM+Improves+Controllable+Text+Generation4. Simple and Effective Masked Diffusion Language Models — Subham Sekhar Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan, Edgar Marroquin, Justin T. Chiu, Alexander M. Rush, Volodymyr Kuleshov, 2024
https://scholar.google.com/scholar?q=Simple+and+Effective+Masked+Diffusion+Language+Models5. Large Language Diffusion Models — Shen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang, Jingyang Ou, Jun Hu, Jun Zhou, Yankai Lin, Ji-Rong Wen, Chongxuan Li, 2025
https://scholar.google.com/scholar?q=Large+Language+Diffusion+Models6. Flow Matching for Generative Modeling — Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, Matthew Le, 2023
https://scholar.google.com/scholar?q=Flow+Matching+for+Generative+Modeling7. Improving and Generalizing Flow-Based Generative Models with Minibatch Optimal Transport — Alexander Tong, Kilian Fatras, Nikolay Malkin, Guillaume Huguet, Yanlei Zhang, Jarrid Rector-Brooks, Guy Wolf, Yoshua Bengio, 2024
https://scholar.google.com/scholar?q=Improving+and+Generalizing+Flow-Based+Generative+Models+with+Minibatch+Optimal+Transport8. Scaling Rectified Flow Transformers for High-Resolution Image Synthesis — Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Muller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yannik Marek, Robin Rombach, 2024
https://scholar.google.com/scholar?q=Scaling+Rectified+Flow+Transformers+for+High-Resolution+Image+Synthesis9. Discrete Flow Matching — Itai Gat, Tal Remez, Neta Shaul, Felix Kreuk, Ricky T. Q. Chen, Gabriel Synnaeve, Yossi Adi, Yaron Lipman, 2024
https://scholar.google.com/scholar?q=Discrete+Flow+Matching10. Self-conditioned Embedding Diffusion for Text Generation — Robin Strudel, Corentin Tallec, Florent Altche, Yilun Du, Yaroslav Ganin, Arthur Mensch, Will Grathwohl, Nikolay Savinov, Sander Dieleman, Laurent Sifre, Remi Leblond, 2022
https://scholar.google.com/scholar?q=Self-conditioned+Embedding+Diffusion+for+Text+Generation11. Difformer: Empowering Diffusion Models on the Embedding Space for Text Generation — Zhujin Gao, Junliang Guo, Xu Tan, Yongxin Zhu, Fang Zhang, Jiang Bian, Linli Xu, 2022
https://scholar.google.com/scholar?q=Difformer%3A+Empowering+Diffusion+Models+on+the+Embedding+Space+for+Text+Generation12. LangFlow: Continuous Diffusion Rivals Discrete in Language Modeling — Yuxin Chen, Chumeng Liang, Hangke Sui, Ruihan Guo, Chaoran Cheng, Jiaxuan You, Ge Liu, 2026
https://scholar.google.com/scholar?q=LangFlow%3A+Continuous+Diffusion+Rivals+Discrete+in+Language+Modeling13. Classifier-Free Diffusion Guidance — Jonathan Ho, Tim Salimans, 2021
https://scholar.google.com/scholar?q=Classifier-Free+Diffusion+Guidance14. GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models — Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, Mark Chen, 2021
https://scholar.google.com/scholar?q=GLIDE%3A+Towards+Photorealistic+Image+Generation+and+Editing+with+Text-Guided+Diffusion+Models15. Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding — Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J. Fleet, Mohammad Norouzi, 2022
https://scholar.google.com/scholar?q=Photorealistic+Text-to-Image+Diffusion+Models+with+Deep+Language+Understanding16. High-Resolution Image Synthesis with Latent Diffusion Models — Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Bjorn Ommer, 2021
https://scholar.google.com/scholar?q=High-Resolution+Image+Synthesis+with+Latent+Diffusion+Models17. MDLM: Masked Diffusion Language Models — likely the MDLM authors cited as [56] in the paper, 2024
https://scholar.google.com/scholar?q=MDLM%3A+Masked+Diffusion+Language+Models18. Duo — likely the Duo authors cited as [57] in the paper, 2025
https://scholar.google.com/scholar?q=Duo19. Latent Diffusion Models — Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Bjorn Ommer, 2022
https://scholar.google.com/scholar?q=Latent+Diffusion+Models20. LangFlow — the LangFlow authors cited as [10] in the paper, 2026
https://scholar.google.com/scholar?q=LangFlow21. FLM — the FLM authors cited as [30] in the paper, 2026
https://scholar.google.com/scholar?q=FLM22. Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution — Aaron Lou, Chenlin Meng, Stefano Ermon, 2024
https://scholar.google.com/scholar?q=Discrete+Diffusion+Modeling+by+Estimating+the+Ratios+of+the+Data+Distribution23. Scaling Behavior of Discrete Diffusion Language Models — Dimitri von Rutte, Janis Fluri, Omead Pooladzandi, Bernhard Scholkopf, Thomas Hofmann, Antonio Orvieto, 2025
https://scholar.google.com/scholar?q=Scaling+Behavior+of+Discrete+Diffusion+Language+Models24. Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner — Cai Zhou, Chenxiao Yang, Yi Hu, Chenyu Wang, Chubin Zhang, Muhan Zhang, Lester Mackey, Tommi Jaakkola, Stephen Bates, Dinghuai Zhang, 2025
https://scholar.google.com/scholar?q=Coevolutionary+Continuous+Discrete+Diffusion%3A+Make+Your+Diffusion+Language+Model+a+Latent+Reasoner25. Stay on Topic with Classifier-Free Guidance — Guillaume Sanchez, Honglu Fan, Alexander Spangher, Elad Levi, Pawan Sasanka Ammanamanchi, Stella Biderman, 2023
https://scholar.google.com/scholar?q=Stay+on+Topic+with+Classifier-Free+Guidance26. Adaptive Classifier-Free Guidance via Dynamic Low-Confidence Masking — Pengxiang Li, Shilin Yan, Joey Tsai, Renrui Zhang, Ruichuan An, Ziyu Guo, Xiaowei Gao, 2025
https://scholar.google.com/scholar?q=Adaptive+Classifier-Free+Guidance+via+Dynamic+Low-Confidence+Masking27. Studying Classifier(-Free) Guidance From a Classifier-Centric Perspective — Xiaoming Zhao, Alexander G. Schwing, 2025
https://scholar.google.com/scholar?q=Studying+Classifier%28-Free%29+Guidance+From+a+Classifier-Centric+Perspective28. DEPT: Decoupled Embeddings for Pre-training Language Models — Alex Iacob, Lorenzo Sani, Meghdad Kurmanji, William F. Shen, Xinchi Qiu, Dongqi Cai, Yan Gao, Nicholas D. Lane, 2024
https://scholar.google.com/scholar?q=DEPT%3A+Decoupled+Embeddings+for+Pre-training+Language+Models29. AI Post Transformers: Generative Modeling via Drifting in One Step — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-generative-modeling-via-drifting-in-one-671da0.mp330. AI Post Transformers: VL-JEPA for Vision-Language Semantic Prediction — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-vl-jepa-for-vision-language-semantic-pre-69c9f4.mp331. AI Post Transformers: Why Transformers Fail at Counting — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-08-why-transformers-fail-at-counting-137924.mp332. AI Post Transformers: Mamba-3 for Efficient Sequence Modeling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-mamba-3-for-efficient-sequence-modeling-97a22a.mp3Interactive Visualization: ELF and Continuous Language Diffusion