1 00:00:01,000 --> 00:00:09,300 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. 2 00:00:09,300 --> 00:00:34,075 [Hal Turing] Today's paper is RelayS2S: Dual-Path Speculative Generation for Real-Time Dialogue, by Long Mai at Trinity College Dublin and Junli Liang at University College Dublin, from 2026. The question is whether a small speech model can say the first few words of a reply while a much stronger pipeline writes the rest, so you get a fast start without a big drop in answer quality. 3 00:00:34,075 --> 00:00:47,850 [Dr. Ada Shannon] The bet is that the opening of a reply is cheap to get right, and the content is expensive. If that holds, the slow model's latency mostly stops mattering. Whether it holds on real speech is what we need to check. 4 00:00:47,850 --> 00:00:52,475 [Hal Turing] Start with the target. How fast is fast enough for a voice system? 5 00:00:52,475 --> 00:01:13,900 [Dr. Ada Shannon] About 200 milliseconds. Levinson and Torreira, in a 2015 review, pulled together evidence that the gap between turns in human conversation averages around that, across many languages. Stivers and colleagues measured something similar in ten languages in 2009. Planning a sentence from scratch takes longer than 200 milliseconds, so people must be planning while the other person is still talking. 6 00:01:13,900 --> 00:01:17,550 [Hal Turing] And the usual voice assistant doesn't do that. 7 00:01:17,550 --> 00:01:45,200 [Dr. Ada Shannon] No. The standard design is a cascade. A speech recognizer turns audio into text. Here that's Whisper, from Radford and colleagues at OpenAI in 2022. A text LLM writes the reply, and a text-to-speech model speaks it. Every stage waits for the one before. The answers are good, because you can plug in the strongest text model available, and you can log and audit the text in the middle. But the wait grows with the LLM. A bigger model takes longer to produce its first tokens. 8 00:01:45,200 --> 00:01:49,300 [Hal Turing] So the smart option is slow. What's the fast option? 9 00:01:49,300 --> 00:02:19,925 [Dr. Ada Shannon] Full-duplex speech-to-speech models. One network listens and speaks at the same time, like a phone call, with no separate turn detector. Moshi, from Défossez and colleagues at Kyutai in 2024, is the best-known open example. It models the user's audio and its own audio as parallel streams. It also predicts text tokens alongside the audio, which Kyutai calls an inner monologue, to help the language quality. Its reported latency is around 160 to 200 milliseconds. It can say 'mm-hm', and it can stop when you cut in. 10 00:02:19,925 --> 00:02:23,350 [Hal Turing] So why isn't everyone just using that? 11 00:02:23,350 --> 00:02:45,150 [Dr. Ada Shannon] The answers are weaker. These models train on less and noisier conversation data than text LLMs, and they're usually smaller. Reasoning, facts and instruction following all suffer. There's also no text in the middle to inspect, which enterprises care about. One correction on scope: the paper never runs Moshi as a baseline. It uses its own small half-billion-parameter fast path as the stand-in for this camp. 12 00:02:45,150 --> 00:02:51,900 [Hal Turing] Here's the part that makes the design work for me. If the full answers are weak, why trust the opening? 13 00:02:51,900 --> 00:03:12,025 [Dr. Ada Shannon] That's the paper's empirical seed. The authors found that 82.5 to 95 percent of five-word prefixes from a weak speech model were judged contextually appropriate, even when the full answer wasn't. Think of 'Well, so...' or 'Sure, the thing is...'. People open that way and work out the content while talking. Note that it's a range, and it's their own analysis. 14 00:03:12,025 --> 00:03:15,250 [Hal Turing] Which sounds like speculative decoding. 15 00:03:15,250 --> 00:03:45,150 [Dr. Ada Shannon] It borrows the shape, and the difference matters more than the resemblance. In speculative decoding, a small draft model guesses several tokens and the big model checks them in one parallel pass. Leviathan, Kalman and Matias at Google, 'Fast Inference from Transformers via Speculative Decoding', 2023. Chen, Borgeaud, Irving and colleagues at DeepMind, 'Accelerating Large Language Model Decoding with Speculative Sampling', the same year. Both use rejection sampling, so the output distribution is exactly the big model's. It's lossless. 16 00:03:45,150 --> 00:04:04,750 [Dr. Ada Shannon] RelayS2S is not lossless. The draft comes from a different model family, in a different modality, with no shared vocabulary. The check is a learned classifier, so it can be wrong. And once the prefix is spoken, you can't roll it back. In text speculation a rejected token just vanishes. Here the listener has already heard it. 17 00:04:04,750 --> 00:04:09,125 [Hal Turing] So the only real control is the gate before you speak. 18 00:04:09,125 --> 00:04:17,050 [Dr. Ada Shannon] Right. Fallback is the other lever. If the gate doesn't trust the prefix, the system runs the plain cascade instead. 19 00:04:17,050 --> 00:04:22,500 [Hal Turing] There's a limit on the speed, too. Why five words? Why not start on word one? 20 00:04:22,500 --> 00:04:40,550 [Dr. Ada Shannon] Streaming text-to-speech needs a few words buffered before it can render audio. The paper uses five, which is about two seconds of speech. That sets a floor on when sound can start. It also gives the slow path about two seconds of runway to produce its continuation before the prefix runs out. 21 00:04:40,550 --> 00:04:45,425 [Hal Turing] Which raises a measurement question. Latency from when? To what? 22 00:04:45,425 --> 00:04:52,950 [Dr. Ada Shannon] Hold that. We'll come back to it once we've seen the paper's own definition. The clock-start turns out to matter a lot. 23 00:04:52,950 --> 00:05:01,800 [Hal Turing] One more bit of history. Speaking before you know the whole sentence, then repairing it mid-utterance, isn't new in dialogue research. 24 00:05:01,800 --> 00:05:30,400 [Dr. Ada Shannon] No. Skantze and Hjalmarsson at KTH, in 2010, worked on incremental speech generation, where a system starts talking and can revise itself as it goes. Hough followed in 2011 with self-repair in generation. RelayS2S's version is learned and neural, but the instinct is old: start, then fix it in the continuation. Filler-based tricks like 'hmm, let me think' hide latency too. The difference here is that the prefix carries real content, so the continuation has to fit what's already been said. 25 00:05:30,400 --> 00:05:36,100 [Hal Turing] So walk me through one turn. The user stops talking. What fires, in order? 26 00:05:36,100 --> 00:05:59,550 [Dr. Ada Shannon] Turn detection fires, and two paths start at once. The fast path is a small duplex model that drafts the first few words. A verifier then decides whether to commit that draft. If it commits, the slow path, Whisper plus the big LLM, continues from the exact words already spoken. If it rejects, the cascade answers the whole turn. Both paths write text into one shared buffer feeding one streaming TTS session. Nothing gets spliced in the audio domain. 27 00:05:59,550 --> 00:06:02,350 [Hal Turing] What's inside the fast path? 28 00:06:02,350 --> 00:06:24,925 [Dr. Ada Shannon] A streaming conformer encoder, then a causal convolution adapter. Together they emit one speech representation every 160 milliseconds. The backbone is Qwen2.5-0.5B, with five control tokens added to its vocabulary. [SIL] means keep listening. [BOC] starts a backchannel, like 'uh-huh'. [BOS] starts a real answer. [STP] means stop, because the user barged in. [EOS] ends the response. 29 00:06:24,925 --> 00:06:35,425 [Hal Turing] Here's what bothers me. That model only ticks every 160 milliseconds. Five words at one token per tick is already most of a second. 30 00:06:35,425 --> 00:07:05,525 [Dr. Ada Shannon] Right, and that's the problem forked generation solves. Tick-synchronous decoding would cost at least the chunk size times 160 milliseconds. For five to eight words, that's roughly 0.8 to 1.3 seconds before audio can start. So after [BOS], the model forks. The main stream keeps reading live audio every 160 milliseconds and keeps the authoritative state. The speculative stream copies the decoder state, zeroes out the speech embedding so it hears nothing more, and decodes text at full speed. 31 00:07:05,525 --> 00:07:09,100 [Hal Turing] So the drafting branch is deaf on purpose. 32 00:07:09,100 --> 00:07:17,700 [Dr. Ada Shannon] Yes. Only the main stream can notice a barge-in. If it emits [STP], playback halts and the rest of the draft is discarded. 33 00:07:17,700 --> 00:07:22,350 [Hal Turing] And the verifier decides whether the draft gets spoken at all. 34 00:07:22,350 --> 00:07:43,575 [Dr. Ada Shannon] It's a small classifier reading the draft's decoder hidden states, plus three calibration signals per token: entropy, log-probability, and the margin between the top two logits. It outputs a confidence score, and the default cutoff is 0.50. The labels came from a three-LLM majority vote on whether the prefix was sensible given the context. Humans labeled the validation and test sets. 35 00:07:43,575 --> 00:07:46,375 [Hal Turing] What does moving that cutoff do? 36 00:07:46,375 --> 00:08:04,825 [Dr. Ada Shannon] It trades coverage for safety. On synthetic dialogue, going from 0.50 to 0.75 rejects almost every bad prefix, 98 percent, but only 70 percent of good prefixes survive. More turns go to the slow cascade. Lowering it commits nearly everything. 37 00:08:04,825 --> 00:08:09,500 [Hal Turing] Then the slow path has to finish a sentence it didn't start. 38 00:08:09,500 --> 00:08:30,425 [Dr. Ada Shannon] It treats the spoken words as a hard constraint. Local LLMs get the prefix prefilled as the start of the assistant message. API models get it as already-spoken context with an instruction to emit only the continuation. It's also told to repair. If the prefix was wrong, it can say 'wait, I mean' and correct course in the continuation. 39 00:08:30,425 --> 00:08:32,275 [Hal Turing] Training data? 40 00:08:32,275 --> 00:08:56,575 [Dr. Ada Shannon] 194,000 conversations, 2,437 hours. About 104,000 are fully synthetic, rendered with CosyVoice2 and mixed with urban noise. The other 90,000 pair real human audio from GigaSpeech, Common Voice and Switchboard with GPT-5.4 assistant replies. They inject backchannels, interruptions and pauses, then train in three stages on two L40S GPUs: encoder, full duplex, then verifier. 41 00:08:56,575 --> 00:09:00,700 [Hal Turing] On turn-taking itself, does the little model behave? 42 00:09:00,700 --> 00:09:21,575 [Dr. Ada Shannon] On the core events, yes. Staying silent when it should scores 99.7 F1, and stopping on a barge-in scores 93.2. Start-speaking recall is 96.1 but precision is 78.9, so it rarely misses a chance to talk but sometimes starts when it shouldn't. Backchannels are weak at 48.4 F1, which the paper treats as auxiliary. 43 00:09:21,575 --> 00:09:26,725 [Hal Turing] Before numbers, pin down the clock. First-chunk latency starts when? 44 00:09:26,725 --> 00:09:46,350 [Dr. Ada Shannon] At the agent's decision to speak, and it ends when the first five response words are available to the TTS. For the cascade that's ASR time plus LLM time to five words. For a committed RelayS2S turn it's the small model's fifth word plus verifier time. It is not time to audible sound. 45 00:09:46,350 --> 00:09:49,925 [Hal Turing] Fine. What does it buy on the synthetic test set? 46 00:09:49,925 --> 00:10:08,975 [Dr. Ada Shannon] At the slowest one-in-ten turns, about 81 milliseconds, against a little over a second for the GPT-4.1 cascade at 1,006. And it's 81 for every back-end. The cascade's tail ranges from 418 milliseconds with a 0.5B model up to that 1,006. 47 00:10:08,975 --> 00:10:13,125 [Hal Turing] Why identical across back-ends? That looks too clean. 48 00:10:13,125 --> 00:10:29,351 [Dr. Ada Shannon] Because the slow path isn't in the timed path for committed turns. The fallback rate there is 8.1 percent, under 10, so the ninetieth percentile still lands on a committed turn. The moment fallback exceeds 10 percent, that number stops being true. 49 00:10:29,351 --> 00:10:31,651 [Hal Turing] Does quality survive? 50 00:10:31,651 --> 00:10:53,277 [Dr. Ada Shannon] Only with the verifier. On real voice with GPT-4.1, without the gate about 14 answers in 100 are noticeably bad at five words. With it, about 11 in 100, against roughly 10 for the plain cascade. Gating removes most of the damage, not all. The quality scorer is Gemini-3-Flash on text only, and the intervals overlap, with no significance test reported. 51 00:10:53,277 --> 00:10:57,627 [Hal Turing] Why not use the three-word prefix, then? It scored best. 52 00:10:57,627 --> 00:11:18,452 [Dr. Ada Shannon] Stalls. The relay margin is the prefix's audio duration minus how long the slow path takes to deliver its first continuation chunk. Negative means a silent gap. With GPT-4.1 at three words, 27.5 percent of committed turns stall. At five words, 1.9 percent. Five words is about two seconds of runway. 53 00:11:18,452 --> 00:11:21,552 [Hal Turing] Is this the same idea as KAME? 54 00:11:21,552 --> 00:11:42,852 [Dr. Ada Shannon] KAME, from Kuroki and colleagues, 2026, is a tandem design. A speech model talks first and a back-end LLM corrects it, with no gate. The ungated ablation is the closest controlled analogue. The paper also ran KAME's released checkpoint and got 2.49 average against 4.81. But that's a different checkpoint, data and task, so it isn't evidence against the architecture. 55 00:11:42,852 --> 00:11:53,227 [Hal Turing] The part I can't shake is the clock. Everything so far is first-chunk latency. What's the time from the user stopping to the listener hearing something? 56 00:11:53,227 --> 00:12:14,527 [Dr. Ada Shannon] The paper doesn't measure it, for any system. It has no end-to-end audible number, with ElevenLabs Flash or anything else. Three terms are missing. First, the detector that fires the turn. RelayS2S reuses its own start-speaking decision, and the cascade's end-of-turn detection isn't timed. Second, the TTS time to first audio. Third, network. I won't guess figures for those. 57 00:12:14,527 --> 00:12:19,727 [Hal Turing] If those terms are the same for both systems, what happens to the gain? 58 00:12:19,727 --> 00:12:45,752 [Dr. Ada Shannon] The absolute gap stays, but the relative gain shrinks. On real speech with GPT-4.1 the average goes from 748 to 269 milliseconds, a 479 millisecond cut. That's a smaller fraction of a total that also includes detection, rendering and network. The terms may not be equal either. The cascade's API latency is measured at the client, network included. The committed fast path runs locally on one GPU. So the two sides aren't timed on the same footing. 59 00:12:45,752 --> 00:12:50,202 [Hal Turing] And audio can't start before five words exist anyway. 60 00:12:50,202 --> 00:13:10,527 [Dr. Ada Shannon] Right, and some of those starts are false, given the start-speaking precision. Now the real-voice tail. The abstract leads with the synthetic 81 milliseconds and the real-voice average. With GPT-4.1 on real speech, the slowest tenth of turns goes from 823 to 655 milliseconds. That's about a fifth faster, not a tenth of the wait. 61 00:13:10,527 --> 00:13:14,927 [Hal Turing] Because 29.3% of turns fall back. 62 00:13:14,927 --> 00:13:37,727 [Dr. Ada Shannon] Yes. The average rewards the many easy turns, which really do get faster. The tail measures the hard turns, and those barely move. The 81 is what you get when fallback stays under ten percent, which happened on synthetic audio. I wouldn't expect it from real speech. There's also a curation issue. The real set is real user audio paired with synthetic assistant replies, two turns long, English only. Messy multi-turn speech is exactly what would stress the draft. 63 00:13:37,727 --> 00:13:41,252 [Hal Turing] Why do the fallbacks happen? Bad transcription? 64 00:13:41,252 --> 00:14:01,702 [Dr. Ada Shannon] Partly. Word error on fallback turns is 8.7% against 2.8% on committed ones. But over half of fallback turns have perfect transcripts. They involve rare names and technical terms that a 0.5 billion parameter model gets wrong. So the turns that fall back are the harder ones, and those are the ones users notice. 65 00:14:01,702 --> 00:14:10,477 [Hal Turing] Then there's irreversibility. Lossless speculative decoding can reject a draft for free. Here the audio is already out. 66 00:14:10,477 --> 00:14:31,302 [Dr. Ada Shannon] That's the core difference. Those papers guarantee the output distribution is unchanged. Here there's no rollback, so the paper repairs in the continuation, a spoken 'wait, I mean'. Repair recovered 63.2% of bad committed prefixes, but that's 48 of 76 cases, a small sample. Roughly 2% of committed turns fail on both paths. 67 00:14:31,302 --> 00:14:33,452 [Hal Turing] Who does the judging? 68 00:14:33,452 --> 00:14:59,052 [Dr. Ada Shannon] Gemini-3-Flash, on text only, agreeing with humans at kappa 0.70. No one listened. There's no user study on perceived responsiveness, handoff prosody or how an in-flight continuation behaves when the user barges in. The verifier labels also come from language models asking whether a prefix is sensible, which isn't quite what a listener hears as wrong. The 1.1 point quality gap to the cascade comes with overlapping intervals and no significance test. I wouldn't read it as a tie or a loss. 69 00:14:59,052 --> 00:15:04,502 [Hal Turing] What about baselines? A streaming cascade seems like a stronger opponent. 70 00:15:04,502 --> 00:15:26,302 [Dr. Ada Shannon] The baseline is sequential Whisper plus an LLM. Streaming cascades like PredGen, from Li and Grover in 2025, start the LLM on partial transcripts, and ChipChat from Likhomanenko and colleagues in 2026 is an optimized cascade. The paper calls them complementary and never runs them. Moshi isn't run either, and the KAME comparison uses a released checkpoint on different data. The paper says so itself. 71 00:15:26,302 --> 00:15:28,452 [Hal Turing] Is it still useful? 72 00:15:28,452 --> 00:15:49,177 [Dr. Ada Shannon] For teams with a cascade on a large or hosted LLM, it's a plausible add-on. The text handoff doesn't touch either component. But you'd train the fast path and verifier yourself, and it's English-specific, since the five-word budget is. The authors advise against medical, legal and mental health use, which fits: a 0.5B model is saying the first words. 73 00:15:49,177 --> 00:15:52,077 [Hal Turing] What would settle the open questions? 74 00:15:52,077 --> 00:16:02,202 [Dr. Ada Shannon] Three things: a real-voice end-to-end measurement including TTS and network, a P90 on real speech against the cascade, and a listening study. 75 00:16:02,202 --> 00:16:22,377 [Hal Turing] So, takeaways. A small speech model can speak the first words of most real turns, and a stronger model can finish them. The evidence supports a large average speedup on curated English audio. It doesn't support a fast worst case, or any claim about audible latency. Thanks for listening, everyone.