RelayS2S: Dual-Path Speculative Generation for Real-Time Dialogue

Long Mai (Trinity College Dublin), Junli Liang (University College Dublin) · 2026
arXiv 2603.23346 · small speech model says the first 5 words, big pipeline writes the rest

One turn, step by step

Design space (qualitative, illustrative placement)

Positions are schematic, not measured. Moshi and PredGen/ChipChat were not run as baselines.

Why fork? Tick-synchronous vs forked decoding

Words in prefix:

Fast path emits one speech frame per 160 ms. Forked branch is deaf and decodes text at full speed (mock ≈ 12 ms/word).

Control tokens

Hover a token.

Fast-path turn-taking F1 (%)

Verifier gate: coverage vs safety

Threshold:

Anchored: at 0.75, 98% of bad prefixes rejected, 70% of good kept. Other points are illustrative curves.

Verifier inputs per draft token (mock)

Hover a cell. Hidden states feed the classifier alongside these signals.

Prefix usefulness seed

82.5–95% of 5-word prefixes from a weak speech model were judged contextually appropriate, even when the full answer was not. Authors' own analysis.

First-chunk latency (ms)

Answer quality: bad answers per 100 (GPT-4.1, real voice)

Intervals overlap; no significance test reported. Judge: Gemini-3-Flash on text only.

Fallback rate vs the P90 guarantee

The 81 ms P90 holds only while fallback stays under 10%.

Why five words: stalls vs runway

Prefix length:

Stall rate at 3 and 5 words is from the paper (27.5%, 1.9%); 4, 6, 7 are interpolated for display. ≈ 2 s of speech at 5 words.

Where does the latency clock start and stop?

Not lossless: speculative decoding vs RelayS2S

Why turns fall back

Over half of fallback turns have perfect transcripts: rare names and technical terms trip the 0.5B model. Repair recovered 48 of 76 bad committed prefixes (small sample).

References

  1. RelayS2S — Mai, Liang, 2026
  2. Fast Inference from Transformers via Speculative Decoding — Leviathan, Kalman, Matias, 2023
  3. Speculative Sampling — Chen, Borgeaud, Irving et al., 2023
  4. Medusa — Cai et al., 2024
  5. Moshi — Défossez et al., 2024
  6. dGSLM — Nguyen et al., 2022/23
  7. LSLM — Ma et al., 2024
  8. Timing in turn-taking — Levinson, Torreira, 2015
  9. Whisper — Radford et al., 2022
  10. Talking Turns — Arora et al., 2025
  11. Full-Duplex-Bench — Lin et al., 2025
  12. Universals in turn-taking — Stivers et al., 2009
  13. Synchronous LLMs as Full-Duplex Agents — Veluri et al., 2024
  14. LLaMA-Omni — Fang et al., 2024
  15. Freeze-Omni — Wang et al., 2025
  16. KAME — Kuroki et al., 2026
  17. DDTSR — Liu et al., 2026
  18. PredGen — Li, Grover, 2025
  19. ChipChat — Likhomanenko et al., 2026
  20. SpeakStream — Bai et al., 2025
  21. Selective Classification for Deep Neural Networks — Geifman, El-Yaniv, 2017