The main visual question: where reasoning stops and real-time emission begins. Toggle the input mix to see how timestamped omnimodal streams feed the Thinker while the Talker keeps speech responsive.
Speech is harder than text because the model must decide both what to say and when to emit. Compare an older dual-track design against ARIA-style adaptive interleaving.
Why use Hybrid Attention + MoE instead of a plain dense transformer? The answer is not just “efficiency” — it is a moving tradeoff among memory, routing overhead, communication, and quality retention.
The episode’s caution: benchmark wins and demo polish are not the same thing. This tab separates areas described as benchmarked from areas described as capabilities or open questions.
Compact source list for the visuals and claims discussed in the episode.