A Lightmatter-authored paper claims 3D-integrated photonics can cut inference prefill latency by up to 8.5x for long-context Mixture-of-Experts workloads. This companion visualizes the physics behind the claim, the simulator that produced the numbers, and where the headline multiplier does and doesn't hold up.
Copper SerDes must exit through the die's physical perimeter. Lightmatter's Passage platform stacks the electronic die on a photonic layer ~100µm below it, so light can exit from anywhere across the whole chip surface.
Per Lightmatter's published Passage spec sheet — not independently measured.
Tighter pitch, and one bidirectional fiber replaces four copper wires — roughly an 8x radix jump.
At 224 Gbps/lane, passive copper is reliable to ~1m — the reason scale-up pods get stuck at a single rack and everything beyond forces multi-rack scale-out.
Prefill chews the entire input at once to build the KV cache (compute-bound). Decode generates one token at a time, pulling full model weights from HBM every step (memory-bandwidth-bound).
Same model, opposite bottleneck. This is why an interconnect win in one phase doesn't automatically transfer to the other.
MoE adds a router that sends each token to a handful of specialist experts out of potentially hundreds — tokens get shuffled across devices in an "all-to-all," stacking communication on top of raw compute.
Each cell is the maximum multiplier found anywhere in the device/batch sweep — not a typical operating point. Hover a cell for detail. R4 is a speculative quad-die config extrapolated from a paywalled report about an unshipped chip.
The spread itself is informative: gains scale with how communication-bound a regime is.
Optical scale-up speeds prefill dispatch — but that feeds the single decode worker faster and in tighter bursts, growing its steady-state batch.
TTFT gains and TPOT regression roughly cancel — the user experiences a shifted bottleneck, not a clear win.
3 of 4 authors are Lightmatter employees; the paper promotes Lightmatter's own Passage platform.
Every reported multiplier is the maximum across the device × batch sweep. An operator's actual fixed configuration lands well below the reported line.