This page treats JANUS as a systems diagram, not a paper summary. The visuals focus on why sparse MoE math still turns into dense operational pain: expert memory stays large, demand is skewed, and token-level latency targets force the runtime to make fast placement and scaling decisions.
PaperDisaggregating Attention and Experts for Scalable MoE Inference
AuthorsZhang, Wang, Zhao, Xiao, Yang et al.
Year2025
Known Source IDs2512.13525
Transcript Scanno additional DDDD.DDDDD IDs found
Headline Idea
Split It
Attention GPUs and expert GPUs stop scaling as one indivisible replica.
Latency Target
TPOT
Time per output token dominates the user-visible serving experience.
Reported Gain
4.7×
Paper framing versus monolithic serving on decode-heavy traces.
Dynamic Cost
-40%
Mocked from the episode narrative of independent pool elasticity.
One Model, Two Resource Shapes
Toggle between the monolithic replica view and JANUS’s split serving graph.
Hover the bars. Decode-heavy traffic grows expert pressure faster than it grows attention pressure.
Monolith mode
The whole serving unit has to expand when either decode pressure or expert hotspots rise. That couples sparse routing pain to dense replica cost.
Router Skew Matrix
Hover cells. A few experts dominate dispatch volume even when each token only activates a small top-k set.
Why Sparse Compute Still Feels Expensive
JANUS is motivated by a mismatch between what the model activates and what the infrastructure must keep live.
MemoryExpert weights stay resident even when most are idle on a given token.
SkewHot experts create local queue buildup and p95 TPOT spikes.
ElasticityAttention and experts need different scaling granularity and timing.
Adaptive Two-Phase Communication
Step through how JANUS turns a tiny-message storm into fewer node-level transfers.
Transfer Overhead Heatmap
Hover the matrix. Each cell is mock traffic between an attention node and an expert node.
Pairwise mode
Many small transfers create bad overhead even if total bytes are not enormous.
Activated-Expert Pressure Map
Hover the grid. JANUS’s scheduler proxy is simple: avoid feeding more tokens into already-hot expert workers.
Scheduler Effect
Before and after balancing. Lower spread matters because MoE kernels are short and memory-bound.
Performance Under a Token SLO
Mock values follow the episode framing: big gains versus monolithic serving, smaller but still meaningful gains versus a disaggregated baseline.
MonolithDisAgg baselineJANUS
Latency Curve vs Offered Load
The dashed line is the target. JANUS pushes the knee of the curve to the right.
4.7×Throughput gain vs monolith in the paper framing.
3.3×Representative gain vs a weaker disaggregated comparison.
p95The point is tail control, not just average tokens per second.
Cost Surface
Hover cells. Lower values mean fewer normalized GPU-hours while staying inside the target band.
What the Workload Bias Hides
The transcript emphasized short prompts and long decodes. Shift the prompt mix to see the architectural bet get less one-sided.
SLO-Aware Independent Scaling
Change the pressure profile. Attention and expert pools can move independently instead of dragging each other around.
Expert Placement Matrix
Hover the matrix. Higher opacity means more residency and expected load on that worker-shard pair.
Decode spike
Expert GPUs rise fastest because token-by-token decode keeps hitting the routed MoE path.
Control Loop
JANUS behaves like a runtime controller: observe TPOT, estimate pressure, adjust pools and placement, repeat.
Interpretation Guardrail
The episode’s skepticism matters. JANUS looks strong as a systems prototype, but the evidence is still workload-shaped and baseline-sensitive.
Trace biasShort prompts make decode-side wins loom larger.
Cluster sizeFour nodes is real engineering, but not the final production stress test.
AblationsArchitecture, scheduler, and scaler each deserve cleaner isolation.
References
Compact source list for the episode and supporting comparisons.
JANUS Disaggregating Attention and Experts for Scalable MoE Inference. arXiv:2512.13525
Switch Transformers Scaling to Trillion Parameter Models with Simple and Efficient Sparsity.
FastMoE A Fast Mixture-of-Expert Training System.
DistServe Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving.
MegaScale-Infer Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism.
Related AI Post Transformers Episodes Batch-Aware Expert Routing, Splitwise, infrastructure trade-offs, optimization for serving, LAPS, DeepSeek-V4, and FengHuang.