AI Post Transformers arXiv 2512.13525 Visual Companion MoE Serving

JANUS for Scalable MoE Inference

This page treats JANUS as a systems diagram, not a paper summary. The visuals focus on why sparse MoE math still turns into dense operational pain: expert memory stays large, demand is skewed, and token-level latency targets force the runtime to make fast placement and scaling decisions.
PaperDisaggregating Attention and Experts for Scalable MoE Inference
AuthorsZhang, Wang, Zhao, Xiao, Yang et al.
Year2025
Known Source IDs2512.13525
Transcript Scanno additional DDDD.DDDDD IDs found
Headline Idea
Split It
Attention GPUs and expert GPUs stop scaling as one indivisible replica.
Latency Target
TPOT
Time per output token dominates the user-visible serving experience.
Reported Gain
4.7×
Paper framing versus monolithic serving on decode-heavy traces.
Dynamic Cost
-40%
Mocked from the episode narrative of independent pool elasticity.

One Model, Two Resource Shapes

Toggle between the monolithic replica view and JANUS’s split serving graph.

Attention / KV path Expert memory path Control + SLO edge

Pressure by Layer Family

Hover the bars. Decode-heavy traffic grows expert pressure faster than it grows attention pressure.

Monolith mode
The whole serving unit has to expand when either decode pressure or expert hotspots rise. That couples sparse routing pain to dense replica cost.

Router Skew Matrix

Hover cells. A few experts dominate dispatch volume even when each token only activates a small top-k set.

Why Sparse Compute Still Feels Expensive

JANUS is motivated by a mismatch between what the model activates and what the infrastructure must keep live.

Memory Expert weights stay resident even when most are idle on a given token.
Skew Hot experts create local queue buildup and p95 TPOT spikes.
Elasticity Attention and experts need different scaling granularity and timing.

Adaptive Two-Phase Communication

Step through how JANUS turns a tiny-message storm into fewer node-level transfers.

Transfer Overhead Heatmap

Hover the matrix. Each cell is mock traffic between an attention node and an expert node.

Pairwise mode
Many small transfers create bad overhead even if total bytes are not enormous.

Activated-Expert Pressure Map

Hover the grid. JANUS’s scheduler proxy is simple: avoid feeding more tokens into already-hot expert workers.

Scheduler Effect

Before and after balancing. Lower spread matters because MoE kernels are short and memory-bound.

Performance Under a Token SLO

Mock values follow the episode framing: big gains versus monolithic serving, smaller but still meaningful gains versus a disaggregated baseline.

Monolith DisAgg baseline JANUS

Latency Curve vs Offered Load

The dashed line is the target. JANUS pushes the knee of the curve to the right.

4.7×Throughput gain vs monolith in the paper framing.
3.3×Representative gain vs a weaker disaggregated comparison.
p95The point is tail control, not just average tokens per second.

Cost Surface

Hover cells. Lower values mean fewer normalized GPU-hours while staying inside the target band.

What the Workload Bias Hides

The transcript emphasized short prompts and long decodes. Shift the prompt mix to see the architectural bet get less one-sided.

SLO-Aware Independent Scaling

Change the pressure profile. Attention and expert pools can move independently instead of dragging each other around.

Expert Placement Matrix

Hover the matrix. Higher opacity means more residency and expected load on that worker-shard pair.

Decode spike
Expert GPUs rise fastest because token-by-token decode keeps hitting the routed MoE path.

Control Loop

JANUS behaves like a runtime controller: observe TPOT, estimate pressure, adjust pools and placement, repeat.

Interpretation Guardrail

The episode’s skepticism matters. JANUS looks strong as a systems prototype, but the evidence is still workload-shaped and baseline-sensitive.

Trace biasShort prompts make decode-side wins loom larger.
Cluster sizeFour nodes is real engineering, but not the final production stress test.
AblationsArchitecture, scheduler, and scaler each deserve cleaner isolation.

References

Compact source list for the episode and supporting comparisons.

JANUS
Disaggregating Attention and Experts for Scalable MoE Inference.
arXiv:2512.13525
Switch Transformers
Scaling to Trillion Parameter Models with Simple and Efficient Sparsity.
FastMoE
A Fast Mixture-of-Expert Training System.
DistServe
Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving.
MegaScale-Infer
Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism.
Pre-gated MoE
arXiv:2308.12066
Latency-Optimized Expert Placement
arXiv:2508.12851
Related AI Post Transformers Episodes
Batch-Aware Expert Routing, Splitwise, infrastructure trade-offs, optimization for serving, LAPS, DeepSeek-V4, and FengHuang.