Interactive visual companion

VL-JEPA for Vision-Language Semantic Prediction

arXiv: 2512.10942 Delong Chen et al. Meta FAIR · HKUST · Sorbonne · NYU Topic: semantic prediction vs token generation

What this page visualizes

The episode argues that VL-JEPA shifts supervision from next-token cross-entropy to conditional semantic embedding prediction. Instead of spelling out every answer during training, the model predicts an answer meaning vector from visual input + query, then decodes only when needed.
conditional prediction y-encoder teacher space selective decoding streaming video skepticism: target space quality

Episode lens

Not “text summary,” but a systems view: where computation goes, what the training target changes, and why low-latency video workloads may benefit if decoding becomes sparse and event-driven.
Trainable params
~50% less
Decode ops
2.85× fewer
Main target
latent Ŝy
Core caveat
teacher space

Conditional semantic prediction pipeline

VL-JEPA predicts the embedding of the answer given visual evidence and a query. The text-side target encoder defines the semantic training target; decoding becomes optional readout instead of the core supervision path.

Mode comparison

visual stream semantic target / predictor decode / output

Loss geometry: token sequence vs semantic space

Similarity matrix heatmap

Mock answer variants: semantic spaces cluster paraphrases, while exact-token training treats wording mismatch as a harder error signal.

Teacher-space dependence

The episode’s main skepticism: are gains from a fundamentally better objective, or from using a particularly strong text-side target embedding space?

Streaming video timeline with selective decoding

Compute budget allocation

Selective decoding moves more work into cheap continuous updates and fewer expensive serial readout events.

Window-by-window semantic change heatmap

Hover cells to inspect semantic drift intensity across a 3-minute mock stream. Red bands indicate moments where decode triggers are likely justified.

Mock benchmark profile

Claim strength radar

Strong where workloads are latent-first, weaker where exact wording, OCR fidelity, counting precision, or structured generation matter.

What the episode says is still missing

References