This page visualizes the episode’s core claim: chart understanding fails when models can read text but cannot bind marks, axes, legends, values, and reasoning into one coherent object. ChartNet’s bet is that large-scale, code-guided multimodal alignment can teach that binding more reliably.
The chart below mixes reported dataset scale from the episode with modality coverage. Bubble area tracks sample count; horizontal placement tracks how much of the chart object a dataset exposes beyond image + QA.
Semantic meaning sits in geometry. A bar is not text, a legend is not decoration, and an axis tick is part of the answer function.
It aligns image, code, table, summary, QA, reasoning, grounding, safety, and real-world subsets around the same chart instance.
Bigger synthetic families can still inherit narrow priors from their seed generator. Scale is not the same thing as semantic diversity.
Step through the pipeline to see where ChartNet gains fidelity and where error propagation can still creep in. The active step highlights its inputs, outputs, and failure mode.
Select a step to inspect the transformation.
The heatmaps show the episode’s core point: chart failure is usually a binding failure between visual marks, axis scales, legend keys, table values, and language reasoning.
Higher values mean the artifacts reinforce one another. Code↔table and axes↔values should be strong if the pipeline is internally consistent.
Hot cells mark where one local error spreads globally. Legend-color swaps and axis-scale mistakes are especially destructive.
Grounding is different from answer-only QA. It asks the model to point to where the evidence lives on the chart, not just output a sentence.
The bars illustrate the episode’s framing, not exact paper numbers. Switch between in-family evaluation and harder external transfer to see why the paper feels useful now but not fully settled.
ChartNet is credible supervised fine-tuning fuel for smaller chart-capable VLMs. It packages a much richer training object than earlier narrow benchmarks.
Robustness claims outrun evidence if evaluation stays too close to the synthetic universe or uses judge setups that may flatter reconstruction quality.
Real-world QA transfer on ChartQAPro, OpenCQA, and EvoChart; explicit dependence on TinyChart seeds; and deeper reasoning-diversity analysis rather than style diversity alone.