AI Post Transformers • Interactive Visualization

Trace Rewriting Against Unauthorized LLM Distillation

A visual companion focused on one question: can a model rewrite its own reasoning traces so humans still benefit, while a copied student learns less from the same outputs? The graphics below compare the defense pipeline, the rewrite signal itself, reported utility-vs-degradation tradeoffs, and the unresolved question of whether this is true poisoning or mostly style mismatch from a stronger ghostwriter.

arXiv:2602.15143 arXiv:2502.11598 Live Viz URL Transcript IDs found: 2602.15143, 2502.11598

Black-Box Distillation vs Online Rewrite Defense

The provider cannot stop queries. The defense acts at response time, rewriting the chain-of-thought before it leaves the API.
teacher utility path attacker training path watermark / rewrite channel

Defense Families

Prompt-based trace rewriting is easier to deploy. Gradient-based rewriting is closer to classic poisoning logic.

Threat Pressure Map

Reasoning tasks create the richest leakage surface because intermediate traces carry process, not just answers.

Step-by-Step Trace Surgery

Toggle the displayed trace to see how a legible human explanation can be re-authored into a form that preserves the answer but alters the training signal.

Reasoning Token Heatmap

Mock salience map across 12 reasoning steps. Hover cells to inspect how semantic preservation and student learnability pull in opposite directions.

Watermark Carrier Pattern

Structural fingerprints can live above the token level: ordering, clause rhythm, trace segmentation, and recurring rewrite motifs.

Utility vs Student Degradation

Interactive mode switch compares reasoning-heavy tasks and shows the paper’s central tradeoff: keep the teacher strong while reducing what a 1B-3B student absorbs.

Rewrite Model Scale Sensitivity

The anti-distillation effect strengthens as the rewriting model gets stronger, which is why the “security by ghostwriter” interpretation matters.

Detection Snapshot

Watermark detection looks cleaner than token-statistical baselines, but the student-dependent detection rate remains uneven.
Few-shot verification
K = 5
highlighted in discussion
Example true detect
0.55
reported weak case on 3B student
False alarms
~0
claimed near-zero in favored setup
Open question
Removal Attacks
scrubbing and stronger distillers still under-tested
Numbers here are illustrative but aligned to the episode’s framing: detection is promising, not settled.

Poisoning or Better Editing?

The core ambiguity is causal. If teacher accuracy rises after rewriting, maybe the rewriter is fixing reasoning quality rather than injecting a harmful training signal.

Ablation Checklist

The next experiments need to separate correction, style mismatch, and true anti-learning effects under stronger student training regimes.

Attacker Upgrade Ladder

Benchmarked thieves are small and modestly trained. Real attackers can mix full finetuning, continued pretraining, answer-only corpora, and multi-stage pipelines.

Key Papers

Compact references tied to the visuals above.
Ding et al. (2025), Information-Preserving Reformulation of Reasoning Traces for Antidistillation.
Li et al. (2025), DOGe: Defensive Output Generation for LLM Protection Against Knowledge Distillation.
Jiang (2026), DistillGuard: Evaluating Defenses Against LLM Knowledge Distillation.

Related Context

Watermarking, poisoning, and prior episode context.
Goldblum et al. (2020), Dataset Security for Machine Learning: Data Poisoning, Backdoor Attacks, and Defenses.
Wan et al. (2023), Poisoning Language Models During Instruction Tuning.