TFGN: Replay-Free, Task-Free Continual Pre-Training at Scale

arXiv:2605.15053 Anurup Ganguli · 2026 Independent Researcher Solo-authored preprint

TFGN claims to clear four constraints on catastrophic forgetting at once — replay-free, task-free, LLM-scale, and free of external penalty terms — via a dense, input-conditioned overlay inside each transformer block. This page visualizes the architecture, the reported backward-transfer numbers, and the verification gaps the hosts flagged around the NDA-gated mechanism.

The Four-Constraint Problem

No published method has cleared all four simultaneously. TFGN's whole pitch is that it does — visualized below as a pipeline from token to backward-transfer metric.

Dense Overlay vs. Sparse Mixture-of-Experts

Click to compare how a token is processed. TFGN's forward pass stays fully dense — every parameter fires on every token. Production MoE systems (Mixtral, DeepSeek-V3) route each token through a subset of experts instead.

Cross-Domain Gradient Orthogonality

Cross-domain gradients stay ≥99.59% L2-orthogonal in every TFGN condition, with no orthogonality loss term, projection operator, or task-boundary hook added — a structural signature, not a regularizer. Hover a cell for the pair value. Illustrative reconstruction of a 6-domain phase sequence consistent with the reported floor.

Extension A — System A / System M Loop

A closed-loop layer stacked on the same substrate: sensing, an internal predictive model, gating, and consolidation — all reading signals the network already computes.

Extension B — Plan Vector vs. Activation Steering

Activation steering nudges the residual stream at inference. The plan vector reshapes the decoder's effective forward-pass operator directly.

Operator Reshape Fidelity

Cosine similarity against the target domain's native operator across 30 source→target pairs. Toggle scale below. Illustrative per-pair reconstruction matching the reported 0.9996 / 0.9995 averages.

Sub-task Injection Is Patchier

Backward Transfer (BWT) — Lower Magnitude Is Better

Zero BWT means no forgetting of prior domains. LoRA r=256 lands within ~15% of standard fine-tuning's BWT — parameter-efficiency and forgetting-resistance turn out to be unrelated properties.

Held-out JavaScript Perplexity, After Python-only Training

Orthogonal writes aren't just neutral — training on Python alone drops held-out JS perplexity, a synergy nobody engineered directly.

Same Prompt, Two Fine-Tuning Regimes

Prompt: "The history of artificial intelligence began in", asked right after a Python training phase.

The Eight-Axis Scorecard (Table 3)

TFGN is the only row passing all eight axes. Hover a row — the two closest competitors each fail at least two, mostly on domain count and covering only one training regime.

Reproduction Chain — Gated Behind an NDA

Architecture, source code, weights, and training recipe are gated behind a signed NDA pending patent prosecution. "Falsifiable in the operational sense" means reproducible only after signing.

Single-Seed Point Estimates

Every BWT, HellaSwag, and orthogonality number is a single-seed point estimate — no confidence intervals. The 14x tighter-than-Standard-FT headline rests on this gap.

What's Actually Verified

References