TFGN claims to clear four constraints on catastrophic forgetting at once — replay-free, task-free, LLM-scale, and free of external penalty terms — via a dense, input-conditioned overlay inside each transformer block. This page visualizes the architecture, the reported backward-transfer numbers, and the verification gaps the hosts flagged around the NDA-gated mechanism.
No published method has cleared all four simultaneously. TFGN's whole pitch is that it does — visualized below as a pipeline from token to backward-transfer metric.
Click to compare how a token is processed. TFGN's forward pass stays fully dense — every parameter fires on every token. Production MoE systems (Mixtral, DeepSeek-V3) route each token through a subset of experts instead.
Cross-domain gradients stay ≥99.59% L2-orthogonal in every TFGN condition, with no orthogonality loss term, projection operator, or task-boundary hook added — a structural signature, not a regularizer. Hover a cell for the pair value. Illustrative reconstruction of a 6-domain phase sequence consistent with the reported floor.
A closed-loop layer stacked on the same substrate: sensing, an internal predictive model, gating, and consolidation — all reading signals the network already computes.
Activation steering nudges the residual stream at inference. The plan vector reshapes the decoder's effective forward-pass operator directly.
Cosine similarity against the target domain's native operator across 30 source→target pairs. Toggle scale below. Illustrative per-pair reconstruction matching the reported 0.9996 / 0.9995 averages.
Zero BWT means no forgetting of prior domains. LoRA r=256 lands within ~15% of standard fine-tuning's BWT — parameter-efficiency and forgetting-resistance turn out to be unrelated properties.
Orthogonal writes aren't just neutral — training on Python alone drops held-out JS perplexity, a synergy nobody engineered directly.
Prompt: "The history of artificial intelligence began in",
asked right after a Python training phase.
TFGN is the only row passing all eight axes. Hover a row — the two closest competitors each fail at least two, mostly on domain count and covering only one training regime.
Architecture, source code, weights, and training recipe are gated behind a signed NDA pending patent prosecution. "Falsifiable in the operational sense" means reproducible only after signing.
Every BWT, HellaSwag, and orthogonality number is a single-seed point estimate — no confidence intervals. The 14x tighter-than-Standard-FT headline rests on this gap.