TIDE asks a sharp architectural question: what if deeper transformer layers should be able to re-check the original token identity instead of reconstructing it from context alone? This page turns that argument into interactive diagrams, routing maps, collapse heatfields, and benchmark-style comparisons.
One path carries contextual meaning upward. The other path lets each layer re-access a token-linked memory sketch, so rare identifiers do not have to survive purely as residue inside the hidden state.
These matrices use mock similarity values to illustrate the paper’s concern: in similar contexts, distinct rare tokens can become harder to separate. Toggle the model to compare ordinary contextual-only behavior against TIDE-style identity reinjection.
Each layer produces a softmax over memory banks plus a null route. Step through token types and watch which banks light up, how much identity gets re-injected, and when the layer chooses to mostly ignore the side channel.
These benchmark-style charts use plausible mock values to express the paper’s reported shape: larger gains on rarer vocabulary slices, modest but broad task improvements, and stronger wins on jargon-heavy domains.
Compact links for the paper cluster behind this episode. arXiv links are direct where known.