Rotation-based quantizers pair one activation-side matrix P with its exact inverse on the weights. FreeAct lets different token types use different activation-side matrices while the weight side keeps a single static transform.
Both problem settings force heterogeneous tokens through one shared transform in the baseline methods.
A handful of channels carry values 10-100x larger than the rest (Dettmers et al., LLM.int8, 2022). Hover a cell to see its relative magnitude. Outlier columns wreck per-tensor resolution once you drop to 4 bits.
Fewer bits means fewer representable levels for the same outlier-skewed distribution.
P·P
P and P′ are slices of one P̃, indexed by token ID (the [MASK] id, or the image-token id) โ not separately stored matrices.
Figure 5: collapsing the transform to d/32 or d/64 still converges to the full-rank ceiling โ evidence rank-deficiency is real, not assumed. Table 2: dropping the clip threshold hurts every model, Dream's HumanEval hardest.
A theoretically sound idea, validated narrowly: one scale band, simulated quantization, no specialist baselines.
Named in related work, absent from Table 1.