FreeAct: Rethinking One-to-One Transforms for LLM Quantization

Xiaohao Liu, Xiaobo Xia, Manyi Zhang, Ji-Fu Li, Xianzhi Yu, Fei Shen, Xiu Su, See-Kiong Ng, Tat-Seng Chua · 2026
National University of Singapore · Huawei Technology · Central South University
arXiv:2603.01776 W4A4 Quantization Diffusion LLMs + MLLMs ๐ŸŽง AI Post Transformers

From one-to-one to many-to-one

Rotation-based quantizers pair one activation-side matrix P with its exact inverse on the weights. FreeAct lets different token types use different activation-side matrices while the weight side keeps a single static transform.

Activation path Weight path Token-type split

Two flavors of the same mismatch

Both problem settings force heterogeneous tokens through one shared transform in the baseline methods.

Why activations are the hard part

A handful of channels carry values 10-100x larger than the rest (Dettmers et al., LLM.int8, 2022). Hover a cell to see its relative magnitude. Outlier columns wreck per-tensor resolution once you drop to 4 bits.

Bit-width squeeze

Fewer bits means fewer representable levels for the same outlier-skewed distribution.

Proposition 1: rank-deficiency buys wiggle room

P·P⁻¹ = I forces one-to-one. If activations are rank-deficient, a whole family of matrices satisfies the equivalence, not just the exact inverse. That slack lets each token type get a private subspace stacked on a shared basis U, zero-padded so types never leak into each other. On the weight side everything stitches into one static P̃.

Shared basis U Type A private subspace Type B private subspace Zero-padded

Cost: three lines of code, zero extra memory

P and P′ are slices of one P̃, indexed by token ID (the [MASK] id, or the image-token id) โ€” not separately stored matrices.

W4A4 accuracy across four models

+5.3%
max relative gain vs FlatQuant (avg)
4
models: LLaDA, Dream, Qwen2.5-VL, InternVL2.5
2
benchmarks where FlatQuant still wins

Ablations back the theory

Figure 5: collapsing the transform to d/32 or d/64 still converges to the full-rank ceiling โ€” evidence rank-deficiency is real, not assumed. Table 2: dropping the clip threshold hurts every model, Dream's HumanEval hardest.

What Table 1 doesn't show

A theoretically sound idea, validated narrowly: one scale band, simulated quantization, no specialist baselines.

Domain-specific competitors never tabled

Named in related work, absent from Table 1.

References