AI Post Transformers · Episode Companion

TwinQuant: Manifold-Constrained Low-Rank Decomposition for 4-Bit Quantization

arXiv:2606.01556 ICML 2026 Haodong Wang et al. · HKUST & Sun Yat-sen University June 1, 2026

TwinQuant challenges the core assumption behind SVDQuant-style 4-bit quantization: that a weight matrix's important information sits in a small, fixed set of directions. For LLMs it doesn't — it's spread across hundreds. TwinQuant learns the low-rank split itself via manifold optimization: a true orthogonal (Stiefel) rotation that folds into RMSNorm, plus a more flexible invertible (general linear) transform for per-layer residual handling.

References