TwinQuant challenges the core assumption behind SVDQuant-style 4-bit quantization: that a weight matrix's important information sits in a small, fixed set of directions. For LLMs it doesn't — it's spread across hundreds. TwinQuant learns the low-rank split itself via manifold optimization: a true orthogonal (Stiefel) rotation that folds into RMSNorm, plus a more flexible invertible (general linear) transform for per-layer residual handling.