Hyper-Connections replace the fixed residual connection with a learned wiring matrix across n parallel streams. It initializes as exactly Pre-Norm, so it strictly generalizes the Pre-Norm / Post-Norm seesaw — at the cost of real activation memory the paper doesn't headline. This page visualizes the mechanism, the reported results, and where the evidence gets thin.