AI Post Transformers • SVG-first visual companion
A frozen ChatGLM reads image, video, and speech through learned bridges. This page visualizes the wiring, the compression bottlenecks, the staged training recipe, and why the method feels distinctly 2023 when compared with later end-to-end multimodal systems.
Compact arXiv links for the papers doing most of the work in the episode.