AI Post Transformers • SVG-first visual companion

X-LLM: Treating Multimodalities as Foreign Languages

A frozen ChatGLM reads image, video, and speech through learned bridges. This page visualizes the wiring, the compression bottlenecks, the staged training recipe, and why the method feels distinctly 2023 when compared with later end-to-end multimodal systems.

arXiv 2305.04160 Posted May 7, 2023 Revised May 22, 2023 Backbone ChatGLM Transcript regex scan: no extra arXiv IDs
3
modality bridges feeding one dialogue core
256 -> 32
illustrative visual token squeeze before ChatGLM
~10k
joint multimodal instruction examples in the final stage
84.5%
paper-reported relative score on a small Chinese synthetic benchmark
References

Core paper trail

Compact arXiv links for the papers doing most of the work in the episode.

Core cited papers