AI Post Transformers · Episode Companion

Deep Native Structural Reasoning for Proteins, Molecules, and Crystals

SciReasoner, a 29-author foundation model from Shanghai AI Laboratory and collaborators, reasons natively over protein, molecule, and crystal structures using domain-specific tokenizers instead of flattened text — compressing a molecule's SMILES string from 31 meaningless BPE fragments into 14 structurally meaningful tokens.

29 Authors Shanghai AI Laboratory Qwen3-14B Backbone 86 Benchmarks · SOTA on 67 arXiv:2607.07708 ↗

One Vocabulary, Three Scientific Domains

Proteins, molecules, and crystals are tokenized by domain-specific encoders into a shared discrete vocabulary, then fed through a single 14B-parameter backbone that reasons over addressable structural evidence instead of a black-box score.

Why Not Just Use Text?

A protein's function depends on its 3D fold, not just its amino-acid letters. Standard sub-word tokenizers treat structure as a character string and slice through chemically meaningful units — rings, stereocenters, backbone geometry — wherever compression happens to land.

The Core Claim

Can one model represent proteins, molecules, and crystals as native discrete tokens, so it reasons over evidence it can point to — instead of producing a score with no way to check its work?

31 Meaningless Fragments vs. 14 Structural Tokens

The same molecule, tokenized two ways. A standard BPE tokenizer (the kind Qwen or GPT use) shatters the SMILES string into fragments that mean nothing in isolation. SciReasoner's structure-aware tokenizer compresses it into fewer tokens, each preserving real chemical meaning. Hover any cell for detail.

BPE Sub-word Tokenizer — 31 fragments

Structure-Aware Tokenizer — 14 tokens

Token Count Comparison

Five-Stage Training Pipeline

Three pretraining stages teach the 14B backbone to read structural tokens; two post-training stages teach it to reason with them. Click a stage number.

Performance Against Baselines

SOTA Coverage

SciReasoner is state-of-the-art on 67 of 86 benchmarks spanning proteins, DNA, RNA, molecules, and crystals. Materials formation-energy prediction hits R² = 0.895 against ground truth.

Five-Product Retrosynthesis Showcase

Top-3 predictions on five USPTO-50K products. SciReasoner's rank-one correctly identifies a copper-catalyzed azide-alkyne cycloaddition that Opus-4.7 mistakes for a different coupling reaction.

Ablation: Pull the Structural Tokens Out

The cleanest evidence in the paper. Strip structural tokens and scores drop across every domain — and the reasoning trace itself degrades from citing real structural evidence to guessing from superficial cues.

MAD/MAE Ratio: One Outlier Too Many

Most materials benchmarks show a believable 2×–5× ratio. Two jump past 20× and 100× — log scale, outliers in orange.

Human Eval: 98% Tie-or-Better

~1,800 case-judgments vs. DeepSeek-V4-Pro, explicitly called a pilot. Inter-rater agreement is not reported, and most automated reasoning-quality scores come from GPT-5.5 — itself one of the benchmarked models.

Open Evaluation Questions

References