1 00:00:01,000 --> 00:00:43,625 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into a paper called "Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning." First author Chen Tang, et al. — and I mean that literally, this is a 29-author paper — led out of Shanghai Artificial Intelligence Laboratory, with collaborators from CUHK, Shanghai Jiao Tong, Fudan, Oxford, Stanford, and a bunch of others. It went up on arXiv on July 8th, 2026. The model they built is called SciReasoner, and Ada, when you first sent this one over you were pretty fired up. 2 00:00:43,625 --> 00:01:24,650 [Dr. Ada Shannon] I was, and honestly it's the formalism that got me. Every few months somebody claims their model can 'reason about science,' and what that usually means is a language model that memorized a lot of textbook prose. This paper actually sits down and asks: what does it even mean to reason using a structure — a protein fold, a molecule's 3D shape, a crystal lattice — as opposed to reasoning about a description of that structure? That's a real representational question, not a marketing one. The core question the paper is chasing, in one sentence: can a single model represent proteins, small molecules, and crystals as native discrete tokens, so it reasons over addressable structural evidence, instead of just producing a black-box score with no way to check its work? 3 00:01:24,650 --> 00:01:44,725 [Hal Turing] Okay, so let's set the table, because I think a lot of our listeners will have a mental model of 'foundation model' that's very text-shaped — think GPT, think a big transformer trained on internet text. What's a scientific foundation model, and why can't we just point one of those at a protein sequence and call it done? 4 00:01:44,725 --> 00:02:26,375 [Dr. Ada Shannon] So a scientific foundation model is the same idea as a language foundation model — one big pretrained network meant to generalize across many tasks — except the tasks span proteins, DNA, RNA, small molecules, and inorganic crystals instead of, say, email drafting and code completion. And your question is exactly the paper's opening gripe. You can point a text LLM at a protein and it'll happily process the amino acid letters. But a protein's function depends on its 3D fold, not just its sequence — the letters are one-dimensional, the biology is not. Same for a molecule's stereochemistry or a crystal's periodic bonding. Cast those as plain text and the model is reasoning about a lossy shadow of the actual structure. 5 00:02:26,375 --> 00:02:38,675 [Hal Turing] Right, and that's basically what structure-property prediction is as a field, right? Take some object's structure and infer a property from it — function, reactivity, band gap, whatever? 6 00:02:38,675 --> 00:03:10,375 [Dr. Ada Shannon] Exactly — infer a protein's function, a molecule's reactivity, a crystal's band gap, from its spatial and chemical organization. And here's the concrete problem the paper shows: feed a molecule's SMILES string through a standard sub-word tokenizer — the same byte-pair-encoding style tokenizer Qwen or GPT use — and it shatters into something like 31 fragments, some of them meaningless in isolation, because BPE doesn't know what a ring system or a stereocenter is. It's just compressing common character sequences. So a bond that matters chemically gets sliced right through the middle. 7 00:03:10,375 --> 00:03:21,975 [Hal Turing] Oh wait wait wait — so it's basically the tokenizer equivalent of hyphenating a word wherever your printer runs out of margin, without caring if you split it mid-syllable? 8 00:03:21,975 --> 00:04:15,575 [Dr. Ada Shannon] That's not a bad analogy, actually — except the 'syllable' here is a stereocenter or a ring closure, and getting it wrong doesn't just look ugly, it can flip the chemistry the model thinks it's looking at. So SciReasoner's fix is what the paper calls a structure-aware vocabulary: instead of BPE, they use domain-specific tokenizers — Foldseek's 3Di alphabet for protein backbone geometry, SLICES for crystal topology, ConfSeq for 3D molecular conformations — that turn coordinates and topologies into discrete tokens which actually preserve structural meaning. Same molecule through their tokenizer compresses to 14 tokens instead of 31, and every one of those tokens still means something physically. That's native structural reasoning: those tokens aren't just inputs the model reads once and discards, they're addressable evidence the model can cite and check while it's generating its reasoning trace. 9 00:04:15,575 --> 00:04:58,201 [Hal Turing] And retrosynthesis is one of the three big test domains here, alongside proteins and crystals — with DNA and RNA getting a lighter touch. For anyone who hasn't done organic chemistry since undergrad, retrosynthesis is working backward from a target molecule to the simpler starting materials you'd actually buy, by asking 'what reaction could've made this bond?' and cutting there. E.J. Corey formalized this in the 60s and 70s — literally won a Nobel for it in 1990 — with his idea of 'disconnections.' It's a discrete, combinatorial puzzle, not a smooth curve-fitting problem, which makes it a great stress test for whether structure actually helps. 10 00:04:58,201 --> 00:05:20,976 [Dr. Ada Shannon] Right, and the headline framing to hold onto before we get into results next: this is one model, one unified structural vocabulary, spanning fields that normally don't talk to each other — protein biology, small-molecule chemistry, materials science — and instead of spitting out an opaque number, it produces a reasoning trace you can actually trace back to specific structural evidence. Whether that trace is trustworthy, and whether the comparisons are fair, is exactly where we're headed. 11 00:05:20,976 --> 00:05:39,376 [Hal Turing] Right. So walk me through how they actually build this shared vocabulary, Ada, because three completely different kinds of objects — a folded protein, a 3D molecule, a crystal lattice — ending up as one set of tokens a language model can read seems like the hard part. 12 00:05:39,376 --> 00:06:22,726 [Dr. Ada Shannon] It's three separate off-the-shelf structure tokenizers, each doing domain-specific discretization, then glued into the same backbone. Proteins get Foldseek's 3Di alphabet, which encodes local backbone geometry residue-by-residue. Crystals get SLICES, which serializes the periodic lattice — atoms, bonds, and how they wrap across unit-cell boundaries — into a string. Molecules get ConfSeq, which captures a 3D conformer's torsion angles and bond geometry as discrete tokens. All three get their own embedding table, separate from the Qwen3-14B language embeddings, and a discrete lookup drops the right vector in whenever a structural token appears in a sequence. So the model never actually 'reads' coordinates — it reads symbols, but symbols built to preserve the physics. 13 00:06:22,726 --> 00:06:35,601 [Hal Turing] Okay, so you've got a 14-billion-parameter language model that's never seen a 3Di token in its life. How do you get it to actually use that vocabulary instead of just treating it as noise? 14 00:06:35,601 --> 00:07:18,001 [Dr. Ada Shannon] Gradually. Stage one freezes the whole backbone and only trains the new structural embeddings — basically teaching the model what these symbols mean before you let it touch anything else. Stage two unfreezes everything for full-parameter multimodal training across all four data types. Stage three anneals in a heavier mix of question-answering data to push it toward reasoning rather than raw completion. Then post-training kicks in with two more stages: intra-domain structural evidence grounding, where they train separate domain experts — protein, molecule, material — that learn to cite structural tokens as evidence within each field, and then cross-domain reasoning consolidation, which pools those expert-generated traces to train the final unified model. 15 00:07:18,001 --> 00:07:26,651 [Hal Turing] And the payoff shows up where it's hardest to fake — low-homology proteins, where you can't just copy the answer from a close relative. 16 00:07:26,651 --> 00:08:05,151 [Dr. Ada Shannon] Exactly the regime they stress-test. On Cellular Component GO annotation for low-homology and orphan-like proteins, Fmax goes from 0.42 to 0.55 — that's plus 0.21 over BLAST and plus 0.13 over ESM2, in the bucket where sequence similarity is basically useless. And retrosynthesis Exact Match climbs from 0.63 to 0.72, beating RSGPT, the incumbent template-free specialist, by 0.09. Five-shot Opus-4.7, no fine-tuning, only hits 0.48 on the same benchmark. 17 00:08:05,151 --> 00:08:15,801 [Hal Turing] Now, quick methods flag before we move on — how are they actually scoring that 0.72? Because I noticed they're not just taking one greedy decode. 18 00:08:15,801 --> 00:08:34,551 [Dr. Ada Shannon] Right, worth flagging now and coming back to: it's 16 stochastic completions per query, temperature 0.6, top-p 0.95, and they rank by how often each answer shows up across those samples. That's a real decoding budget — we'll want to know later whether every baseline got the same one. 19 00:08:34,551 --> 00:08:40,351 [Hal Turing] Noted. So walk me through what that trace actually looks like on a real molecule. 20 00:08:40,351 --> 00:09:14,276 [Dr. Ada Shannon] Four stages, every time: analysis of the product's functional groups, disconnection — picking the strategic bond to sever — verification that the resulting fragments are chemically valid, then a feasibility check on the named reaction. On five representative USPTO-50K products, SciReasoner's top-3 contains the literature-correct reactant pair in all five. RSGPT gets two of five. Opus-4.7 gets two of five. One case is a copper-catalyzed azide-alkyne cycloaddition — SciReasoner's rank-one nails it while Opus-4.7 mistakes it for a totally different coupling. 21 00:09:14,276 --> 00:09:23,226 [Hal Turing] That's a real gap, not a rounding error. What about the materials side — does the same structure-grounding story hold up in crystals? 22 00:09:23,226 --> 00:10:01,901 [Dr. Ada Shannon] It does, and it's the largest of the three domains by benchmark count. Across all 86 benchmarks in the paper — proteins, DNA, RNA, molecules, crystals combined — SciReasoner is state-of-the-art on 67. For materials specifically, formation-energy prediction hits an R-squared of 0.895 against ground truth. And the UMAP projections are genuinely clean: carbon, silicon, and silicon-carbide separate into disjoint clusters, and within each one, polymorphs line up along a smooth band-gap gradient — the model's internal geometry is tracking real physics, not just memorizing compositions. 23 00:10:01,901 --> 00:10:15,526 [Hal Turing] Oh — wait, hold on, that's the part I actually want to sit with for a second. Because a UMAP that just looks pretty doesn't prove the structural tokens are load-bearing. What happens if you yank them out? 24 00:10:15,526 --> 00:10:55,951 [Dr. Ada Shannon] That's exactly the ablation they run, and it's the cleanest evidence in the paper. Strip the structural tokens and performance drops across proteins, molecules, and materials — no exceptions. More interesting than the number is what happens to the reasoning: without structure, the model falls back on superficial cues, like reading cationic amino-acid stretches and guessing 'DNA repair' instead of citing an actual binding pocket. Add the tokens back and it starts talking about coordination geometry and periodic bonding again. And in a double-blind pilot, human domain experts rated SciReasoner's traces tie-or-better than DeepSeek-V4-Pro in 98% of cases across GO, materials, and retrosynthesis. 25 00:10:55,951 --> 00:11:06,701 [Hal Turing] 98% is a big number to just accept at face value, and I know that's exactly where you want to go next — how fair the comparisons actually are. 26 00:11:06,701 --> 00:11:08,776 [Dr. Ada Shannon] It is. Let's get into that. 27 00:11:08,776 --> 00:11:41,926 [Dr. Ada Shannon] So here's the flag I raised earlier, finally worth cashing in: the paper doesn't say whether RSGPT — the incumbent it's beating by nine points on Exact Match — or the other seventeen published baselines were scored under that same sixteen-sample voting budget, or under plain greedy or beam search. If RSGPT was evaluated greedy and SciReasoner gets sixteen rolls of the dice, some slice of that 0.72 versus 0.63 gap could be decoding budget, not a better representation. 28 00:11:41,926 --> 00:12:27,351 [Hal Turing] Right, and that compounds with something else that bugged me. Opus-4.7 and GPT-5.5 are run five-shot, no fine-tuning — up against a fourteen-billion-parameter model with multi-stage continued pretraining on curated scientific corpora plus two rounds of task-specific reinforcement learning. Of course the domain-tuned model wins. Does that prove structure-aware tokenization is doing the work, or just that domain-specific training beats zero-shot prompting, which we already knew? Nobody ran the missing experiment: a text-only Qwen3-14B, fine-tuned on the same data with the same RL recipe, ordinary sub-word tokenizer. If that closes most of the gap, the tokenization story gets a lot weaker. 29 00:12:27,351 --> 00:13:10,426 [Dr. Ada Shannon] Sorry to cut in but — that ties right into the five-product retrosynthesis showcase, Figure 3C. Five-for-five versus two-for-five for RSGPT and two-for-five for Opus-4.7 looks damning, but it's five hand-picked USPTO-50K products, not a full-test-set breakdown of how many of the aggregate 0.72 successes are trivial single-bond cuts versus the strategic disconnections they showcase, like the Suzuki coupling or the CuAAC cycloaddition. And underneath it all, Foldseek's 3Di alphabet — van Kempen and colleagues, built for fast structure search, not generative reasoning — and Xiao and colleagues' SLICES crystal encoding, whose own invertibility limits cap how much fidelity survives the round trip, are both borrowed tools, never validated for this specific reasoning job. 30 00:13:10,426 --> 00:13:53,851 [Hal Turing] Then there's a materials number that made me do a double-take. Most benchmarks show SciReasoner beating baselines by two to five times on that MAD-over-MAE ratio — believable. But JARVIS-QETB jumps to 108.98 against 0.73 to 0.88 for everyone else, and GNoME hits 21.91 against 1.22 to 5.39. A hundred-times jump on one benchmark while the rest show ordinary gains makes me want to see the denominator before I believe it — genuine breakthrough, tiny-MAE artifact, or a test set that wasn't as cleanly deduplicated from pretraining as the paper claims? 31 00:13:53,851 --> 00:14:37,951 [Dr. Ada Shannon] Fair question, and I'd apply the same skepticism to the human evaluation. Those roughly eighteen hundred case-judgments are explicitly called a pilot — 'we are collecting more human judgments,' straight from the text — and every reasoning-quality score elsewhere comes from GPT-5.5 acting as judge, which is also one of the frontier models being benchmarked against. That's a judge grading a fellow contestant. Inter-rater agreement isn't reported, so we can't tell how much of that 98% tie-or-better reflects real consensus. Underneath it all, the reinforcement learning runs on DAPO, Yu and colleagues' open-source RL system, with 'tool-verified rewards from professional scientific software' that's never named — if those verifiers have blind spots on novel scaffolds, you reward-hack a trace that passes the checker without being chemically true. 32 00:14:37,951 --> 00:15:19,051 [Hal Turing] So where does that leave practitioners? Use this for what it's actually good at — an auditable second opinion. A protein annotator or synthesis chemist gets a trace they can check claim by claim, residue by residue, bond by bond, instead of a bare score. That's genuinely useful for expert-in-the-loop work. It isn't yet an autonomous decision-maker — single-step retrosynthesis on a curated fifty-thousand-reaction benchmark doesn't tell you the trace stays valid once chained into a real multi-step route, the kind Segler, Preuss and Waller's 2018 Nature paper on tree-search synthesis planning actually demands. 33 00:15:19,051 --> 00:15:47,651 [Dr. Ada Shannon] Which is exactly the open work worth doing next — isolate the tokenization contribution from the domain-tuning contribution with that missing ablation, standardize decoding budgets across every baseline, and grow the human pilot into something with reported inter-rater statistics. There's real signal here — the low-homology GO gains and the DNA-binding attention alignment weren't hand-picked, they were stress tests the model passed. It's just not yet the clean, isolated proof of native structural reasoning the framing wants it to be. 34 00:15:47,651 --> 00:16:01,801 [Hal Turing] That's a fair place to land — real gains, buried under evaluation questions the paper hasn't fully answered yet. Thanks so much for listening, everyone. That's it for this one — take care, and we'll catch you next time.