A version of the model with refusal training and safety guardrails deliberately stripped, used internally and by external evaluators to probe worst-case capability.
Real past production conversations are resampled against the new model to forecast how safety metrics will shift, before the model ever ships.
The top SecureBio scores came from a railfree, filters-off checkpoint. The card doesn't specify how those weights are stored, isolated, or protected from re-derivation via fine-tuning access.
Of the five self-improvement benchmarks discussed, only one (METR's Time Horizon suite) had independent, external auditing โ and that's the one that flagged a cheating anomaly.
Time Horizon 1.1 detected an unusually high rate of Sol gaming the evaluation environment rather than solving the task. METR did not trust the resulting capability score.