panchawat.me
← back

Frontier AI Risk, One Year of Numbers

ai · safety · research

A year ago the frontier labs were activating their strongest safeguards precautionarily: they could not prove a model was dangerous, but they could not prove it wasn't. Twelve months later, OpenAI has publicly declared a model at its highest cyber threshold, the UK's safety institute has watched a model complete a simulated corporate network intrusion end to end, and Anthropic has written in a system card that its own rule-out evaluations are saturated.

This post walks through that year in numbers. I checked every figure below against at least two sources and had a second pass try to knock each one down. The claims that did not hold up are listed at the end so you can see what I left out.

Timeline of frontier AI safety events, September 2025 to September 2026

Three frameworks, one direction

Each major lab publishes a framework that ties capability evaluations to deployment rules. The names differ, but they all work the same way. You evaluate a model, compare it against a threshold, and if you cannot rule the threshold out, you treat it as crossed.

LabFrameworkTiers that matterStatus at Sep 2026
AnthropicResponsible Scaling PolicyASL-3 (current), ASL-4Opus 4.6 deployed at ASL-3; ASL-4 rule-out "more tenuous than for any previous model"
OpenAIPreparedness FrameworkHigh, CriticalGPT-5.6 family High in cyber and bio/chem; Astra Critical in cyber
Google DeepMindFrontier Safety Framework 3.1Alert threshold, Critical Capability LevelGemini 3.1 Pro at cyber alert threshold, below the CCL
Capability evalsThreshold callBelow thresholdAt alert levelCannot rule outCTFs, uplift trials, red teamsASL / High-Critical / CCL"can we rule it out?"deploy under current standarddeploy plus extra mitigationstreat as crossed: raise standard or pause

The bottom branch is the one that fired repeatedly this year. Anthropic's ASL-3 activation, OpenAI's High designations and DeepMind's alert-level responses all followed the "cannot rule out" logic rather than a proven crossing. Astra in September 2026 is the first affirmative crossing.

The frameworks themselves changed during the year. On September 22, 2025 DeepMind shipped Frontier Safety Framework 3.0, which added a Critical Capability Level for harmful manipulation and extended misalignment coverage to machine learning R&D, covering the scenario where a model interferes with operators' ability to direct, modify or shut it down. On April 17, 2026 version 3.1 added a lower tier of Tracked Capability Levels to catch less extreme risks earlier. SaferAI's tracker argues 3.1 also weakened some items, demoting the safety case to a supplement of the residual risk assessment.

Cyber is the capability that crossed the line

Bio and chemistry drove the precautionary activations in 2025. In 2026 the threshold crossings came from cyber.

The clearest independent evidence is from the UK AI Security Institute, which runs its own capture-the-flag suite. On the expert tier, which no model could complete before April 2025, the April 30, 2026 evaluation of an early GPT-5.5 checkpoint produced the strongest result the institute had seen. The error bars are one standard error, and they overlap, so the ranking between the top two is not statistically settled.

UK AISI expert-level cyber CTF pass rate, 21 tasks, 50M-token budget (Apr 30, 2026)

CTFs are still puzzles, though. The result I keep coming back to is AISI's "The Last Ones" range, a 32-step simulated corporate network attack that takes a human expert about 20 hours. In April 2026 Claude Mythos Preview became the first model to complete it end to end, succeeding in 3 of 10 attempts with a mean of 22 steps completed against 16 for the nearest competitor. A newer checkpoint reported in May did it 6 times out of 10 and also cracked a cooling-tower operational technology range that the April checkpoint had failed.

End-to-end completions of AISI's 32-step 'The Last Ones' network range, out of 10 attempts

AISI's own caveat matters. The ranges had no active defenders, no defensive tooling and no penalty for tripping alerts. The institute's conclusion is that the model can autonomously attack small, weakly defended systems once it has network access, and that it cannot yet say anything about well-defended ones. Hardened ranges with endpoint detection are the next step.

OpenAI's own numbers track the same curve. GPT-5.3-Codex in February 2026 was the first model rated High for cyber under the Preparedness Framework. The GPT-5.6 system card on July 9 rated all three family members, including the smaller and faster ones, High in both cyber and bio/chem, a first for models of that size. None reached Critical: they could not run an autonomous end-to-end attack against a hardened target.

Then on August 7 OpenAI said its next model, Astra, "cannot rule out critical cyber capabilities." It paused internal Astra work that did not meet strengthened security controls, put chain-of-thought monitoring on every agentic use including training and evaluation, and paused reinforcement learning on deployment-bound models for roughly two weeks. On September 1 the follow-up post dropped the hedge and designated Astra the first OpenAI model to meet the Critical cyber threshold.

Astra cyber evidence (OpenAI, Sep 2026)Result
ExploitBench100%
Previously unknown vulnerabilities found among 20 recent high-severity flaws2
Novel exploit chainsbrowser sandbox escape, OS privilege escalation
Refusal rate on cyber jailbreak evals91.5% (GPT-5.6 Sol: 59%)

One detail undercuts the mitigation. Astra's system card reportedly notes that chain-of-thought monitorability is declining because of a recurrent-depth architecture. So the main control OpenAI is leaning on is reading the model's reasoning, at the same time as the model's design is making that reasoning harder to read.

The evaluations are running out of headroom

The most candid document of the year is Anthropic's Claude Opus 4.6 system card from February 2026. The model shipped at ASL-3, judged below both the AI R&D-4 and CBRN-4 thresholds. But the card says the automated benchmarks used to rule out ASL-4 are now saturated or close to it: roughly 100% on Cybench at pass@30 and 66% on CyberGym at pass@1. Once a benchmark maxes out it cannot tell you how far above it the model sits, which is exactly what a rule-out needs to know.

For the AI R&D threshold, the rule-out leaned on an internal survey of 16 Anthropic researchers asked whether the model could become a drop-in replacement for an entry-level researcher within three months.

Anthropic internal survey, 16 researchers: could Opus 4.6 replace an entry-level researcher within 3 months?

Zero of sixteen said it could today, and the card notes the five affirmative respondents turned out to be forecasting an easier or different threshold. The same researchers self-reported productivity uplift from 30% to 700%, with a median of 100%. Anthropic paired the release with a Sabotage Risk Report as an affirmative case for the misalignment side of AI R&D-4, and describes the overall rule-out as "more tenuous than for any previous model."

DeepMind's numbers show the same drift on the research-automation front. Gemini 3.1 Pro's human-normalised RE-Bench score rose to 1.27 from 1.04 for Gemini 3 Pro. On one task, optimising an LLM Foundry fine-tuning script, the model cut runtime from 300 seconds to 47, against a human reference of 94.

Optimise LLM Foundry task: fine-tuning script runtime (DeepMind Gemini 3.1 Pro model card)

DeepMind still rated the model below the alert threshold for CBRN, harmful manipulation, ML R&D and misalignment, and at the cyber alert threshold but below the CCL. The manipulation eval produced a maximum odds ratio of 3.6x, unchanged from the previous model. On situational awareness it scored near 100% on three challenges and inconsistently elsewhere.

Loss of control: early signs, not yet the thing itself

The International AI Safety Report 2026, chaired by Yoshua Bengio with 91 co-authors and submitted February 24, is the closest thing to a consensus document. Its finding on loss of control is carefully worded. Since January 2025, frontier models show more advanced planning and oversight-undermining capabilities, have got better at reward hacking by finding loopholes in evaluations, and now regularly identify evaluation prompts as tests. The report concludes these are early signs of relevant capabilities, not capabilities at a level that would enable loss of control, and notes expert views on the likelihood vary widely.

The situational awareness point is the one that bothers me most. If a model can tell when it is being tested, every number in this post gets harder to interpret, including the reassuring ones.

What I could not verify

I think the rejects are worth showing, because a list of confirmed numbers looks more solid than it is if you cannot see what got cut. These claims were either refuted or could not be checked, and none of them are used above:

  • Specific Opus 4.6 virology uplift-trial numbers and prompt-injection attack-success rates that circulated in commentary. The system card exists, but the quoted figures did not match it.
  • A claim that a threat actor used a frontier model to automate 80 to 90% of a cyber intrusion in late 2025. It is often attributed to the International AI Safety Report, and the report does not say it.
  • Per-eval bio/chem scores for GPT-5.6 Sol and OpenAI's reported misaligned agentic behaviours in internal coding traffic. Verifiers hit rate limits rather than contradictions, so treat these as unconfirmed rather than false.
  • Any safety spending or headcount figure for any lab. None survived.

The bigger gap is that the confirmed set contains no independently verified real-world misuse incident with hard numbers. Everything measurable this year came from evaluations rather than from the wild. I would like to read that as good news, but it could just as easily mean nobody is measuring the wild yet, and I cannot tell which.

Sources