All news & analysis
Explainer 2026-09-06

Frontier AI: where the field actually stands

Training compute passed 1e26 FLOP and grows 4–5× a year, and in 2026 a frontier model was export-controlled like a strategic technology for the first time. The benchmarks keep moving: ARC-AGI-3 launched with every model under 1%, then one reached 30.2%.

Frontier AI carries four metrics and 17 milestones here, and it is the field where measuring progress is itself contested.

The cleanest number is training compute. GPT-4 was the first model at the 1e25 FLOP scale in 2023; frontier models passed 1e26 FLOP in 2025, a tenfold jump, with compute growing roughly 4–5× per year per Epoch AI. Over thirty models from twelve developers have now crossed 1e25. Annual private AI investment reached $109.1B in the US against China's $9.3B in 2024, and published research runs about 103,146 papers a year — by far the largest output of any field on this tracker.

Capability is harder to measure, and the tracker uses ARC-AGI precisely because it keeps resisting saturation. OpenAI's o3 scored 76–88% on ARC-AGI-1 in December 2024, the first AI to move past memorization on it, against a human baseline near 85%. ARC-AGI-2 followed in March 2025 as a harder successor. ARC-AGI-3 launched in March 2026 as the first fully interactive benchmark — hand-built game environments with no instructions, where agents must discover the rules — and every frontier model scored under 1% at launch, best 0.37%, while humans solve them all. In July, ARC Prize independently administered Claude Opus 5 and recorded 30.2%, close to four times the previous best of 7.8%, clearing several environments no model had beaten. Worth flagging because a different number circulated more widely: a 96.2% figure came from an independent developer's own harness over 25 public levels, not from ARC Prize's administered evaluation. The tracker records the administered one.

Two structural shifts stand out in the 2026 record. The first is compute sovereignty: DeepSeek's 1.6T-parameter V4 runs on Huawei Ascend silicon, and a Huawei-led team completed full-parameter post-training on about 1,000 Ascend 910Cs. Pre-training hardware remains undisclosed, so 'trained without Nvidia' is not established — a distinction the field's coverage frequently loses. DeepSeek moved to raise about $7.4B at up to $59B in its first external round.

The second is that frontier models became objects of state control. In June 2026 Anthropic released Claude Fable 5 under a tiered safety deployment, with the unrestricted Mythos 5 going only to vetted cyberdefenders via a US government program. Days later the US Commerce Department export-controlled both models, barring all foreign-national access; unable to enforce that selectively in real time, Anthropic shut both off worldwide. It is the first time a deployed frontier AI model was export-controlled like a strategic technology.

Efficiency moved too. CoreWeave trained DeepSeek-V3, 671B parameters, in 2.02 minutes on 8,192 GB300 GPUs in MLPerf v6.0. Anthropic's Claude Sonnet 5 pushed near-Opus-class agentic performance to a fraction of the cost. And 2025's lesson — that reasoning and agentic training, not raw scale alone, drive the frontier — held.

The locked milestone is AGI itself, with no date attached and the definition explicitly marked contested. That is the honest placement.

Related metric Training compute
Share

More on this topic