Artificial Analysis Intelligence Index v4.3 · Oct 2024 – Sep 2026

Every time OpenAI or Anthropic beat themselves

Two years of frontier releases, filtered to one rule: a model appears only if its best configuration matched or beat its own lab's previous best. Everything is scored on a single modern scale, so the staircase is comparable end to end.

Scale AA Intelligence Index v4.3 Evals 10, weighted across agents / coding / science / general Line step — a record holds until it is broken

Applies to the three time-series charts. Arena and OSWorld 2.0 are snapshots, not series, and get none.

Frontier index

Computer use

OSWorld-Verified: 369 real desktop tasks, scored on whether the end state is correct. Anthropic shipped the first computer-use model in the month this window opens and was alone on the benchmark for nearly a year — OpenAI's first published score lands in December 2025. Every point here is a measured run, not an estimate.

Both series stop in mid-2026, and neither lab's newest models are here. OSWorld-Verified has no run for GPT-5.6, GPT-6 Astra, Claude Opus 5 or Claude Fable 5.1 — they were benchmarked on OSWorld 2.0 instead. The gap is symmetric, not an omission on one side, and the end of each line says which models are missing from it.

The successor benchmark disagrees violently. OSWorld 2.0 replaces the task set with 108 long-horizon workflows, and it is the only scale the newest models have been run on. It is not a continuation of the series above — Claude Opus 4.8 scores 83.4 on Verified and 20.6 here — so it gets its own panel rather than a splice onto the same line.

Arena

A snapshot, not a series — Arena removes deprecated models, so there is no retrievable history, and it re-baselined scores in July 2026, which makes older published figures incomparable. This is the live text leaderboard as fetched on 9 September 2026. The finding is the shape of it: no OpenAI model appears in the top ten, and GPT-6 Astra has entered Agent Arena with scores still pending.

Valuations

Post-money valuation at each announced primary round, both companies from one source so the comparison is like-for-like. Neither has listed: Anthropic filed a confidential S-1 on 1 June 2026 and is targeting an October Nasdaq debut, OpenAI has signalled 2027. The line steps because a valuation holds until the next round reprices it.

Log makes the two growth rates comparable; linear shows the absolute gap.

Tables

The record table

Every point on the chart, plus the near-misses. Est. marks a score Artificial Analysis publishes as an estimate pending independent evaluation — which covers essentially everything released before mid-2026.

Best published variant per model, AA Intelligence Index v4.3
Model Released Index Confidence Status

Effort ladders

Reasoning-effort variants exist only for models benchmarked from mid-2026 onward — nothing earlier has a published ladder, and none is estimated here. Because every one of them shipped inside an eight-week window, the effort view spaces models in release order rather than on the time axis; on a two-year axis the ladders sit on top of each other.

Published index score by reasoning effort
Model None Low Medium High xHigh Max Spread

Computer use, both scales

The two OSWorld versions side by side. A model with figures in both columns shows how far apart the scales are — this is why they are never drawn on one line.

OSWorld-Verified vs OSWorld 2.0
Model Released Verified 2.0

Arena and valuations

Arena text leaderboard, 9 September 2026
Rank Model Lab Score 95% CI
Post-money valuation by round
Company Round Date Raised Post-money

How to read this

The index is rescaled every version. Claude Opus 4.7 scored 57 on the index current at its launch and 41 on v4.3. Plotting historical headline numbers would show a staircase that is partly an artifact of rescaling, so every score here is the v4.3 value, including backfilled scores for 2024 and 2025 models.

Old models sit low by construction. v4.3 is built from Terminal-Bench v4.0, GDPval-AA v2, AutomationBench-AA, Humanity's Last Exam and similar — evaluations that did not exist in 2024. GPT-4o scoring 8 does not mean it was useless in 2024; it means the current test set is far outside what it was built for.

Ties count. GPT-5.5 matched GPT-5.4 at 39 rather than exceeding it, and is plotted as a diamond on the flat section of the line.

Every chart on this page has the same failure mode. The intelligence index rescales between versions, OSWorld replaced its task set, and Arena re-baselined its scores in July 2026. In each case the fix is the same: pick one version, use it for every model, and put the incompatible successor in its own panel instead of joining the line.

The projection is a curve fit, not a forecast. It is an ordinary least-squares quadratic through each lab's plotted points, extended six months past today and drawn dashed over a shaded band. A quadratic has no theory behind it here — it will happily send a benchmark through 100% or a valuation through the roof, and it takes every point at equal weight including the estimated ones. Where a series has only three points the fit passes exactly through them and has nothing left to disagree with, so it is drawn thinner and labelled; OpenAI's valuation line is the case in point.

The fit runs in whichever space is on screen. On the valuation chart that means the linear and log toggles produce genuinely different projections — a quadratic in dollars against a quadratic in growth rate — and Anthropic's ends at roughly $3T on one and $10T on the other. Neither is more correct; the spread between them is a fair measure of how much weight the exercise carries.

Two figures carry an asterisk. Claude Sonnet 5 is published with a June 2026 release month but no day, and is plotted mid-month. The top Arena row was rendered by the source with its leading digit missing ("507"); it is treated as 1507, consistent with sitting above the 1505 below it.

Model scores from Artificial Analysis Intelligence Index v4.3, its model leaderboard and the individual model pages. Computer use from the OSWorld-Verified and OSWorld 2.0 leaderboards. Arena from the live text leaderboard. Valuations from a single funding history covering both companies, so neither is measured against a friendlier source than the other. All retrieved 9 September 2026.