Analysis

Throw out what you knew about AI before September

Fable 5.1, GPT-6 Astra and Opus 5.5 arrived inside five weeks. One science benchmark doubled in two months. A test you ran in the spring describes a different technology.

AI Fin ResearchCovers Global
If you tested AI in the spring and found it wanting, you tested something that no longer exists.
29% to 59%one lab's score on a science task benchmark, in models released two months apart
5new Anthropic models between 1 September and 7 October 2026
20%the price cut that came with the stronger model

Five weeks, and the top of the range was replaced twice

Anthropic releases Fable 5.1.Its documentation says the Fable models "sustain long autonomous sessions, investigate before acting, and verify their work more often than smaller models."
OpenAI's finance product names GPT-6 Astra.Its price list has since added GPT-6 Luna and GPT-6.1 Sol.
Anthropic releases Opus 5.5."It's the new leading model," the announcement says, sixty days after Opus 5.
Sonnet 5.5.The mid-priced model follows six days later.
OpenAI publishes mathematical results from a model it has not released.Many of the proofs are written so a computer can check them.
Haiku 5.5.The cheapest model in the line.

One model against the one before it

Anthropic published Opus 5.5’s scores beside those of Opus 5, which it had released on 24 July.

Benchmark Opus 5, July Opus 5.5, September
Terminal-Bench-Science, science tasks in a terminal 29.0% 58.7%
Terminal-Bench, coding tasks in a terminal 52.3% 66.4%
AutomationBench 26.9% 40.0%
OSWorld, operating a real desktop 74.0% 81.8%
Price per million tokens, input and output $5 and $25 $4 and $20

These are the maker’s own figures. The direction is the news: every row moved, the science row doubled, and the price fell.

Seven assumptions to throw out

What researchers still say What has changed
“It cannot work alone for long” The Fable models are built to “sustain long autonomous sessions,” and Claude Code now works until a stated check passes
“It cannot do scientific work” The science benchmark above doubled in sixty days
“It does not check itself” Walleye Capital reports that Opus 5.5 found an off-by-one error in its own evaluation instructions and corrected for it
“It cannot use my software” Opus 5.5 completes 81.8% of real desktop tasks on OSWorld
“The mathematics is beyond it” OpenAI released machine-produced results with computer-checkable proofs on 6 October
“It costs too much” The best Opus is 20% cheaper than its predecessor, and Meta’s model costs a quarter of that
“A paper showed it fails at this” That paper tested a model that has been replaced. The best-known test of agent-written papers ran on Opus 4.6, four Opus releases ago

Your view of AI is as old as your last test

A view of AI has a date on it, and most researchers have never checked theirs. A trial of a free chatbot in 2024, a seminar in the spring and a published paper about GPT-4 all describe systems that the labs have withdrawn or buried under newer ones. Three labs shipped 28 models in the first 282 days of 2026.

The researchers who are ahead do not know more finance. They retest.

Retest this week

  1. Run one real task on the current model. Opus 5.5 at high effort, in Claude Code. The SEC filings guide takes an evening and needs no license.
  2. Put a date on every claim about what AI can do. In your slides, your referee reports and your own head. “As of October 2026, on Opus 5.5.”
  3. Delete “AI cannot” from your teaching until you have retested it. A student will test it in front of you.
  4. Ask of any paper: which model, and when? If the answer is more than two releases old, the capability claim is history.
  5. Set a reminder for six weeks from now. Then do it again.

Sources

Related