Throw out what you knew about AI before September
Fable 5.1, GPT-6 Astra and Opus 5.5 arrived inside five weeks. One science benchmark doubled in two months. A test you ran in the spring describes a different technology.
Five weeks, and the top of the range was replaced twice
One model against the one before it
Anthropic published Opus 5.5’s scores beside those of Opus 5, which it had released on 24 July.
| Benchmark | Opus 5, July | Opus 5.5, September |
|---|---|---|
| Terminal-Bench-Science, science tasks in a terminal | 29.0% | 58.7% |
| Terminal-Bench, coding tasks in a terminal | 52.3% | 66.4% |
| AutomationBench | 26.9% | 40.0% |
| OSWorld, operating a real desktop | 74.0% | 81.8% |
| Price per million tokens, input and output | $5 and $25 | $4 and $20 |
These are the maker’s own figures. The direction is the news: every row moved, the science row doubled, and the price fell.
Seven assumptions to throw out
| What researchers still say | What has changed |
|---|---|
| “It cannot work alone for long” | The Fable models are built to “sustain long autonomous sessions,” and Claude Code now works until a stated check passes |
| “It cannot do scientific work” | The science benchmark above doubled in sixty days |
| “It does not check itself” | Walleye Capital reports that Opus 5.5 found an off-by-one error in its own evaluation instructions and corrected for it |
| “It cannot use my software” | Opus 5.5 completes 81.8% of real desktop tasks on OSWorld |
| “The mathematics is beyond it” | OpenAI released machine-produced results with computer-checkable proofs on 6 October |
| “It costs too much” | The best Opus is 20% cheaper than its predecessor, and Meta’s model costs a quarter of that |
| “A paper showed it fails at this” | That paper tested a model that has been replaced. The best-known test of agent-written papers ran on Opus 4.6, four Opus releases ago |
Your view of AI is as old as your last test
A view of AI has a date on it, and most researchers have never checked theirs. A trial of a free chatbot in 2024, a seminar in the spring and a published paper about GPT-4 all describe systems that the labs have withdrawn or buried under newer ones. Three labs shipped 28 models in the first 282 days of 2026.
The researchers who are ahead do not know more finance. They retest.
Retest this week
- Run one real task on the current model. Opus 5.5 at high effort, in Claude Code. The SEC filings guide takes an evening and needs no license.
- Put a date on every claim about what AI can do. In your slides, your referee reports and your own head. “As of October 2026, on Opus 5.5.”
- Delete “AI cannot” from your teaching until you have retested it. A student will test it in front of you.
- Ask of any paper: which model, and when? If the answer is more than two releases old, the capability claim is history.
- Set a reminder for six weeks from now. Then do it again.
Sources
- Anthropic, Introducing Claude Opus 5.5 anthropic.com
- Anthropic, Claude Platform release notes platform.claude.com
- Claude Code documentation, model configuration code.claude.com
- OpenAI, ChatGPT for Financial Services openai.com
- OpenAI, Sharing AI progress in mathematics openai.com
- OSWorld os-world.github.io