🔍 Read the full analysis: A September 2026 Look At My AI Tools And Workflow on ThorstenMeyerAI.com
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
TL;DR
A 29 September 2026 workflow review by ThorstenMeyerAI.com argues six frontier models now sit within ~20 index points while task costs differ by ~100×, shifting selection from capability to cost efficiency. The author’s stack pairs Claude Opus 5.5 for building with the newly released GPT-6.1 Sol for cheap routine review.
The AI model market has consolidated into a price curve rather than a capability race, with six frontier models now clustered within roughly 20 index points of each other while their cost per task differs by about 100×, according to a workflow review published on 29 September 2026 by ThorstenMeyerAI.com. The author, Thorsten Meyer, says the practical question has shifted from “which model is smartest?” to “which model clears my quality bar at the lowest cost per task?” — and his answer pairs Claude Opus 5.5 as the main builder with the same-day-released GPT-6.1 Sol as a low-cost reviewer.
All capability figures in the review come from the Artificial Analysis Intelligence Index v4.3.x. At the top of the table, Claude Opus 5.5 (released 22 September) scores 58 at its max setting and costs $5.98 per task, or about 17 tasks per $100. At the bottom, GPT-6 Luna scores 37 but costs $0.07 per task — roughly 1,429 tasks per $100 — which is why Meyer assigns it classification, extraction and routing work.
Three findings shape the ranking, per the review. Opus 5.5 outscores its more expensive sibling Claude Fable 5.1 by 5 points while costing less per task. Sonnet 5.5 at max effort costs more per task than Opus at max while scoring 2 points lower, which Meyer says makes it a poor fit at that setting. And GPT-6.1 Sol costs about one-eighth of GPT-6 Astra and one-twentieth of Fable per task for a score only 1 to 2 points lower.
GPT-6.1 Sol, released 29 September at $2/$10 per 1M input/output tokens, scores 51 at xhigh for $0.39 per task. Its catches are real: high and xhigh settings take 57 to 69 seconds to produce a first token, making it unsuitable for interactive use, and Opus 5.5 still leads it by 5 points at xhigh. Artificial Analysis has not yet published low or max settings for the model, and Meyer cautions that one index point falls inside measurement noise.
Opus builds. Sol reviews. Jev decides.
One price tape, six models
Score against cost, at every effort setting
The effort dial moves the bill more than the model
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol: near-Astra scores at a fraction of the price
Three published settings
| Setting | Index | Cost per task | Output tokens | First token |
|---|---|---|---|---|
| medium | 48 | $0.21 | 15M | 5.3 s |
| high | 50 | $0.32 | 25M | 57 s |
| xhigh | 51 | $0.39 | 36M | 69 s |
Same score band, very different bill
My stack: who builds, who reviews
Cheaper tokens are not cheaper work
Read the numbers with four warnings
Part 2: Jev, the model that decides instead of writing
One call in, typed answers out
Three question types
Confidence is the superpower
Three uses running in my publishing operation
The fit test, then the shadow test
- Replay 300 to 500 past decisions
- Compare overall and per confidence band
- Read 20 disagreements, decide who was right
- High band at 95% or better?
- Own flag, off by default
- Canary on 5 to 10 units
- Roll out in the confident band only
24 use cases, sorted by how well they fit
Proven in production
- 1Relevance gate
- 2Language check
- 3Classifier fallback
Publishing and content
- 4Thin-source detector
- 5Same-event dedupe
- 6Product fits roundup
- 7Disclosure present
- 8Headline quality
- 9Comment moderation
Commerce and support
- 10Support-ticket routing
- 11Return-reason coding
- 12Review to feature complaints
- 13Catalogue taxonomy
- 14Order-fraud pre-triage
Software and AI systems
- 15LLM guardrail
- 16RAG passage filter
- 17Citation check
- 18Tool and intent routing
- 19Log-line triage
- 20PR risk triage
Business ops and home
- 21Inbox triage
- 22Expense categorisation
- 23Lead qualification
- 24Smart-home intent
Limits, cost and one hard rule
Effort Settings Now Outweigh Model Choice
The review’s central argument is that choosing the effort level now moves costs more than choosing between models. On Opus 5.5, going from xhigh to max adds 2 index points for 73% more cost per task; from medium to max, cost rises 4.46× for 7 points. That is why Meyer runs Opus at high or xhigh for development — high delivers 54 points for $1.82 — and reserves max for rare cases.
The second takeaway concerns independent review at commodity prices. A review pass at $0.32 to $0.39 per task, Meyer writes, is “cheap enough to be routine,” allowing a different model family to check every meaningful change rather than having Opus review its own output. He pairs this with four operating rules, including that effort is not capability, that a second model reading the same flawed spec is not an independent review, and that passing tests are not approval to ship.
The review also pushes back on pure price-cutting: halving model price saves about 12.5% of real cost in Meyer’s illustrative example, and a single extra minute of human review can erase the saving — a figure he labels illustrative, not measured.
A Month of Frontier Releases Reshaped the Stack
September 2026 saw near-weekly releases: Claude Fable 5.1 on 1 September, GPT-6 Astra on 3 September, Opus 5.5 and GPT-6 Luna on 22 September, Sonnet 5.5 on 28 September, and GPT-6.1 Sol on 29 September — Sol arriving just a week after its GPT-6 Sol predecessor at the same $2/$10 token pricing. Even Sol’s medium setting matches GPT-6 Sol’s score of 48 at one-fifth of its $1.06 per-task cost, according to the review.
Meyer’s previous approach is not detailed in the piece, but the framing — the frontier “stopped being a leaderboard and became a price curve” — signals the shift from picking a single best model to assembling a portfolio by role: Opus for building, Sol for detail work and review, Astra or Fable only as tie-breaking second opinions, and Luna plus Sonnet for side work. He also uses Jev, a decision model that cannot write a sentence, for high-volume yes/no and routing judgements.
Sol’s behavior on the index is unusually concise: its high setting used 25M output tokens, against a median of 82M for comparable models.
“The question changed from ‘which model is smartest?’ to ‘which model clears my quality bar at the lowest cost per task?'”
— Thorsten Meyer, ThorstenMeyerAI.com
Watching Sol’s Full Curve and the Field’s Response
Meyer indicates his stack is set for now — Opus 5.5 building, Sol reviewing — but several developments could force revisions. Artificial Analysis is expected to publish low and max effort settings for GPT-6.1 Sol, which will complete its price-performance picture. The October release cycle may bring competitors responding to Sol’s pricing, and Meyer’s stated practice of shadow-testing suggests any model switch would follow his own benchmark-before-switch rule rather than index scores alone. Readers tracking this stack should also watch whether Sol’s 57-to-69-second time-to-first-token at high settings improves, since that latency is the main barrier to interactive use.
Key Questions
What is the main claim of the September 2026 workflow review?
That six frontier models now sit within roughly 20 index points of each other while their cost per task differs by about 100×, so model selection should be driven by cost per task at a given quality bar, not by raw benchmark rank.
Which models does the author actually use day to day?
Claude Opus 5.5 at high or xhigh effort as the main builder; GPT-6.1 Sol at high or xhigh for detail work and review passes; Sonnet 5.5 at high and GPT-6 Luna for scoped subtasks and bulk work; Astra or Fable only as second opinions; and Jev, a decision-only model, for routing judgements.
Why is GPT-6.1 Sol cheap but not a replacement for Opus 5.5?
Per the review, Sol scores 51 at xhigh against Opus’s 56, and its high and xhigh settings take 57 to 69 seconds to produce a first token, making it unsuitable for interactive building. Its value is running review and detail passes at $0.32 to $0.39 per task.
What benchmark do the scores come from, and how reliable is it?
All scores come from the Artificial Analysis Intelligence Index v4.3.x. The author cautions it measures general capability, that one index point falls inside noise, and that teams should shadow-test models on their own workloads before switching.
Does turning up a model’s effort setting make it smarter?
No, according to the author. Higher effort raises cost sharply — Opus 5.5 goes from $1.34 per task at medium to $5.98 at max — but Meyer states effort is not capability and cannot fill in missing requirements.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
