🔍 Read the full analysis: A Closer Look At My September 2026 AI Workflow on ThorstenMeyerAI.com
Get smart everyday buys delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Thorsten Meyer’s Sept. 29 account describes a workflow that uses Claude Opus 5.5 for most software building and newly released GPT-6.1 Sol for detailed review. His model comparisons use Artificial Analysis Intelligence Index v4.3.x scores and estimated cost per task; he says teams should test models on their own work before switching.
Meyer’s comparison lists Opus 5.5 at 58 on the index’s top setting and an estimated $5.98 per task. GPT-6.1 Sol scores 51 at xhigh for $0.39 per task, while GPT-6 Astra scores 53 at its top setting for $3.26. Claude Fable 5.1 scores 53 for $7.63, Sonnet 5.5 scores 56 for $7.60, and GPT-6 Luna scores 37 for $0.07. The source does not provide a full methodology for those per-task estimates in the supplied material.
For development, Meyer places Opus at high effort by default: the source lists that setting at 54 index points and $1.82 per task. He reserves xhigh for work such as architecture, migrations and trust boundaries, at 56 points and $3.46. He assigns Sol high or xhigh to focused examination of files and code changes, and to a second review pass.
The source reports that GPT-6.1 Sol launched Sept. 29 at $2 per million input tokens and $10 per million output tokens. Its index figures range from 48 at medium to 51 at xhigh, with reported first-token times of 5.3 seconds, 57 seconds and 69 seconds at medium, high and xhigh, respectively. Meyer says he uses Sonnet 5.5 and Luna for narrower tasks, while Astra or Fable may provide another opinion when models disagree.
Opus builds. Sol reviews. Jev decides.
One price tape, six models
Score against cost, at every effort setting
The effort dial moves the bill more than the model
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol: near-Astra scores at a fraction of the price
Three published settings
| Setting | Index | Cost per task | Output tokens | First token |
|---|---|---|---|---|
| medium | 48 | $0.21 | 15M | 5.3 s |
| high | 50 | $0.32 | 25M | 57 s |
| xhigh | 51 | $0.39 | 36M | 69 s |
Same score band, very different bill
My stack: who builds, who reviews
Cheaper tokens are not cheaper work
Read the numbers with four warnings
Part 2: Jev, the model that decides instead of writing
One call in, typed answers out
Three question types
Confidence is the superpower
Three uses running in my publishing operation
The fit test, then the shadow test
- Replay 300 to 500 past decisions
- Compare overall and per confidence band
- Read 20 disagreements, decide who was right
- High band at 95% or better?
- Own flag, off by default
- Canary on 5 to 10 units
- Roll out in the confident band only
24 use cases, sorted by how well they fit
Proven in production
- 1Relevance gate
- 2Language check
- 3Classifier fallback
Publishing and content
- 4Thin-source detector
- 5Same-event dedupe
- 6Product fits roundup
- 7Disclosure present
- 8Headline quality
- 9Comment moderation
Commerce and support
- 10Support-ticket routing
- 11Return-reason coding
- 12Review to feature complaints
- 13Catalogue taxonomy
- 14Order-fraud pre-triage
Software and AI systems
- 15LLM guardrail
- 16RAG passage filter
- 17Citation check
- 18Tool and intent routing
- 19Log-line triage
- 20PR risk triage
Business ops and home
- 21Inbox triage
- 22Expense categorisation
- 23Lead qualification
- 24Smart-home intent
Limits, cost and one hard rule
How Task Costs Shape Model Choice
Meyer’s account illustrates a practical change in how some developers evaluate AI models: they may weigh the cost of completing a task alongside benchmark performance. In his figures, Sol xhigh costs $0.39 per task, compared with $3.26 for Astra at its top setting and $7.63 for Fable 5.1. Those estimates could make a second model review more feasible to run routinely, though the source does not establish that other users will see the same costs or results.
He also argues that higher effort settings can raise bills substantially for small score gains. On Opus 5.5, moving from xhigh to max adds two index points and increases the listed cost per task from $3.46 to $5.98. Such comparisons matter to teams choosing defaults, but they do not show which model will perform best on a specific codebase or task. Meyer recommends shadow-testing against existing workflows before switching.
Benchmarks Behind the Workflow
The figures cited by Meyer come primarily from the Artificial Analysis Intelligence Index v4.3.x, which he describes as a map of general capability rather than a verdict on an individual workload. The source gives model release dates ranging from Fable 5.1 on Sept. 1 to GPT-6.1 Sol on Sept. 29, with Opus 5.5 listed as released Sept. 22 and Sonnet 5.5 on Sept. 28.
His account distinguishes token prices from estimated task costs. It lists per-million-token prices of $4/$20 for Opus, $10/$50 for Fable and Astra, $2/$10 for Sol, and $0.10/$0.50 for Luna, in input/output order. Task costs vary with effort and output length, so the token price alone does not capture the expense of a given task.
““The practical reading: Sol is not the model I ask to build. It is the model I can afford to run on everything.””
— Thorsten Meyer
Limits of the Published Comparisons
The figures are benchmark results and cost estimates cited by Meyer, not evidence that every user will obtain the same quality, latency or cost. The source notes that one index point is within the noise and says low and max effort results for GPT-6.1 Sol had not yet been published. It also reports long first-token times at Sol’s high and xhigh settings.
The supplied account cuts off during an illustrative example about model prices and human review. It therefore does not provide the full calculation or its assumptions. The material also does not say how the estimated task costs were derived beyond identifying the index and listing token prices and output figures.
Testing Models on Real Work
Meyer’s next recommended step for teams is to shadow-test models on their own tasks before changing defaults. That would let developers compare quality, latency, review effort and cost against the work they actually need done. The source does not announce a scheduled follow-up or a date for additional benchmark results; GPT-6.1 Sol’s low and max settings remained unlisted in the account.
Key Questions
What model does Meyer use for building?
He names Claude Opus 5.5 at high effort as his main model for development, with xhigh for harder work such as architecture and migrations.
What role does GPT-6.1 Sol play in his workflow?
Meyer uses Sol at high or xhigh for focused inspection and independent review. He says the listed cost is $0.32 to $0.39 per task at those settings.
Are these benchmark scores proof that one model is best for every team?
No. Meyer describes the index as a measure of general capability, not a verdict on a team’s workload, and recommends shadow-testing before switching.
What remains unknown about GPT-6.1 Sol’s benchmark results?
In Meyer’s Sept. 29 account, the index had not yet published Sol’s low or max effort results. He also says high and xhigh have long reported first-token times.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
