🔍 Read the full analysis: My September 2026 Setup For Building, Research, And Decisions on ThorstenMeyerAI.com
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Thorsten Meyer’s September 29, 2026 setup uses Opus 5.5 for building and GPT-6.1 Sol for research and review, with other models reserved for specific tasks. His source material argues that model cost per task and effort settings matter alongside benchmark scores, but the figures come from a general capability index and need workload-specific testing.
Thorsten Meyer says he is using Claude Opus 5.5 for software building and the newly released GPT-6.1 Sol for research and review, in a September 29 account of how he assigns work across AI models. His setup reflects a shift in emphasis from benchmark rank alone to cost per task, though the cited scores measure general capability and do not establish which model is best for every workload.
Meyer’s comparison draws on the Artificial Analysis Intelligence Index v4.3.x. In his table, Opus 5.5 scores 58 at its top setting and costs $5.98 per task; GPT-6.1 Sol at xhigh scores 51 and costs $0.39. The figures are presented as benchmark results, not a guarantee of performance on a reader’s own work. Meyer advises shadow-testing before switching a workflow.
He assigns Opus 5.5 high effort to features, APIs, multi-file changes and refactors, and xhigh to harder work such as architecture and migrations. He uses Sol at high or xhigh to examine specific files and review changes. He lists Astra and Fable as occasional second opinions, while Sonnet 5.5 and Luna are options for scoped tasks and routine classification.
The source says GPT-6.1 Sol launched on September 29 at the same token prices as GPT-6 Sol: $2 per million input tokens and $10 per million output tokens. On the cited index, Sol’s medium setting scores 48 at $0.21 per task; high scores 50 at $0.32; and xhigh scores 51 at $0.39. Meyer says high and xhigh take 57 and 69 seconds, respectively, to produce a first token, which may make them less suitable for interactive use.
Opus builds. Sol reviews. Jev decides.
One price tape, six models
Score against cost, at every effort setting
The effort dial moves the bill more than the model
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol: near-Astra scores at a fraction of the price
Three published settings
| Setting | Index | Cost per task | Output tokens | First token |
|---|---|---|---|---|
| medium | 48 | $0.21 | 15M | 5.3 s |
| high | 50 | $0.32 | 25M | 57 s |
| xhigh | 51 | $0.39 | 36M | 69 s |
Same score band, very different bill
My stack: who builds, who reviews
Cheaper tokens are not cheaper work
Read the numbers with four warnings
Part 2: Jev, the model that decides instead of writing
One call in, typed answers out
Three question types
Confidence is the superpower
Three uses running in my publishing operation
The fit test, then the shadow test
- Replay 300 to 500 past decisions
- Compare overall and per confidence band
- Read 20 disagreements, decide who was right
- High band at 95% or better?
- Own flag, off by default
- Canary on 5 to 10 units
- Roll out in the confident band only
24 use cases, sorted by how well they fit
Proven in production
- 1Relevance gate
- 2Language check
- 3Classifier fallback
Publishing and content
- 4Thin-source detector
- 5Same-event dedupe
- 6Product fits roundup
- 7Disclosure present
- 8Headline quality
- 9Comment moderation
Commerce and support
- 10Support-ticket routing
- 11Return-reason coding
- 12Review to feature complaints
- 13Catalogue taxonomy
- 14Order-fraud pre-triage
Software and AI systems
- 15LLM guardrail
- 16RAG passage filter
- 17Citation check
- 18Tool and intent routing
- 19Log-line triage
- 20PR risk triage
Business ops and home
- 21Inbox triage
- 22Expense categorisation
- 23Lead qualification
- 24Smart-home intent
Limits, cost and one hard rule
Cost Shapes Meyer’s Model Choices
Meyer’s setup highlights how price per task can change the practical value of a model ranking. In his figures, Sol xhigh scores one to two index points below Astra and Fable 5.1, while its stated task cost is substantially lower. That gap may make routine review more affordable, but the comparison applies to the benchmark’s measured tasks and settings.
He also argues that effort level can affect cost as much as model choice. For Opus 5.5, his table puts high at 54 index points and $1.82 per task, while max reaches 58 points at $5.98. Sonnet 5.5 max is listed at $7.60 per task for 56 points. These numbers illustrate a trade-off in the cited measurements; they do not establish that lower effort will suit every project.
The proposed workflow uses a different model family to review Opus’s output. Meyer says that costs $0.32 to $0.39 per task with Sol at high or xhigh, making regular review more feasible for him. A separate model can provide another perspective, but it cannot correct missing requirements shared across the prompt and review.
From Rankings to Task Costs
Meyer frames his comparison around a claimed change in the AI model landscape over the four weeks before publication: six models had clustered within about 20 index points, while their costs per task differed by roughly 100 times. Those are his summary figures; the supplied material does not give the underlying calculation or a full set of task definitions.
The index is a general capability measure, according to Meyer, rather than a direct assessment of coding, research or review on any particular team’s work. His table lists Opus 5.5, Sonnet 5.5, Fable 5.1, GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna. He reports that Opus 5.5 was released September 22, Sonnet 5.5 on September 28, and Sol on September 29. His token-price figures also vary by model, and token rates alone do not capture the full cost of completing and checking a task.
““The question from ‘which model is smartest?’ to ‘which model clears my quality bar at the lowest cost per task?’””
— Thorsten Meyer
Benchmark Limits and Open Questions
The source does not provide independent confirmation of Meyer’s benchmark cost figures or explain all assumptions behind the per-task calculations. The index scores are not workload-specific, and the practical results may differ by prompt, project and review process. Meyer says one index point is within the noise, which limits the meaning of small score gaps.
Some measurements were not yet available in the account: Artificial Analysis had not published Sol’s low or max settings. The source also says Sonnet 5.5 max produced about 193,000 output tokens per task in the index, but gives no further detail here on how that result was measured. It is not clear how the models compare on Meyer’s own completed tasks over time, or whether the described costs include human review.
Testing the Setup on Real Work
Meyer’s stated next step for anyone considering the same assignments is to shadow-test models on their actual workload before making a switch. That would let users compare quality, task completion costs and response times against their current process. The source does not announce a formal follow-up test or a date for updated results.
For his own workflow, Meyer says he uses Opus to build and Sol to examine details and review meaningful changes. He advises sending a failed case and its evidence back for correction, while treating tests as one part of review rather than approval to ship. Whether the balance remains useful will depend on later benchmark updates and results from real projects.
Key Questions
What models does Meyer use for building and review?
He says Opus 5.5 is his main model for building, while GPT-6.1 Sol handles detailed research and review.
Why does Meyer use GPT-6.1 Sol for review?
His cited index figures put Sol high at $0.32 per task and xhigh at $0.39. Meyer says the relatively low task cost makes routine review affordable in his workflow.
Does the index show which model is best for every job?
No. Meyer describes the Artificial Analysis Intelligence Index as a measure of general capability and recommends testing models against the work they would actually perform.
What trade-off does Sol have at high effort?
The source lists Sol high at 50 index points and $0.32 per task, with a 57-second time to first token. Meyer notes that the wait can make it less suitable for interactive use.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
