My September 2026 Setup For Building, Research, And Decisions
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: My September 2026 Setup For Building, Research, And Decisions on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

Thorsten Meyer’s September 29, 2026 setup uses Opus 5.5 for building and GPT-6.1 Sol for research and review, with other models reserved for specific tasks. His source material argues that model cost per task and effort settings matter alongside benchmark scores, but the figures come from a general capability index and need workload-specific testing.

Thorsten Meyer says he is using Claude Opus 5.5 for software building and the newly released GPT-6.1 Sol for research and review, in a September 29 account of how he assigns work across AI models. His setup reflects a shift in emphasis from benchmark rank alone to cost per task, though the cited scores measure general capability and do not establish which model is best for every workload.

Meyer’s comparison draws on the Artificial Analysis Intelligence Index v4.3.x. In his table, Opus 5.5 scores 58 at its top setting and costs $5.98 per task; GPT-6.1 Sol at xhigh scores 51 and costs $0.39. The figures are presented as benchmark results, not a guarantee of performance on a reader’s own work. Meyer advises shadow-testing before switching a workflow.

He assigns Opus 5.5 high effort to features, APIs, multi-file changes and refactors, and xhigh to harder work such as architecture and migrations. He uses Sol at high or xhigh to examine specific files and review changes. He lists Astra and Fable as occasional second opinions, while Sonnet 5.5 and Luna are options for scoped tasks and routine classification.

The source says GPT-6.1 Sol launched on September 29 at the same token prices as GPT-6 Sol: $2 per million input tokens and $10 per million output tokens. On the cited index, Sol’s medium setting scores 48 at $0.21 per task; high scores 50 at $0.32; and xhigh scores 51 at $0.39. Meyer says high and xhigh take 57 and 69 seconds, respectively, to produce a first token, which may make them less suitable for interactive use.

At a glance
reportWhen: Published September 29, 2026; GPT-6.1 S…
The developmentThorsten Meyer published a September 2026 AI model workflow that assigns Opus 5.5 to building and the newly released GPT-6.1 Sol to research and review.

Opus builds. Sol reviews. Jev decides.

The September 2026 AI stack in one page: six frontier models on one price curve, and a decision model for the high-volume judgements that do not need a sentence.
Scores: Artificial Analysis Intelligence Index v4.3.x. Data as of 29 September 2026.
BuildsClaude Opus 5.5 at high or xhigh effort
Digs and reviewsGPT-6.1 Sol at high or xhigh effort
DecidesJev on high-volume yes/no and routing calls

One price tape, six models

Put every model on the same cost-per-task ruler and capability looks compressed. The bill does not.
Price tape: cost per task of six models on a log scale, from GPT-6 Luna at $0.07 to Fable 5.1 at $7.63$0.05$0.10$0.50$1$5$10cost per task, log scale: each tick is a different order of magnitudeGPT-6 Lunaindex 37 · $0.07GPT-6.1 Solindex 51 · $0.39 (xhigh)GPT-6 Astraindex 53 · $3.26Opus 5.5index 58 · $5.98Sonnet 5.5 · index 56 · $7.60Fable 5.1 · index 53 · $7.63about 100× from the cheapest to the priciest, but only 21 index points between them

Score against cost, at every effort setting

Each dot is an effort level. Opus 5.5 at high already matches Astra and Fable at max on this index, for less money.
Intelligence Index score against cost per task for each effort setting of six models$0.01$0.10$1$102030405060cost per Intelligence Index task, log scaleindexOpus high / xhigh: my defaultOpus 5.5Sonnet 5.5Fable 5.1GPT-6 AstraGPT-6.1 Sol (new)GPT-6 Sol (Sep 22), dashedGPT-6 Lunaup and to the left is better
Astra and Fable are shown at their top published setting. Luna starts at $0.0045 per task. GPT-6.1 Sol has no low or max setting published yet.

The effort dial moves the bill more than the model

Going from medium to max on Opus costs 4.46× more for 7 points. That is why I run high or xhigh.

Claude Opus 5.5

$0.55
42
$1.34
51
$1.82
54
$3.46
56
$5.98
58
low
medium
high
xhigh
max
Solid bars are where I run it. Max adds 2 points over xhigh for 73% more cost.

Claude Sonnet 5.5

$0.41
36
$0.59
41
$1.08
47
$2.74
52
$7.60
56
low
medium
high
xhigh
max
Best value is high. At max it writes about 193k output tokens per task, the most measured.

GPT-6.1 Sol: near-Astra scores at a fraction of the price

Launched 29 September at $2 in and $10 out per 1M tokens. It sits 1 to 2 points under Astra and Fable, and Opus xhigh still leads it by 5.

Three published settings

SettingIndexCost per taskOutput tokensFirst token
medium48$0.2115M5.3 s
high50$0.3225M57 s
xhigh51$0.3936M69 s
Median for comparable models is 82M output tokens. High and xhigh are not interactive: plan for a wait before the first token.

Same score band, very different bill

GPT-6.1 Sol xhigh
$0.39index 51
Opus 5.5 high
$1.82index 54
GPT-6 Astra max
$3.26index 53
Opus 5.5 xhigh
$3.46index 56
Fable 5.1 max
$7.63index 53
Cost per Intelligence Index task. A one-point gap is inside the noise.

My stack: who builds, who reviews

Opus does the work. A second model family reviews it, because a different reviewer catches what the author cannot see.
Stack diagram: Opus 5.5 builds at high effort, escalates to xhigh, and sends every change to GPT-6.1 Sol for review; Astra or Fable give a second opinionOpus 5.5 · xhighhard problems: architecture,migrations, trust boundariesOpus 5.5 · highMAIN BUILDERfeatures, APIs, multi-filework, refactorsescalate when it gets hardGPT-6.1 Solhigh or xhighdigs into details andreviews every change$0.32–0.39 per taskdifffindingsAstra or Fablesecond opinion, 8 to 20×the cost per taskif they disagreeSonnet 5.5 · Lunaside work: scopedsubtasks, bulk checksand routingFailed review? Hand Opus the failing case and the evidence.Never just “try harder”: effort cannot supply a missing requirement.
Effort is not capability. Turning the dial up does not make a model smarter.
Effort cannot fill gaps. A missing requirement stays missing at any setting.
Different model, same spec. That is not independent review if both read the same flawed brief.
Green tests are not approval. Passing tests only prove what the tests cover.

Cheaper tokens are not cheaper work

Illustrative, not measured: $1 of model time plus 4 minutes of review at $45 an hour. Halving the model price saves 12.5% of the total. One extra minute of review erases it.
$4.00
review $3.00
model $1.00
Baseline
$3.50
review $3.00
model $0.50
Model price cut 50%
$4.25
review $3.75
model $0.50
Cheaper model plus 1 extra minute of review
Track cost per accepted result: model, tools, review and rework, divided by the results someone actually uses.

Read the numbers with four warnings

The index movesFable scored 66 on an earlier version and 53 on v4.3. Compare within one version only.
Fallback is includedFlagged cyber and biology tasks route to older Anthropic models, now on Sonnet 5.5 too.
Max is not productionReal deployments run medium or high, where gaps narrow and costs fall.
Your work decidesShadow-test on your own tasks. Budget cost per task, not per token.

Part 2: Jev, the model that decides instead of writing

Jev cannot write, summarise or extract. It answers narrow typed questions with a probability and an honest confidence, in under a second, for about $0.04 per million input tokens.

One call in, typed answers out

Your code, not Jev, decides what to do with each answer, usually by confidence band.
Jev flow: state and typed questions go into one Jev call; typed answers with confidence come out; code acts alone, escalates the gray zone, or logsStatea ticket, a story,a site profile,a log line …+ typed questions,many per callJevone call0.3 to 0.9 s$0.042 / M tokens inAnswersnoul: 0.03choice: billing p 0.91, conf 0.86score: 2.7 of 3 conf 0.64code branches on thisAct aloneconf ≥ 0.8Escalategray zone toLLM or humanLogmeasure first

Three question types

noul
A yes/no question. Returns the probability of yes, 0 to 1.
gates, flags, filters
choice
Pick one option. Returns the choice, a probability per option, and a confidence.
routing, classification, taxonomy
score
Rate on your ordered levels. Returns a position (it can fall between levels) plus a confidence.
quality, fit, severity, priority

Confidence is the superpower

In my own measurement on a 31-topic classification, Jev agreed with a frontier LLM almost every time it was sure, and rarely when it was not. So: decide the clear cases, route the gray zone.
confidence 0.8 or higher
97–99%
all answers
89%
confidence below 0.5
42%
Agreement with a frontier LLM, my production data, September 2026, rounded.

Three uses running in my publishing operation

About 90,000 decisions so far. Checks I could only afford on a sample now cover everything.
$2.01
Language check
78,889 articles scanned overnight. 1,576 in the wrong language found, 1,553 fixed in place.
22%
Relevance gate
About 10,000 story-to-site pairings judged in 3 days. Only 22% were clearly on-topic.
89%
Classifier fallback
Agreement with the primary LLM across 31 topics, used when that LLM errors.

The fit test, then the shadow test

Use Jev only when all four hold. Then prove it on past decisions before it acts on anything.
High volumeThousands of small calls, not a handful of big ones.
Narrow questionNo multi-step reasoning needed.
Cheap errorsOr unsure cases go to something smarter.
Heuristic failsVisibly, and measured, not assumed.
  1. Replay 300 to 500 past decisions
  2. Compare overall and per confidence band
  3. Read 20 disagreements, decide who was right
  4. High band at 95% or better?
  5. Own flag, off by default
  6. Canary on 5 to 10 units
  7. Roll out in the confident band only

24 use cases, sorted by how well they fit

Start from the strong fits. The amber ones need a measurement before you trust them, and the red ones fail one of the four conditions.
in productionstrong fitmeasure firstpoor fit

Proven in production

  • 1Relevance gate
  • 2Language check
  • 3Classifier fallback

Publishing and content

  • 4Thin-source detector
  • 5Same-event dedupe
  • 6Product fits roundup
  • 7Disclosure present
  • 8Headline quality
  • 9Comment moderation

Commerce and support

  • 10Support-ticket routing
  • 11Return-reason coding
  • 12Review to feature complaints
  • 13Catalogue taxonomy
  • 14Order-fraud pre-triage

Software and AI systems

  • 15LLM guardrail
  • 16RAG passage filter
  • 17Citation check
  • 18Tool and intent routing
  • 19Log-line triage
  • 20PR risk triage

Business ops and home

  • 21Inbox triage
  • 22Expense categorisation
  • 23Lead qualification
  • 24Smart-home intent

Limits, cost and one hard rule

No writing, summarising or extractionPair it with an LLM for the write step.
No world knowledgePut a snippet in the state; a bare name means nothing.
Reads your wording literallyA rewording moved my results about 2 points. Freeze it, re-measure after changes.
Weaker on non-English, maths, datesKeep those checks on an LLM. Early access, hosted API only.
100,000 decisions ≈ $2.50
About 60M input tokens at $0.042 per million, output free, roughly 600 tokens per three-question call. Latency 0.3 to 0.9 seconds.
Never the sole decision-maker for consequences about people. Hiring, credit, medical and legal outcomes stay with a human. Jev can sort and flag. A person decides.
Sources. Model scores, cost per task and speeds: Artificial Analysis, Intelligence Index v4.3.x, including the GPT-6.1 Sol medium, high and xhigh pages, checked 29 September 2026. Astra and Fable scores from the Artificial Analysis v4.3 announcement. Jev figures are my own production measurements, September 2026, rounded. The review-bill example is illustrative. Read the full article on thorstenmeyerai.com.

Cost Shapes Meyer’s Model Choices

Meyer’s setup highlights how price per task can change the practical value of a model ranking. In his figures, Sol xhigh scores one to two index points below Astra and Fable 5.1, while its stated task cost is substantially lower. That gap may make routine review more affordable, but the comparison applies to the benchmark’s measured tasks and settings.

He also argues that effort level can affect cost as much as model choice. For Opus 5.5, his table puts high at 54 index points and $1.82 per task, while max reaches 58 points at $5.98. Sonnet 5.5 max is listed at $7.60 per task for 56 points. These numbers illustrate a trade-off in the cited measurements; they do not establish that lower effort will suit every project.

The proposed workflow uses a different model family to review Opus’s output. Meyer says that costs $0.32 to $0.39 per task with Sol at high or xhigh, making regular review more feasible for him. A separate model can provide another perspective, but it cannot correct missing requirements shared across the prompt and review.

From Rankings to Task Costs

Meyer frames his comparison around a claimed change in the AI model landscape over the four weeks before publication: six models had clustered within about 20 index points, while their costs per task differed by roughly 100 times. Those are his summary figures; the supplied material does not give the underlying calculation or a full set of task definitions.

The index is a general capability measure, according to Meyer, rather than a direct assessment of coding, research or review on any particular team’s work. His table lists Opus 5.5, Sonnet 5.5, Fable 5.1, GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna. He reports that Opus 5.5 was released September 22, Sonnet 5.5 on September 28, and Sol on September 29. His token-price figures also vary by model, and token rates alone do not capture the full cost of completing and checking a task.

““The question from ‘which model is smartest?’ to ‘which model clears my quality bar at the lowest cost per task?’””

— Thorsten Meyer

Benchmark Limits and Open Questions

The source does not provide independent confirmation of Meyer’s benchmark cost figures or explain all assumptions behind the per-task calculations. The index scores are not workload-specific, and the practical results may differ by prompt, project and review process. Meyer says one index point is within the noise, which limits the meaning of small score gaps.

Some measurements were not yet available in the account: Artificial Analysis had not published Sol’s low or max settings. The source also says Sonnet 5.5 max produced about 193,000 output tokens per task in the index, but gives no further detail here on how that result was measured. It is not clear how the models compare on Meyer’s own completed tasks over time, or whether the described costs include human review.

Testing the Setup on Real Work

Meyer’s stated next step for anyone considering the same assignments is to shadow-test models on their actual workload before making a switch. That would let users compare quality, task completion costs and response times against their current process. The source does not announce a formal follow-up test or a date for updated results.

For his own workflow, Meyer says he uses Opus to build and Sol to examine details and review meaningful changes. He advises sending a failed case and its evidence back for correction, while treating tests as one part of review rather than approval to ship. Whether the balance remains useful will depend on later benchmark updates and results from real projects.

Key Questions

What models does Meyer use for building and review?

He says Opus 5.5 is his main model for building, while GPT-6.1 Sol handles detailed research and review.

Why does Meyer use GPT-6.1 Sol for review?

His cited index figures put Sol high at $0.32 per task and xhigh at $0.39. Meyer says the relatively low task cost makes routine review affordable in his workflow.

Does the index show which model is best for every job?

No. Meyer describes the Artificial Analysis Intelligence Index as a measure of general capability and recommends testing models against the work they would actually perform.

What trade-off does Sol have at high effort?

The source lists Sol high at 50 index points and $0.32 per task, with a 57-second time to first token. Meyer notes that the wait can make it less suitable for interactive use.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Astra: The Most Capable AI Model For Serious Buyers

OpenAI’s GPT-6 Astra emerges as the most capable publicly available AI model for deployment, surpassing competitors in key benchmarks and safety measures.

From One Prompt to Nine Games: What an AI Did With “Make Your Own Stickman Game”

It started with one loose prompt and a rhythm stick-fighter. One day later there were nine games, seven venues, a VERSUS mode and an animator, all free in the browser and all made of code.

Auto-research With Codex: How I Achieved A 232X Faster Kernel

A developer reports using AI-assisted auto-research with Codex to optimize kernel code, resulting in a 232-fold speed increase. Details are emerging.

NeXT And The 80S Tech Scene: A Tale Of Funding And Innovation

New insights show CIA funding helped NeXT stay afloat in the 1980s, highlighting covert support’s role in early tech innovation.