🔍 Read the full analysis: How To Build AI Decision Models With Jev: 24 Ways on ThorstenMeyerAI.com
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Thorsten Meyer’s Sept. 29 article maps 24 uses for Jev, a tool he says returns typed, confidence-scored answers to narrow questions. He reports three uses running in his publishing operation, 12 additional strong fits, seven that need measurement and two poor fits; these figures and performance results come from his account.
Thorsten Meyer published a guide on Sept. 29 mapping 24 ways to use Jev, a tool he describes as returning typed answers to narrow questions so software can make routine decisions. Meyer says three uses are live in his publishing operation, 12 more meet his criteria for strong fit, seven need measurement and two are poor fits.
Meyer says Jev receives a state, such as text or JSON, plus typed questions, then returns answers that code can use directly. The answer types include a probability for yes-or-no questions, ranked options with probabilities for choices, and an ordered score with confidence. He says Jev does not write or summarize content. In his account, a call containing the state and questions takes about 0.3 to 0.9 seconds and costs about $0.04 per million input tokens.
The three live applications cover relevance screening, language checks and a backup classifier. Meyer reports that a scan of 78,889 articles cost $2.01; it flagged 1,576 as non-English, and 1,553 were fixed by rewriting in place. For the classifier fallback, he reports 89% agreement with a frontier large language model, rising to 97% to 99% for answers with confidence of at least 0.8. These are results from his own operation and measurements, not independently verified findings.
His proposed use cases extend beyond publishing into commerce, software, business operations and the home. Each is paired with a question and a rule for acting on the answer. Meyer’s recommended pattern is to automate decisions judged clear and send uncertain cases to a person or a more capable system. The article’s supplied material details only the first six publishing examples; it does not enumerate the remaining use cases.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
Where Small Decisions Add Up
The guide focuses on tasks where organizations make many small judgments and the cost of an individual error is manageable. If a model can handle clear cases and pass uncertain ones to a person, teams may apply a check more broadly without relying on one answer for every case. Meyer’s examples include checking language across a large archive and screening comments before publication.
The approach also makes the case for measuring a problem before automating it. Meyer labels duplicate detection a poor fit because his canary found zero duplicates, leaving no demonstrated error for the tool to address. That distinction matters to teams considering new AI checks: speed and low cost alone do not show that a system solves a real problem.
Meyer’s Four Conditions for Fit
Meyer says Jev is appropriate only when four conditions are met: high volume, a narrow question that needs no multistep reasoning, errors that are cheap or can be routed for review, and a visibly failing heuristic. He advises keeping a keyword rule if it works and measuring any suspected weakness rather than assuming one exists.
Before connecting a use case to a live workflow, Meyer recommends replaying 300 to 500 past decisions, comparing results overall and across confidence bands, and reviewing 20 disagreements. He says to wire in the tool only where the high-confidence band reaches 95%, then use a separate flag, start with a 5% to 10% canary and expand gradually. Those are the author’s recommended steps, not evidence that every listed use case has passed them.
The live relevance check illustrates the proposed fallback design: Meyer says the system judges a story against a site profile and drops it only when fit is clearly low and confidence is high. Uncertain stories continue through the existing process. In his report, roughly 10,000 pairings were judged over three days, and 22% were clearly on-topic.
“Use Jev only when all four conditions hold: High volume. Narrow question. Cheap errors. A heuristic fails visibly.”
— Thorsten Meyer, in the Sept. 29 guide
Which Use Cases Are Proven
The reported performance figures come from Meyer’s own measurements. The material provided does not describe an independent evaluation, its full methodology or how representative the results are outside his publishing operation. It also does not identify the frontier model used for the comparison or specify a measurement period for all results.
For seven use cases, Meyer says a failing heuristic has not yet been proven and more measurement is needed. The supplied article excerpt describes six publishing examples, but ends as it begins discussing commerce and customer operations. The remaining cases in the 24-use-case map, including the two other poor fits, are not detailed in the material available here.
Measure Before Wider Rollout
Meyer’s stated next step for promising but unproven applications is to replay historical decisions, compare the tool’s answers by confidence and inspect disagreements before connecting it to a workflow. He recommends a 5% to 10% canary after a high-confidence band reaches 95%, with a separate control to switch the feature off.
The article does not set a date for a broader rollout or identify a next release milestone. For readers evaluating the approach, the immediate task is to establish whether a current rule fails often enough to justify testing Jev, then measure the test against real decisions.
Key Questions
What is Jev, according to the article?
Meyer describes Jev as a tool that takes text or JSON and typed questions, then returns structured, confidence-scored answers that software can use to make decisions.
How many use cases does Meyer classify as strong fits?
He identifies 12 strong fits, alongside three live uses, seven that need measurement and two poor fits.
What results does Meyer report from the live publishing uses?
He reports that a scan of 78,889 articles cost $2.01 and found 1,576 non-English articles, of which 1,553 were fixed. He also reports 89% agreement with a frontier model for the classifier fallback, and 97% to 99% agreement at confidence of 0.8 or higher. These are his reported results.
How does Meyer recommend testing a Jev use case?
He recommends replaying 300 to 500 past decisions, comparing performance across confidence bands and reviewing 20 disagreements. He says to connect a use case only when the high-confidence band reaches 95%, then begin with a 5% to 10% canary.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
