🔍 Read the full analysis: How To Apply Jev Across 24 AI Decision-Making Scenarios on ThorstenMeyerAI.com
Get smart everyday buys delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A September 29 article from Thorsten Meyer AI maps Jev, a tool for returning typed, confidence-scored answers to narrow questions, to 24 decision-making scenarios. The author says three uses are live, 12 are strong fits, seven need measurement and two are poor fits; the supplied material details only the first nine scenarios.
Thorsten Meyer AI published a guide on September 29 mapping Jev to 24 decision-making scenarios across publishing, commerce, software, business operations and the home. The author reports that three uses are live in a publishing operation, 12 meet the guide’s four-part fit test, seven need further measurement and two are poor fits.
Jev takes a state, such as text or JSON, alongside typed questions, and returns answers that software can use to branch. The guide describes three answer types: yes-or-no probabilities, choices with probabilities and confidence, and scores against ordered levels. It says Jev does not write, summarize or extract prose. A call carrying multiple questions takes about 0.3 to 0.9 seconds and costs about $0.04 per million input tokens, according to the article.
The author’s three reported live applications check whether a story fits a site, whether an article is in English, and which of 31 topics best matches a headline when a primary language model makes an error. In one overnight scan, the author says, Jev checked 78,889 articles for $2.01, found 1,576 non-English articles and led to 1,553 being fixed. The classifier reportedly agreed with a frontier language model 89% of the time overall, and 97% to 99% when Jev’s confidence was at least 0.8.
The guide’s first six publishing examples include a thin-source detector, duplicate-event detection, product matching in roundups, disclosure checks, headline-quality scoring and comment moderation. It labels disclosure checks and moderation strong fits, while thin-source detection, product matching and headline scoring need measurement. Duplicate detection is called a poor fit: the author’s canary test found no duplicates to address.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
Proposed Uses for Confidence Scores
The guide proposes using Jev for many narrow, low-cost decisions and using confidence scores to determine whether software handles an outcome automatically or routes it to a person or another system. The author says this approach could make it practical to apply checks more widely where manual review is costly.
The guide describes different handling for tasks with different consequences. It says missed disclosures can pose a compliance risk, recommends sending uncertain cases for human review, and advises against auto-publishing based on that check alone. For comment moderation, it proposes automatically approving clearly acceptable comments or hiding clearly identified spam, while queuing other cases. These are recommendations in the guide; it does not report verified results across all 24 scenarios.
The author says the duplicate-detection test found no duplicates to address. The guide recommends measuring the existing process before adding Jev, testing it against past decisions and reviewing disagreements.
The Guide’s Four Fit Conditions
The guide says a use case should meet four conditions before Jev is wired into a workflow: high volume, a narrow question without multi-step reasoning, errors that are inexpensive or can be escalated, and a heuristic that has been shown to fail. The last condition is meant to prevent teams from replacing a rule that already works.
For validation, the author recommends replaying 300 to 500 past decisions, comparing results overall and by confidence band, and reviewing 20 disagreements to determine which system was right. The suggested threshold for integration is at least 95% accuracy in the high-confidence band. The guide then proposes putting the feature behind a separate flag, starting with it off, and testing on 5% to 10% of units before broader rollout.
The material supplied for this article describes the operating method and the first nine scenarios, including three live uses and six publishing examples. Although the guide says its full map covers 24 cases, the remaining examples and the evidence supporting their classifications are not included in the supplied text.
“Jev is the right tool wherever a system needs thousands of small judgements and can hand the unclear ones to something smarter.”
— Thorsten Meyer, in the September 29 guide
Evidence Behind the 24 Cases
The supplied material does not identify Jev’s developer, explain how its confidence scores are calibrated, or provide independent verification of the author’s performance and cost figures. It also does not include the full list of 24 scenarios, so the seven cases marked for measurement and two called poor fits cannot all be assessed from the text provided.
The article reports agreement with a frontier language model, but does not name that model or describe the test set, sampling method or how disagreements were judged. Agreement alone does not establish which system was correct. The guide says the 300-to-500-decision replay and review of disagreements are needed to test performance for each workflow.
Measure Before Deployment
The guide recommends beginning with a shadow test on 300 to 500 real past decisions, then reviewing results by confidence and resolving a sample of disagreements. Teams should wire Jev into a workflow only if the high-confidence band reaches the proposed 95% threshold, according to the author.
For workflows that pass, the next proposed steps are to place the feature behind a dedicated flag, leave it off by default, test it on 5% to 10% of units and expand from there. The supplied article does not report deployment plans for the 21 scenarios beyond the three the author says are already live.
Key Questions
What does Jev do?
According to the guide, Jev answers typed questions about supplied text or JSON with probabilities, choices or scores that software can use. The author says it does not generate prose or summaries.
How many scenarios are already live?
The author says three scenarios are running in a publishing operation: site relevance checks, English-language checks and fallback topic classification.
When does the guide recommend using Jev?
It proposes using Jev when a task involves high volume, a narrow question, low-cost or escalatable errors, and a visibly failing heuristic. The guide recommends measuring the existing process before integration.
What does the supplied material say about all 24 scenarios?
It gives the overall fit counts—three live, 12 strong fits, seven needing measurement and two poor fits—but describes only the first nine scenarios. The remaining examples are not present in the supplied text.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
