🔍 Read the full analysis: 24 Ways To Organize AI Decisions With Jev on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
On Sep 29, 2026, ThorstenMeyerAI published a mapped set of 24 use cases for Jev, a tool that returns calibrated yes/no, choice and score answers for programmatic branching. Three are live in the author’s publishing operation covering roughly 90,000 decisions, 12 meet his four-condition fit test, 7 need measurement first, and 2 are rated poor fits.
ThorstenMeyerAI’s Thorsten Meyer has published a structured map of 24 concrete use cases for Jev, a decision engine that returns calibrated answers to typed questions rather than generated prose. According to the author, three uses are live in his own publishing operation and have processed roughly 90,000 decisions, twelve meet all four of his fit conditions, seven need measurement first, and two are rated poor fits. The piece, dated September 29, 2026, is presented by the author as a playbook for teams considering Jev for high-volume, low-stakes automated judgements.
Jev, as described in the report, does not write, summarise or extract. A caller sends a state — text or JSON — plus a set of typed questions, and receives answers a program can branch on without parsing prose. One call carries the state and all questions, takes 0.3 to 0.9 seconds, and costs about $0.04 per million input tokens, according to the author. Three answer types are supported: noul (a probability of yes from 0 to 1), choice (one selected option with per-option probabilities and a confidence), and score (a position on ordered levels plus confidence).
The author identifies confidence as the tool’s central feature. In his measurement on a 31-topic classification task, Jev agreed with a frontier LLM 97 to 99% of the time when its confidence was 0.8 or higher, but only 42% of the time below 0.5. The recurring pattern across the use cases, as described in the report, is “act on the clear cases, route the gray zone” — the calling code, not Jev, decides what to do with each answer, similar to a keep, change, or kill routing strategy.
The three live uses all sit in publishing. A relevance gate judged about 10,000 story-and-site pairings in three days, finding 22% clearly on-topic. A language check scanned 78,889 articles for $2.01 in one night, flagging 1,576 non-English pieces and fixing 1,553 in place. A classifier fallback runs when the primary LLM errors, agreeing with a frontier LLM 89% overall and 97 to 99% at confidence 0.8 or higher.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
The Four-Condition Fit Test
Central to the report is a four-condition fit test the author applies before wiring anything in: high volume (thousands of small calls, not a few big ones), a narrow question with no multi-step reasoning, cheap errors (or unsure cases routed to something smarter), and — the condition he says is most often skipped — a visibly failing heuristic, measured rather than assumed. If a keyword rule already works, he says, it should be kept.
The author reports two rejections among his own candidates, including a same-event dedupe use case that his canary test dropped because it found zero duplicates to fix. The map lists failures alongside successes, and the author cites the economics — a full-site scan for about two dollars — as his basis for arguing that checks previously too expensive to run everywhere may become viable at this price point. Compliance-adjacent uses, such as an affiliate-disclosure detector that routes misses to human review rather than auto-publishing, follow a pattern the author describes as “cheap judgement plus human gray zone.”
How the Decisions Were Validated
The author says each use case is presented with the question to ask, the question type, and the rule that acts on the answer, tagged as Live, Strong fit, Measure first, or Poor fit. The split: 3 live, 12 strong fits, 7 measure-first cases, and 2 poor fits across publishing, commerce, software, business operations and the home.
His prescribed validation path is staged: replay 300 to 500 real past decisions, compare results overall and per confidence band, read 20 disagreements to judge which system was right, and wire the tool in only where the high-confidence band reaches 95%. Rollout follows the same caution — a dedicated feature flag, off by default, a canary on 5 to 10 units, then full deployment. He attributes the live relevance gate’s safety record to its design of combining three answers into one decision and only acting when Jev is confident, leaving the uncertain middle on its existing path.
One measurement cited in the report: 88% of the news items he processes start from a bare headline, which motivated a “thin-source detector” now tagged measure-first, replacing a 300-character rule he says cannot distinguish a dense wire item from a teaser.
“Jev is the right tool wherever a system needs thousands of small judgements and can hand the unclear ones to something smarter.”
— Thorsten Meyer, ThorstenMeyerAI
What Is Still Unmeasured
Seven of the 24 use cases remain tagged measure-first because the fourth fit condition — a demonstrated, measured failure of the existing heuristic — is unproven. These include the thin-source detector, headline quality scoring, and product-fit checks for shopping roundups. The author states these need measured error rates on today’s matching rules before any wiring.
The agreement and confidence figures come from the author’s own measurements on his data — a 31-topic classification and replayed production decisions — and have not been independently verified or peer-reviewed. The reported $0.04 per million input tokens and 0.3-to-0.9-second latency are his observations rather than vendor benchmarks. The source material also cuts off partway through the commerce section, so the full detail of the commerce, software, operations and home use cases beyond publishing is not visible in the provided excerpt.
The Measure-First Candidates and Rollout Path
The stated next step for the seven measure-first candidates is to run the 300-to-500-decision shadow replays against existing rules and check whether the high-confidence band reaches 95%. The two poor fits, including same-event dedupe, are parked unless a duplicate problem is later measured. For the twelve strong fits, the author recommends starting there after reading the fit test, following the same canary-and-flag rollout used for the live publishing uses. According to the report, readers applying the map to their own systems would follow the same sequence: measure the failing heuristic first, shadow-replay, then canary before full deployment.
Key Questions
What is Jev, according to the report?
A decision engine that accepts a state (text or JSON) and typed questions, and returns calibrated answers — yes-probabilities, choices, or scores with confidence — that code can branch on. It does not generate prose to parse. One call takes about 0.3 to 0.9 seconds and costs roughly $0.04 per million input tokens, per the author’s figures.
How many of the 24 use cases are actually running?
Three, all in the author’s publishing operation: a story-site relevance gate, an English-language check, and a classifier fallback. Together they account for roughly 90,000 decisions so far. Twelve more are strong fits not yet wired in.
What is the four-condition fit test?
High volume of small calls, a narrow question without multi-step reasoning, cheap errors or a routing path for unsure cases, and a visibly failing heuristic that has been measured — not assumed. If an existing rule already works, the author says to keep it.
On a 31-topic classification, Jev agreed with a frontier LLM 97 to 99% of the time at confidence 0.8 or higher, and 42% below 0.5. These are the author’s own measurements on his data, not independent benchmarks.
Why were two use cases rejected?
They failed at least one fit condition. Same-event dedupe was dropped after a canary test found zero duplicates to fix, meaning no measurable problem existed for Jev to solve. The second poor fit is named in the full report but its details are outside the provided excerpt.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
