AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: How Jev Supports AI Decision Modeling: 24 Practical Uses on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get audio and creator gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Thorsten Meyer’s Sept. 29 article maps 24 proposed uses for Jev, a tool that returns typed answers to questions about text or JSON for software to act on. Meyer says three publishing uses are live, 12 use cases meet his four-part fit test, seven need measurement and two are poor fits. The results and cost figures are his reported measurements and have not been independently verified.

Thorsten Meyer published a 24-use-case assessment of Jev on Sept. 29, reporting that three applications are running in his publishing operation and that about 90,000 decisions have been made so far. The report describes where he says the tool fits, which uses need further measurement and which he considers poor fits.

Meyer describes Jev as a service that takes a text or JSON state plus typed questions and returns structured answers that software can use. The answer types include a probability for yes-or-no questions, options with probabilities for a choice, and a position with confidence for a score. He says a call containing the state and questions takes about 0.3 to 0.9 seconds and costs about $0.04 per million input tokens. These figures are reported by Meyer; the supplied material does not include independent testing or pricing documentation.

The three live publishing uses are a relevance check, an English-language check and a fallback topic classifier. Meyer says the language check scanned 78,889 articles for $2.01, identified 1,576 as non-English and fixed 1,553. For the relevance check, he reports about 10,000 story-and-site pairings judged over three days, with 22% clearly on-topic. He also reports 89% agreement with a frontier large language model for the fallback classifier, rising to 97%–99% when Jev’s confidence was at least 0.8. The source does not define the evaluation method or independently validate these results.

Across the full set, Meyer labels 12 uses strong fits, seven “measure first” and two poor fits, alongside the three live applications. Examples he gives include detecting missing disclosures and moderating comments as strong fits. He classifies duplicate-event detection as a poor fit, saying his canary found no duplicates to catch. A thin-source detector is among the cases needing measurement because its proposed replacement for a character-count rule has not yet been shown to improve results.

At a glance
reportWhen: Published Sept. 29, 2026
The developmentThorsten Meyer published a map of 24 proposed Jev uses, including three he says are already running in his publishing operation.

24 use cases for Jev at a glance

Publishing, commerce, software, business operations and the home, sorted by fit.

Every use case, coloured by how well it fits

Start in the green. Amber needs a measurement first. Red fails at least one of the four conditions.
livestrong fitmeasure firstpoor fit

Proven in production

1Relevance gate: story and site2Language check3Classifier fallback

Publishing and content

4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderation

Commerce and support

10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triage

Software and AI systems

15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triage

Business ops and home

21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent

15 of 24 are ready to build or already running

3
12
7
2
Live
Strong fit
Measure first
Poor fit
Live: in my fleet today. Strong fit: meets high volume, narrow question, cheap errors and a visibly failing heuristic. Measure first: the failing heuristic is unproven.
From “24 Ways to Use Jev” on thorstenmeyerai.com. Figures are my own production measurements, September 2026, rounded, unless marked illustrative.

Reported Uses and Confidence Thresholds

Meyer reports using Jev for language screening, relevance checks and topic classification in his publishing operation. He describes the tool as suited to narrow decisions and says uncertain cases can be routed to people or other systems. His account includes cost and confidence figures that have not been independently verified.

Meyer says structured answers should not be treated as automatically correct. He recommends replaying historical decisions, reviewing disagreements and acting only on results that meet a confidence threshold. The article documents one practitioner’s account and does not establish that the reported thresholds or accuracy will transfer to other content, systems or datasets.

Meyer’s Four Conditions for Fit

Meyer says a use should pass four tests before Jev is wired into a workflow: high volume, a narrow question that does not require multi-step reasoning, errors that are cheap or can be escalated, and a visibly failing existing heuristic. He advises keeping a keyword rule if it already works rather than adding a model without evidence of a problem.

For measurement, he proposes replaying 300 to 500 past decisions, comparing results overall and by confidence band, and examining 20 disagreements to judge which answer was right. He says to wire the system in only where the high-confidence band reaches 95%, then enable a separate feature flag for a 5%–10% canary before wider rollout. These are Meyer’s recommended procedures, not reported independent standards or outcomes.

The article applies this framework to publishing and content, with the supplied material beginning a later section on commerce and customer operations. It identifies use cases by the question asked, the type of answer and the rule software should apply. Meyer says code chooses the action, while Jev supplies the answer and confidence.

“Jev is the right tool wherever a system needs thousands of small judgements and can hand the unclear ones to something smarter.”

— Thorsten Meyer, in the Sept. 29 article

How Far the Results Generalize

The reported cost, volume and agreement figures come from Meyer’s operation. The supplied source does not provide underlying datasets, a full evaluation protocol, independent replication or details about how the frontier model used for comparison was selected. It is also unclear whether the reported accuracy applies to other publishers or decisions beyond the tested topics.

Meyer says seven use cases still need measurement because a failing heuristic has not been demonstrated. The supplied material names thin-source detection, product fit in roundups and headline quality among these examples. It does not provide the full list of 24 applications: the source text ends during its commerce and customer operations section.

Measure Before Wider Rollout

Meyer recommends testing proposed uses against past decisions, reviewing errors by confidence band and examining disagreements before enabling automation. For applications that meet his thresholds, he recommends a 5%–10% canary controlled by a feature flag, followed by a broader rollout if results support it. The article does not announce a product launch or give a schedule for deploying the seven measure-first cases.

Key Questions

What is Jev used for?

Meyer describes Jev as a tool that returns structured answers to typed questions about text or JSON. Software can use those answers for tasks such as classification, screening and routing.

Which Jev uses does Meyer say are live?

He lists a story-to-site relevance check, an English-language check and a fallback topic classifier in his publishing operation.

How does Meyer recommend handling uncertain answers?

His proposed pattern is to act on clear, high-confidence results and route uncertain cases to a person or a more capable system. The exact action depends on the application.

Are all 24 use cases ready to deploy?

No. Meyer says 12 meet his strong-fit criteria, seven need measurement and two are poor fits. He also identifies three as already live; the supplied source does not include the complete list of 24.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Engineering Is Automated. Research Is the Residual.

Recent advances show AI has automated most engineering tasks in AI research, raising questions about the future of scientific innovation.

Salesforce Expands Agentforce 360 with OpenAI

AIThis post was created with the assistance of artificial intelligence (AI).Executive SummarySalesforce…

The Twelve Real Complaints About AI Tools in 2026 — A Reddit, Twitter, and GitHub Synthesis

In 2026, users on Reddit, Twitter, and GitHub report widespread issues with AI tools, highlighting discrepancies between marketed and actual performance.

Portfolio. The synthesis.

A comprehensive analysis of six European institutional responses to sovereign LLM development, highlighting strategic insights before EU AI Act enforcement begins.