
AI in Action: Testing Jev, an AI Built for Judgment Calls, on 10,000 Sales Leads
Experiments
Experiments

by
Eric Harvey


AI in Action: Testing Jev, an AI Built for Judgment Calls, on 10,000 Sales Leads
With the "this changes everything!" generative AI arms race continuing interminably, I found myself intrigued (& slightly puzzled) by TypeSafe's recent release of Jev, a model that wasn't actually built to generate anything. New AI models arrive every few weeks, and most of them are variations on one idea: a faster or cheaper way to generate text. Jev is instead built to make judgment calls. You hand it some context and a set of questions, and it hands back typed answers with probabilities attached. What's more, it does so at a fraction of the time and cost (input is cheap, output is free) of standard models.
As I scratched my head, looking to figure out where this might be useful in real-world use cases, my mind went straight to B2B sales and marketing. So much of the daily work in a revenue team is exactly that kind of judgment. Is this lead worth a rep's time, or is it someone browsing on a Friday afternoon? Lead scoring and pipeline routing both run on calls like that, and today they're usually made by brittle rules or by people with too little time. A model designed to classify rather than to chat sounded like it might fit that work better than the tools most teams are reaching for.
It also sits close to why The Authentic .AI exists. The lab's tagline is "Digitally Powered. Authentically Human." A lot of the AI showing up in sales and marketing right now adds noise: more emails, and more content that sounds like everyone else's. I'm more interested in AI that clears space for the human part of the job. A model that quietly sorts the inbound pile, and admits when it isn't sure, might do exactly that.
So this experiment takes one real-world use case, inbound lead triage, and tries to learn something honest about where Jev is useful and where it isn't. It's not a product review and it's not a formal benchmark. It's one rough test, written up in full.
The Hypothesis
Generative AI models are built to write. Sales systems are built to route. When go-to-market (GTM) teams ask a chat model to score a big inbound list, they are borrowing a tool designed for one job and pointing it at another. It can work. But every answer comes back as text that has to be parsed, every call takes seconds, and the model's sense of how sure it is tends to be an afterthought.
TypeSafe calls Jev a "System One" model, built for fast judgment calls rather than long reasoning. My hypothesis was that a layer like this could route a high-volume inbound queue reliably, and that its confidence scores would tell me which leads were safe to automate. The second question was whether it would actually beat the cheaper or more familiar options once the leads got harder.
The Tech Stack
Data generation: Python (Faker library)
Orchestration: n8n 2.8.4, self-hosted on my laptop
Classification: TypeSafe Jev 1.13.0
Baselines: a regex rule set (simple keyword matching), Claude Haiku 5.5, Claude Opus 5.5, GPT-6 Luna and GPT-6 Astra
Step 1: Synthesizing the Pipeline Mess
To test at volume, I needed volume. I wrote a Python script to generate 10,000 synthetic inbound leads, each with a name, job title, company size and a free-text reason for reaching out.
The generator used three recipes. 3,141 leads (31.4%) were obvious junk: students, interns, messages like "asdf" or "need for class." 4,908 (49.1%) were ambiguous mid-market interest. The remaining 1,951 (19.5%) were enterprise buyers writing things like "we need to move off Salesforce" or "need a demo for my team this week."
One thing to note: this data was easy. Each recipe had an obvious tell, and that ended up shaping what the experiment could and couldn't prove.
Step 2: Defining the Schema
I loaded the CSV file into n8n, looped over it in batches of 100, and sent each lead to Jev's API as one line of context: name, title, company size and message. Each request asked three typed questions at once.
ICP fit, how closely a lead matches the ideal customer profile, was a Score question on four levels, from 0 (students, freelancers, companies under 50 people) to 3 (an enterprise VP of Sales, CRO or Head of Sales Systems). Jev returns a probability-weighted position on that scale, which I rounded before routing.
Intent was a Choice question with four options: High Intent, Exploratory, Support and Spam.
Sales action required was a yes-or-no likelihood question (TypeSafe calls it a Noul), which returns a probability from 0 to 1 that a statement is true. I ended up not using it. Genuine buyers scored anywhere from 0.49 to 0.83 on it, and some ambiguous leads reached 0.51, so there was no clean line to route on.
Step 3: Building the Confidence Gate
This is where the experiment got operationally interesting. A chat model gives you an answer. Jev gives you an answer, a probability for every possible answer, and a confidence number that summarizes how concentrated those probabilities are.
Right after the API call, I put a gate in n8n. If both the intent confidence and the ICP confidence were at least 0.85, the lead went straight through. High intent at a good fit went to Salesforce. Exploratory interest, or high intent from a weaker fit, went to a Marketo nurture. Spam and unqualified leads went to a CRM disqualify bucket. If either confidence fell below 0.85, the lead went to a Slack queue for a person to look at.

One detail worth knowing: Jev's confidence score is stricter than it looks. To reach 0.85 confidence on intent, Jev had to be about 89% sure of its answer. That's a high bar, and it explains a lot of what follows.
The Run
Jev processed all 10,000 leads in 1 minute 24.7 seconds on a local n8n instance, about 118 leads a second. Every lead came back with a valid, typed response. At TypeSafe's published price of $0.042 per million input tokens, the whole run cost about 23 cents.

Because the synthetic recipes record what each lead really was, I could check every decision. Of the 7,402 leads the gate let through, none went to the wrong place. Without the gate, 98.5% would have been routed correctly. The gate caught the other 1.5%, and it also held back a lot of leads that Jev had gotten right.
The 26 Percent
The number I keep coming back to isn't the speed or the cost. It's the 2,598 leads, roughly one in four, that Jev wasn't sure enough about.
They weren't mostly junk. Nearly half of the genuine enterprise buyers, 966 of 1,951, landed in the review queue instead of Salesforce. That's why fewer than 10% of leads reached sales directly when almost 20% of the list was high intent. The gate was most cautious about exactly the leads that mattered most.
The pattern wasn't random either. Two of the high-intent messages went to review every single time: "our chat models keep breaking" and "current lead scoring is broken, need a fast deployment." Read cold, both sound a little like support tickets. "Budget approved for new GTM tooling, let us talk" went to review 91% of the time. My High Intent definition read: "Explicitly mentions competitive displacement, needing a demo, or consolidating stack." That message does none of those things. My best guess is that Jev took my wording at face value, wavered between High Intent and Exploratory, and the gate did what I'd told it to do. When I later tested the definition without "Explicitly" on a harder set of leads, it helped less with quiet buyers than I expected, so the wording is probably only part of the story.
So the review queue became, at least partly, a map of where my instructions were vague. That was more useful than I expected. If I were deploying this for real, I'd read that queue for a week before touching the threshold, because a lot of what sits in it is feedback on the criteria, not just on the model.
The Baselines
My first draft argued that I didn't need a control group, because everyone in revenue operations (RevOps) already knows how keyword rules and chat models behave. That was the weakest part of the piece. So I ran one.
The first result was humbling. On the 10,000 synthetic leads, a plain regex rule set (keywords like "demo," "budget" and "move off," plus simple title and company-size rules) routed every single lead correctly. The synthetic data was too easy to tell the methods apart.
So I built a harder test: 100 hand-labelled leads where meaning matters and keywords mislead. Buyers who never say "demo" ("our contract renews in March and honestly we are not happy"). Real buyers whose messages start with "asdf" or "test." Existing customers complaining in buying language ("we bought 200 seats last year and the demo workspace you set up for us still won't load"). Senior titles with nothing more than idle curiosity. Students who paste in high-intent phrases word for word. Then I ran every lead through Jev with the exact same request as before, through the regex, and through two Claude models and two GPT-6 models, the large language models (LLMs) behind Claude and ChatGPT. The LLMs got the same criteria, returned structured data (JSON), and were asked for their own confidence in each answer.
In the table, "Skipped review" is the share of leads confident enough to pass the 0.85 gate, and "Accuracy when skipped" is how many of those were routed correctly.

Not all of this flatters Jev.
Claude Opus was the most accurate model in the test. On a sample of 100, the gap between Opus at 91%, Jev at 88% and Haiku at 87% is within the margin of error, but if it leans anywhere it leans toward Opus. Jev's advantage wasn't raw accuracy. It reached the same range as a strong general model at about 1/300th of what Opus cost and around ten times the speed.
I also expected the chat models to break their JSON now and then. They didn't. With structured output turned on, all 400 LLM calls came back valid. That argument against generative models is weaker than it used to be.
The regex did what regexes do. It never hesitated and was wrong 41 times, including six genuine buyers sent straight to disqualify. It caught half of the buyers who didn't use trigger words, and about one in five support complaints written in buying language.
The bigger GPT model did slightly worse than the smaller one. Both seem to have read "explicitly mentions" very literally and called most of the quiet buyers exploratory. Astra got just 6 of those 16 leads right.
The result I find most interesting is what happened at the gate. Every model gave me a confidence number, but the numbers meant very different things.

One caveat first: these aren't quite the same kind of number. Jev's confidence is calculated from its own probabilities. The LLMs' confidence is a number they write down when asked. And the 0.85 bar was set for Jev in the first place, which works against the LLMs.
Jev's lined up with reality. On average it put 0.88 probability on the intent it chose, and it was right 84% of the time.
The Claude models were underconfident. Opus reported 0.78 on average and was right 88% of the time, so the 0.85 gate blocked almost everything it said. Its judgment wasn't the problem. Every lead where Opus gave 0.6 or more on both questions, 64 of them, was routed correctly. You'd just have to find the right threshold yourself, and re-find it whenever the model changes.
The GPT-6 models went the other way. They reported around 0.9 and were right about 70% of the time, and they were the only models that pushed leads through the gate to the wrong destination, five each. For an automated router, that is the failure that hurts: confident, wrong, and nobody looking.
This is the finding that really stands out to me from this test. Accuracy is critical in driving pipeline automation. Confidence you can trust is what lets a team switch off the review step for part of the queue and stop double-checking.
I ran one smaller check as well. Removing the word "Explicitly" from the High Intent definition took Jev from 88% to 91%, but mostly on the support complaints and opt-outs, not the quiet buyers I expected it to help. On 100 leads, that's inside the noise anyway.
What I'd Change Next Time
The 0.85 threshold was too strict for hard leads. On the 100-lead test, only 10 of the 27 genuine buyers reached Salesforce at 0.85. At 0.7, 17 did, with one mistake among the 53 leads that skipped review. I picked that comparison using the same 100 leads I scored it on, so treat it as a hint, not a setting.
Before lowering the gate, I'd rewrite the criteria. A good share of what the review queue caught looked like my own vague wording. And I'd run the whole thing on real inbound leads, where the tells aren't planted by a script.
The Verdict
I went in expecting this to be a story about speed, cost and determinism: a classification model that gives you the same answer every time, set against general-purpose LLMs that might not. Speed and cost held up. Jev was about ten times faster than the LLMs, and running it on 10,000 leads cost less than a cup of coffee.
Determinism held up only partly. When I sent Jev the same 100 leads twice, it picked the same intent 98 times, but its confidence scores shifted, by as much as 0.17 on one lead. Consistent, yes. Deterministic, no. I didn't re-run the LLMs, so I can't yet say whether Jev was more consistent than they were.
The story I didn't expect was about confidence, and knowing when not to trust the machine.
It's easy to think of AI mainly as a way to generate more text. In go-to-market, a tool like Jev earns its place by reducing noise, and by telling you honestly which part of the noise it couldn't sort. When sales development reps (SDRs) are buried under thousands of ambiguous form fills, they fall back on generic automated sequences just to keep up, and there's no room left for a real conversation. A triage layer that clears out the junk, sends the clear buyers straight through, and hands a person a short, honest queue of "I'm not sure" protects the time a rep needs to write something genuine.
This was one rough test, on synthetic data and a hundred hand-built leads, so I'd treat it as a pointer rather than proof. But it changed how I'd build this. I'd stop asking which model is smartest and start asking which one I can set a rule on.
How this test was run
The 10,000 leads are synthetic. The 100 harder leads were written and labelled for this test. I drafted them with help from an AI assistant and reviewed every label before scoring.
I labelled opt-out messages ("remove me from your list") as Spam so they route to disqualify. The LLMs mostly called them Support, which is a fair reading. Excluding those 8 leads, Claude Opus scores 97%, Claude Haiku 93% and Jev 89%.
On 100 leads the margins are wide. The 95% ranges for routing accuracy are Jev 80-93%, Claude Opus 84-95% and Claude Haiku 79-92%. Calibration error, how far stated confidence sits from actual accuracy (lower is better), was about 0.06 for Jev, 0.15 for both Claude models and 0.19 for both GPT-6 models.
The regex rules were written from the 10,000-lead vocabulary and frozen before the hard set was scored.
The LLMs ran with their default reasoning settings and structured JSON output. Gemini was left out of this round.
Costs come from actual token usage at prices published in October 2026, scaled from 100 leads to 10,000. Response times were measured from my laptop with four requests in flight at a time.
Every raw API response is saved, and every number here can be re-derived from them.
AI in Action: Testing Jev, an AI Built for Judgment Calls, on 10,000 Sales Leads
With the "this changes everything!" generative AI arms race continuing interminably, I found myself intrigued (& slightly puzzled) by TypeSafe's recent release of Jev, a model that wasn't actually built to generate anything. New AI models arrive every few weeks, and most of them are variations on one idea: a faster or cheaper way to generate text. Jev is instead built to make judgment calls. You hand it some context and a set of questions, and it hands back typed answers with probabilities attached. What's more, it does so at a fraction of the time and cost (input is cheap, output is free) of standard models.
As I scratched my head, looking to figure out where this might be useful in real-world use cases, my mind went straight to B2B sales and marketing. So much of the daily work in a revenue team is exactly that kind of judgment. Is this lead worth a rep's time, or is it someone browsing on a Friday afternoon? Lead scoring and pipeline routing both run on calls like that, and today they're usually made by brittle rules or by people with too little time. A model designed to classify rather than to chat sounded like it might fit that work better than the tools most teams are reaching for.
It also sits close to why The Authentic .AI exists. The lab's tagline is "Digitally Powered. Authentically Human." A lot of the AI showing up in sales and marketing right now adds noise: more emails, and more content that sounds like everyone else's. I'm more interested in AI that clears space for the human part of the job. A model that quietly sorts the inbound pile, and admits when it isn't sure, might do exactly that.
So this experiment takes one real-world use case, inbound lead triage, and tries to learn something honest about where Jev is useful and where it isn't. It's not a product review and it's not a formal benchmark. It's one rough test, written up in full.
The Hypothesis
Generative AI models are built to write. Sales systems are built to route. When go-to-market (GTM) teams ask a chat model to score a big inbound list, they are borrowing a tool designed for one job and pointing it at another. It can work. But every answer comes back as text that has to be parsed, every call takes seconds, and the model's sense of how sure it is tends to be an afterthought.
TypeSafe calls Jev a "System One" model, built for fast judgment calls rather than long reasoning. My hypothesis was that a layer like this could route a high-volume inbound queue reliably, and that its confidence scores would tell me which leads were safe to automate. The second question was whether it would actually beat the cheaper or more familiar options once the leads got harder.
The Tech Stack
Data generation: Python (Faker library)
Orchestration: n8n 2.8.4, self-hosted on my laptop
Classification: TypeSafe Jev 1.13.0
Baselines: a regex rule set (simple keyword matching), Claude Haiku 5.5, Claude Opus 5.5, GPT-6 Luna and GPT-6 Astra
Step 1: Synthesizing the Pipeline Mess
To test at volume, I needed volume. I wrote a Python script to generate 10,000 synthetic inbound leads, each with a name, job title, company size and a free-text reason for reaching out.
The generator used three recipes. 3,141 leads (31.4%) were obvious junk: students, interns, messages like "asdf" or "need for class." 4,908 (49.1%) were ambiguous mid-market interest. The remaining 1,951 (19.5%) were enterprise buyers writing things like "we need to move off Salesforce" or "need a demo for my team this week."
One thing to note: this data was easy. Each recipe had an obvious tell, and that ended up shaping what the experiment could and couldn't prove.
Step 2: Defining the Schema
I loaded the CSV file into n8n, looped over it in batches of 100, and sent each lead to Jev's API as one line of context: name, title, company size and message. Each request asked three typed questions at once.
ICP fit, how closely a lead matches the ideal customer profile, was a Score question on four levels, from 0 (students, freelancers, companies under 50 people) to 3 (an enterprise VP of Sales, CRO or Head of Sales Systems). Jev returns a probability-weighted position on that scale, which I rounded before routing.
Intent was a Choice question with four options: High Intent, Exploratory, Support and Spam.
Sales action required was a yes-or-no likelihood question (TypeSafe calls it a Noul), which returns a probability from 0 to 1 that a statement is true. I ended up not using it. Genuine buyers scored anywhere from 0.49 to 0.83 on it, and some ambiguous leads reached 0.51, so there was no clean line to route on.
Step 3: Building the Confidence Gate
This is where the experiment got operationally interesting. A chat model gives you an answer. Jev gives you an answer, a probability for every possible answer, and a confidence number that summarizes how concentrated those probabilities are.
Right after the API call, I put a gate in n8n. If both the intent confidence and the ICP confidence were at least 0.85, the lead went straight through. High intent at a good fit went to Salesforce. Exploratory interest, or high intent from a weaker fit, went to a Marketo nurture. Spam and unqualified leads went to a CRM disqualify bucket. If either confidence fell below 0.85, the lead went to a Slack queue for a person to look at.

One detail worth knowing: Jev's confidence score is stricter than it looks. To reach 0.85 confidence on intent, Jev had to be about 89% sure of its answer. That's a high bar, and it explains a lot of what follows.
The Run
Jev processed all 10,000 leads in 1 minute 24.7 seconds on a local n8n instance, about 118 leads a second. Every lead came back with a valid, typed response. At TypeSafe's published price of $0.042 per million input tokens, the whole run cost about 23 cents.

Because the synthetic recipes record what each lead really was, I could check every decision. Of the 7,402 leads the gate let through, none went to the wrong place. Without the gate, 98.5% would have been routed correctly. The gate caught the other 1.5%, and it also held back a lot of leads that Jev had gotten right.
The 26 Percent
The number I keep coming back to isn't the speed or the cost. It's the 2,598 leads, roughly one in four, that Jev wasn't sure enough about.
They weren't mostly junk. Nearly half of the genuine enterprise buyers, 966 of 1,951, landed in the review queue instead of Salesforce. That's why fewer than 10% of leads reached sales directly when almost 20% of the list was high intent. The gate was most cautious about exactly the leads that mattered most.
The pattern wasn't random either. Two of the high-intent messages went to review every single time: "our chat models keep breaking" and "current lead scoring is broken, need a fast deployment." Read cold, both sound a little like support tickets. "Budget approved for new GTM tooling, let us talk" went to review 91% of the time. My High Intent definition read: "Explicitly mentions competitive displacement, needing a demo, or consolidating stack." That message does none of those things. My best guess is that Jev took my wording at face value, wavered between High Intent and Exploratory, and the gate did what I'd told it to do. When I later tested the definition without "Explicitly" on a harder set of leads, it helped less with quiet buyers than I expected, so the wording is probably only part of the story.
So the review queue became, at least partly, a map of where my instructions were vague. That was more useful than I expected. If I were deploying this for real, I'd read that queue for a week before touching the threshold, because a lot of what sits in it is feedback on the criteria, not just on the model.
The Baselines
My first draft argued that I didn't need a control group, because everyone in revenue operations (RevOps) already knows how keyword rules and chat models behave. That was the weakest part of the piece. So I ran one.
The first result was humbling. On the 10,000 synthetic leads, a plain regex rule set (keywords like "demo," "budget" and "move off," plus simple title and company-size rules) routed every single lead correctly. The synthetic data was too easy to tell the methods apart.
So I built a harder test: 100 hand-labelled leads where meaning matters and keywords mislead. Buyers who never say "demo" ("our contract renews in March and honestly we are not happy"). Real buyers whose messages start with "asdf" or "test." Existing customers complaining in buying language ("we bought 200 seats last year and the demo workspace you set up for us still won't load"). Senior titles with nothing more than idle curiosity. Students who paste in high-intent phrases word for word. Then I ran every lead through Jev with the exact same request as before, through the regex, and through two Claude models and two GPT-6 models, the large language models (LLMs) behind Claude and ChatGPT. The LLMs got the same criteria, returned structured data (JSON), and were asked for their own confidence in each answer.
In the table, "Skipped review" is the share of leads confident enough to pass the 0.85 gate, and "Accuracy when skipped" is how many of those were routed correctly.

Not all of this flatters Jev.
Claude Opus was the most accurate model in the test. On a sample of 100, the gap between Opus at 91%, Jev at 88% and Haiku at 87% is within the margin of error, but if it leans anywhere it leans toward Opus. Jev's advantage wasn't raw accuracy. It reached the same range as a strong general model at about 1/300th of what Opus cost and around ten times the speed.
I also expected the chat models to break their JSON now and then. They didn't. With structured output turned on, all 400 LLM calls came back valid. That argument against generative models is weaker than it used to be.
The regex did what regexes do. It never hesitated and was wrong 41 times, including six genuine buyers sent straight to disqualify. It caught half of the buyers who didn't use trigger words, and about one in five support complaints written in buying language.
The bigger GPT model did slightly worse than the smaller one. Both seem to have read "explicitly mentions" very literally and called most of the quiet buyers exploratory. Astra got just 6 of those 16 leads right.
The result I find most interesting is what happened at the gate. Every model gave me a confidence number, but the numbers meant very different things.

One caveat first: these aren't quite the same kind of number. Jev's confidence is calculated from its own probabilities. The LLMs' confidence is a number they write down when asked. And the 0.85 bar was set for Jev in the first place, which works against the LLMs.
Jev's lined up with reality. On average it put 0.88 probability on the intent it chose, and it was right 84% of the time.
The Claude models were underconfident. Opus reported 0.78 on average and was right 88% of the time, so the 0.85 gate blocked almost everything it said. Its judgment wasn't the problem. Every lead where Opus gave 0.6 or more on both questions, 64 of them, was routed correctly. You'd just have to find the right threshold yourself, and re-find it whenever the model changes.
The GPT-6 models went the other way. They reported around 0.9 and were right about 70% of the time, and they were the only models that pushed leads through the gate to the wrong destination, five each. For an automated router, that is the failure that hurts: confident, wrong, and nobody looking.
This is the finding that really stands out to me from this test. Accuracy is critical in driving pipeline automation. Confidence you can trust is what lets a team switch off the review step for part of the queue and stop double-checking.
I ran one smaller check as well. Removing the word "Explicitly" from the High Intent definition took Jev from 88% to 91%, but mostly on the support complaints and opt-outs, not the quiet buyers I expected it to help. On 100 leads, that's inside the noise anyway.
What I'd Change Next Time
The 0.85 threshold was too strict for hard leads. On the 100-lead test, only 10 of the 27 genuine buyers reached Salesforce at 0.85. At 0.7, 17 did, with one mistake among the 53 leads that skipped review. I picked that comparison using the same 100 leads I scored it on, so treat it as a hint, not a setting.
Before lowering the gate, I'd rewrite the criteria. A good share of what the review queue caught looked like my own vague wording. And I'd run the whole thing on real inbound leads, where the tells aren't planted by a script.
The Verdict
I went in expecting this to be a story about speed, cost and determinism: a classification model that gives you the same answer every time, set against general-purpose LLMs that might not. Speed and cost held up. Jev was about ten times faster than the LLMs, and running it on 10,000 leads cost less than a cup of coffee.
Determinism held up only partly. When I sent Jev the same 100 leads twice, it picked the same intent 98 times, but its confidence scores shifted, by as much as 0.17 on one lead. Consistent, yes. Deterministic, no. I didn't re-run the LLMs, so I can't yet say whether Jev was more consistent than they were.
The story I didn't expect was about confidence, and knowing when not to trust the machine.
It's easy to think of AI mainly as a way to generate more text. In go-to-market, a tool like Jev earns its place by reducing noise, and by telling you honestly which part of the noise it couldn't sort. When sales development reps (SDRs) are buried under thousands of ambiguous form fills, they fall back on generic automated sequences just to keep up, and there's no room left for a real conversation. A triage layer that clears out the junk, sends the clear buyers straight through, and hands a person a short, honest queue of "I'm not sure" protects the time a rep needs to write something genuine.
This was one rough test, on synthetic data and a hundred hand-built leads, so I'd treat it as a pointer rather than proof. But it changed how I'd build this. I'd stop asking which model is smartest and start asking which one I can set a rule on.
How this test was run
The 10,000 leads are synthetic. The 100 harder leads were written and labelled for this test. I drafted them with help from an AI assistant and reviewed every label before scoring.
I labelled opt-out messages ("remove me from your list") as Spam so they route to disqualify. The LLMs mostly called them Support, which is a fair reading. Excluding those 8 leads, Claude Opus scores 97%, Claude Haiku 93% and Jev 89%.
On 100 leads the margins are wide. The 95% ranges for routing accuracy are Jev 80-93%, Claude Opus 84-95% and Claude Haiku 79-92%. Calibration error, how far stated confidence sits from actual accuracy (lower is better), was about 0.06 for Jev, 0.15 for both Claude models and 0.19 for both GPT-6 models.
The regex rules were written from the 10,000-lead vocabulary and frozen before the hard set was scored.
The LLMs ran with their default reasoning settings and structured JSON output. Gemini was left out of this round.
Costs come from actual token usage at prices published in October 2026, scaled from 100 leads to 10,000. Response times were measured from my laptop with four requests in flight at a time.
Every raw API response is saved, and every number here can be re-derived from them.
The Authentic .AI
Sign up to our newsletter
© The Authentic .AI 2026. All rights reserved.
The Authentic .AI is an independent research lab exploring how AI can stay authentically human. New thinking every Monday.
The Authentic .AI
Sign up to our newsletter
© The Authentic .AI 2026. All rights reserved.
The Authentic .AI is an independent research lab exploring how AI can stay authentically human. New thinking every Monday.
The Authentic .AI
Sign up to our newsletter
© The Authentic .AI 2026. All rights reserved.
The Authentic .AI is an independent research lab exploring how AI can stay authentically human. New thinking every Monday.


