I began with a simple thesis:
For decisions, accuracy matters more than speed.
JEV is fast because it constrains a task to a bounded decision contract. I assumed that frontier LLMs would be more accurate, and that their extra latency would be an obvious trade worth making.
So I built a reproducible benchmark to test that assumption against JEV on the same decision contracts:
Choice)Noul)Score)The result was more interesting than “JEV wins” or “frontier models win.”
On intent routing, the best tested frontier models did outperform JEV:
|
Model |
Intent-routing accuracy |
|---|---|
|
JEV |
81.0% |
|
GPT-5.6 Luna |
85.6% |
|
GPT-5.6 Sol |
86.1% |
That supports my original concern. If a decision is wrong, being fast does not make it useful.
But the advantage did not generalize cleanly.
On HateCheck moderation, JEV reached 98.2%; the tested models ranged from 95.1% to 100.0%. On TREC search relevance, JEV’s 47.4% nearest-tier accuracy tied the best tested result.
So JEV was right about something I had underestimated: a bounded decision system can be very fast without automatically giving up meaningful accuracy.
For the relevance task, the input is intentionally small and explicit: a query, a candidate passage, and a rubric.
{
"state": {
"query": "what is the capital of france",
"passage": "Paris is the capital and most populous city of France."
},
"levels": [
"Not Relevant",
"Related",
"Highly Relevant",
"Perfect"
]
}
The model has to return both a continuous score and its belief across the four levels:
{
"score": 2.99,
"probabilities": {
"0": 0.0,
"1": 0.0,
"2": 0.01,
"3": 0.99
}
}
That response is coherent: the probability-weighted score is also approximately 2.99.
This is where my original framing broke down.
A model can return valid JSON, a plausible final score, and a plausible-looking probability distribution—yet have those two fields contradict one another:
{
"score": 2.99,
"probabilities": {
"0": 0.10,
"1": 0.60,
"2": 0.25,
"3": 0.05
}
}
The score says “almost Perfect.” The distribution says “mostly Related.” Both fields cannot be true at once.
That is not merely a formatting issue. It means a downstream system has to decide which answer to trust: the final decision, or the model’s stated uncertainty.
On the relevance experiment, score/distribution consistency was:
|
Model |
Consistency |
|---|---|
|
JEV |
100.0% |
|
GPT-5.6 Sol |
99.8% |
|
GPT-5.6 Luna |
98.6% |
|
DeepSeek V3.2 |
65.0% |
|
Claude Haiku 4.5 |
59.3% |
|
Claude Sonnet 5 |
14.4% |
These are results for the tested configurations, not universal claims about the models. Output mode and effort settings matter.
But they exposed a problem I was not measuring before.
A system needs at least three things:
I set out to show that bounded decisions were too limiting. Instead, I found that the hard problem is not only choosing the right answer—it is producing an answer you can consistently trust.
The benchmark, decision contracts, published comparison summaries, chart source, and export tooling are available here.
The published results are intended to be inspectable and reproducible—not taken as a universal leaderboard. You can use the same workflow with your own model configurations or decision datasets.