Jev as a judge (TypeSafe AI): accept when confident, escalate when unsure
When the pool checked search suggestions on , Google offered "jev llm as a judge" and "jev as a judge". Bing and DuckDuckGo also offered "jev as a judge", and all three offered "jev evals". A suggestion shows that people search for something; it does not show how many. This page reads the source code of four tools that use Jev to grade or judge content. It shows where an item ends up when its Jev call fails, what each tool does with an unsure verdict, and what the three published judge studies report. It ends with a procedure: accept the verdict when Jev is confident, escalate it when Jev is unsure, and count every failure.
Of 4 Jev judge tools read in source, DeepEval stops the run on a Jev error; the other 3 drop the item from the score, and openlayer jevals still reports it as passed.
Picked from the judge tools named in earlier pool research (DeepEval, Jevals, openlayer jevals) and two projects found in this run's 24-hour scan that use Jev to judge content; not a survey (how the four were chosen). The pool installed none of them, ran no eval and made no Jev call. The Jevals harness is listed too, but its code is still not public.
Where a failed item ends up
One item, one Jev call that fails (HTTP error or timeout after retries), default settings
-
DeepEval JevEval
a200ece· library metric- Jev erroror timeout (180 s per task, async path)
- then
- Run abortsfail-closed: no score for the run. With
ignore_errors=Truethe test case is a FAIL
-
openlayer jevals
0a8f895· library- Jev error15 s per attempt, 4 retries
- then
- Dropped from the scoremissing score, left out of the denominator, counted under
errors - then
- Reported as passedfail-open: the report's
passedtreats an error as not a failure. The runtime gate allows by default
-
did-they-answer
e0ee60c· daily judging job- Jev error60 s, 5 tries on 429 or 5xx
- then
- Dropped from the scoreno-decision: the item is not written and is retried on the next daily run
-
competitor-hunter
40613ea· local app- Jev errorcaller default 60 s, 4 retries
- then
- Dropped from the rankingno-decision, while errors stay at or below 5% of items
- Run aborts above 5%fail-closed:
MAX_JEV_ERROR_SHARE = 0.05
- 1/4run aborts by default (DeepEval)
- 3/4drop the item from the score (openlayer, did-they-answer, competitor-hunter)
- 1/4count it as a pass (openlayer's report
passed; its gate also allows) - 0/4count it as a fail by default (DeepEval only with
ignore_errors=True)
errors); did-they-answer and competitor-hunter do not show one next to the score. Colours follow the shared failure vocabulary: green fail-closed, red counts as a pass, gray no-decision. Documented source at pinned commits, 4 Oct 2026.Four judge tools and the Jevals harness: what each does when Jev fails or is unsure
Open a situation to highlight its column. Opening one closes the others. Every cell for the four tools comes from source at the commit shown; the Jevals row is Reported from its methodology page because its harness code is not public. Failure terms use the pool's shared vocabulary: fail-closed (stop or refuse), fail-open (counts as a pass or acts anyway), fallback-<what> (a stated substitute), no-decision (no result; the caller must handle it). Below a threshold: held, acts-anyway, no-threshold. Documented Read 4 Oct 2026, 05:15–05:19 UTC; spot-checked 05:20–05:26 UTC
The Jev call failsHTTP error or timeout
DeepEval stops the run by default (1 of 4). openlayer, did-they-answer and competitor-hunter drop the item from the score (3 of 4); competitor-hunter also stops the run when more than 5% of items error. openlayer's report still says passed (1 of 4). None of the 4 counts the item as a fail by default.
Malformed answerMissing answer or fields, non-numeric values
DeepEval raises, which is handled like an error. openlayer skips a missing answer and reads a non-numeric Noul as 0.5. did-they-answer drops an item with missing usage, but missing answer keys stop the run (traced, not executed). competitor-hunter marks the item unsure; what happens next was not recorded.
Jev is unsureBelow a confidence threshold, and escalation
Escalation for unsure verdicts is coded in 1 of 4 (openlayer, off by default); 0 of 4 escalate by default. 3 of 4 have no confidence cut (no-threshold). competitor-hunter flags unsure items and still ranks them (acts-anyway, 1 of 4).
No API keyand fallbacks
All 4 refuse to run when Jev is named explicitly (fail-closed, 4 of 4). openlayer's auto-detect mode tries other backends and then an LLM (gpt-4.1-mini).
| Tool (commit) | What is judged; question type | What is sent | Jev error or timeout | Malformed answer | Below threshold; escalation | No key | Tests |
|---|---|---|---|---|---|---|---|
DeepEval JevEval (a200ece)Version 4.2.8 in pyproject.toml and on PyPI; no matching git tag (newest tag v4.1.8). Library metric, Apache-2.0 | Your own Noul, Score and Choice questions about test-case fields or an agent trace, combined into a weighted mean from 0 to 1 | Chosen test-case fields or the compacted trace, plus your question texts. No truncation; a pre-check raises above about 32k state or 64k request tokens | fail-closed: run abortsDefault ignore_errors=False. With ignore_errors=True the test case is a FAIL. 180 s per task (async path) | fail-closedA missing answer raises and is handled like an error | no-thresholdMinimum confidence is recorded, not acted on. Pass = score ≥ your threshold (default 0.5). Escalation: not coded | fail-closed"JevEval needs Jev in every mode; there is no LLM to fall back to" | Yes, mocked (FakeSystemOneModel) |
openlayer jevals (0a8f895)Version 0.1.4 in pyproject.toml (PyPI not checked). Evals plus runtime gates, MIT | Agent, RAG-quality and security evals; Noul, Choice and Score | Per eval. Example: each context cut to 3,000 characters; Correctness sends question 1,500, answer 4,000 and reference 4,000 characters | fail-open: counts as passedThe item becomes an error result. Dataset summary: missing score, left out of the denominator, counted under errors (no-decision). The report's passed treats an error as not a failure; the runtime gate defaults to on_error="allow". 15 s per attempt, 4 retries | fallback-0.5A non-numeric Noul is read as 0.5; a missing answer is skipped | no-threshold (default)Evals: score cut only. Gate: held only if you set escalate_below or escalate_above (default none) and pass an on_escalate handler; without a handler the tool runs. Escalation: yes, opt-in | not recorded (default branch)Jev backend named: error (fail-closed). Auto-detect: other backends, then fallback-LLM (gpt-4.1-mini). Which of the two applies when no backend is passed was not recorded | Yes, mocked. A test asserts "errors are not failures" |
did-they-answer (e0ee60c)Daily judging job plus a site; created 4 Oct 2026. No licence in the repository | Irish parliamentary written answers: Choice clarity (3 options), Choice technique (9 options), Noul "referred" | Full question and reply; no truncation | no-decision: droppedThe item is not written and is retried on the next daily run. 60 s, 5 tries on 429 or 5xx | fail-closedMissing answer keys fail outside the error handler, so the run aborts (traced, not executed). A missing usage field alone: no-decision | no-thresholdThe top option is stored; "referred" is split at 0.5. Escalation: not coded | fail-closedExits without a key; with only an OpenRouter key it uses the same Jev model through OpenRouter | No unit tests in the tree |
competitor-hunter (40613ea)Local app; created 3 Oct 2026, MIT. Loads the user's installed jev.py (claude-x-jev c8662b0 read), so the Jev caller version is not pinned | Competitors' social posts: Choice topic and hook type; Score ICP fit and proof (4 levels); Noul replicable, sponsored, newsjack | Creator profile, platform, format, duration; title up to 300 and description up to 700 characters | no-decision: droppedErrored items are left out of the ranking. Above 5% errors the run aborts (fail-closed). Caller default 60 s, 4 retries | not recordedMissing fields become empty and the item is marked unsure; later rank arithmetic is unguarded; the outcome was not determined | acts-anywayUnsure items are flagged and counted but still ranked. Cuts from the preset: 0.6, 0.5, 0.7. Escalation: flag only | fail-closedOpenRouter key required by the caller | None |
| Jevals harnessHarness code still not public. Jevals/Jevals has a README only; jevals-data has data and no code. Suite 0.1.0 | Reported Noul, Choice and Score boards graded against human labels | not recorded | not recorded in codeReported transport failures "are retried with backoff and never scored"; the run resumes them | not recorded in codeReported retried twice, then scored as uniform and a wrong pick | not recorded | not recorded | Harness code not public |
- fail-closed: stops the run or refuses
- fail-open: counts as a pass, or acts anyway
- no-decision: no score; the caller must handle it
- fallback-<what>: a stated substitute value
- no-threshold, or not recorded (the cell says which)
Change, 7 Oct 2026: three cells that read "conditional" (openlayer below threshold and no key, did-they-answer malformed) now show the term counted on When Jev fails, with the condition kept in the note: the default configuration counts. The facts in the cells are unchanged. Rechecked 7 Oct 2026: deepeval 4.2.8 and jevals 0.1.4 are still the latest releases.
Three details that change a score
- "Not applicable" can score 1.0. In DeepEval's JevEval, a test case where every Choice question is 0.5 or more "not applicable" scores 1.0.
- openlayer's gate allows on error. The runtime gate defaults to
on_error="allow", so a Jev failure lets the guarded tool run. Changingon_erroris covered in step 6 of the procedure. - Escalation is rare in code. 1 of 4 tools codes escalation for unsure verdicts (openlayer, an
"escalate"action and anon_escalatehook), and it is off by default. 0 of 4 escalate by default.
Other counts over the 4 tools read in source: no key fails closed in 4 of 4 when Jev is named; mocked tests in 2 of 4 (DeepEval, openlayer), none in 2 of 4. The Jevals row is not in any count.
Evidence summary: what the judge studies report
The pool ran no judge benchmark and no eval. The three rows below are copied unchanged from the graded benchmarks ledger (graded 30 Sep 2026, run 3), with their row IDs. They are not averaged or re-ranked; read each one on the ledger with its materials. Reported
| Row and operator | Task, n and baseline | Result as reported | Grade |
|---|---|---|---|
| B02 Arize, "Jev vs LLM-as-a-Judge" (23 Sep) Arize AI, an AI observability and evaluation vendor; no Jev stake stated. | Hallucination detection (RAGTruth, 2,675 responses) and summary quality (SummEval, 1,700 summaries). Human labels. 23,325 judgments in total. Baselines: Claude Opus 5 and GPT-5.6 Terra at high reasoning effort. | RAGTruth at a 0.5 cutoff: Opus 5 83%, Jev 76%. With cutoffs tuned on half the data: Jev 87%, Opus 5 87%, Terra 80%; "the intervals overlap". SummEval rank correlation: Jev 0.74, Opus 5 0.70. $0.05 vs $14.30 per 1,000 judgments. | B public-data |
| B06 Jevals, release 2026-09-18 One unnamed maintainer; states no affiliation with TypeSafe, gateways or labs and pays list price. | PubMedQA (yes/no), Banking77 (77 intents), HelpSteer2 (5-level score). Human labels; 300 items × 5 runs per task. Via Vercel AI Gateway, ai@7.0.106.Baselines: Gemini 3.8 Flash, GLM-5.3, DeepSeek V4.1 Flash and others; a label-prior baseline. | PubMedQA: "tied" with Gemini 3.8 Flash (Decision Score 69.0 vs 73.0; 95% range of the difference −2.5 to +10.4); accuracy 91.3%; calibration error 5.0 points; $0.029 vs $0.80 per 1,000. Banking77: Gemini "significantly ahead". HelpSteer2: "no model clearly beats answering with the label base rates". | B public-data one row recomputed |
| B08 arXiv 2609.26550 v3, "JEV-as-a-Judge" (v3 29 Sep) Carnegie Mellon University; thanks OpenAI's Researcher Access Program for API credits. | Judging: 5,172 base judgments (RewardBench, JudgeBench, HaluEval, final-answer checks); blinded human adjudication of disputes. Baselines: 16 judges, including GPT-6 Astra and two reward models. | "within three points of GPT-6 wherever a verdict can be read off the text, at 0.36% of its fee"; behind "where the verdict must be derived, as in math, code, and logic"; 12 points behind on JudgeBench. A cascade with a threshold frozen in advance escalates 31% of 1,610 pairs and is 0.9 points more accurate than GPT-6 at 41% of its fee. | B public-data v3 replaced earlier results |
Procedure: accept when confident, escalate when unsure
This is a design reading of the source above, not a measured result. It gives no cutoff number: a threshold has to be set on your own labelled sample (how to choose a confidence threshold).
- Decide what an error means before you run.DeepEval: stops the runopenlayer: drops, shows an error countdid-they-answer, competitor-hunter: drop, no count next to the score
Write down whether a failed Jev call makes the item a fail, a missing score or a stopped run. Do not drop it silently from the denominator: a score with dropped items covers fewer items than it seems to.
- Count errors separately, and record each one.openlayer:
For each failed item, record the item, what failed (HTTP error, timeout or malformed answer), how many attempts were made, and what happened to it. Report items scored out of items in the set (for example 480/500), the error count, and the escalation count next to the score. Below-threshold counts are not error counts; keep them apart. Retries and status codes: errors and rate limits.
errorscountothers: not shown - Accept when confident.DeepEval: records minimum confidence, does not act on itopenlayer gate: acts only if
Read Jev's probability or confidence for each verdict, not only the label. Accept the verdict when it clears the cutoff you set on your own data.
escalate_belowis set - Escalate when unsure.openlayer: opt-incompetitor-hunter: flags, still ranksDeepEval, did-they-answer: not coded
Send verdicts below your cutoff to a stronger LLM judge or a person, and count the escalations. B08's cascade, with a threshold frozen in advance, escalated 31% of 1,610 pairs; it is one Reported design, not a recommendation.
- Check how "not applicable" is scored.DeepEval
In DeepEval's JevEval, a test case where every Choice question is mostly "not applicable" scores 1.0.
- Never let an error become a pass.openlayer defaults
openlayer's report-level
passedand its default gateon_error="allow"both treat an error as not failing. Seton_errorto "block" or "escalate" if a judge failure must stop a release.
To test this path without a key or a live call, see testing without a key. To compare failure handling with the other use cases, see when Jev fails.
What belongs on this page and what does not
- Whether to judge with Jev, an LLM or a trained classifier at all: Jev vs an LLM vs a classifier.
- Graded accuracy figures and their materials: benchmarks. Rows are copied here unchanged, not re-graded.
- Failure handling compared across use cases: when Jev fails. From 7 Oct 2026 the four judge tools read in source (DeepEval JevEval, openlayer jevals, did-they-answer, competitor-hunter) are counted on When Jev fails as the "LLM as a judge" domain; the Jevals harness is not, because its code is not public. There, openlayer jevals is 1 of the 8 implementations (of 64) that let the action go ahead unchecked on a Jev error: an errored eval counts as passed.
- How to choose a cutoff: confidence thresholds.
How the four were chosen and read
The pool picked the judge tools named in its earlier research (DeepEval, Jevals and openlayer jevals) and two projects found in this run's 24-hour scan that use Jev to judge content (did-they-answer and competitor-hunter). This is not a survey, and no GitHub search was used for this page. openlayer jevals was presented on Hacker News on 20 Sep 2026 ("Show HN: jevals – replacing LLM judges with typed Jev decisions").
Each tool was pinned to one commit. A helper agent read the source on 4 Oct 2026, 05:15–05:19 UTC. The runner then spot-checked 12 line ranges in the 4 repositories at the same commits (05:20–05:26 UTC); all matched. The manager rechecked DeepEval's evaluate/configs.py L46 in its review (05:38 UTC). Nothing was installed or run, no eval was run and no Jev call was made. DeepEval JevEval and openlayer jevals are library features, so they are not counted as products; did-they-answer and competitor-hunter are added as rows to the products ledger in this release.
Not read: mchgood/jev-eval at a5ce01d (it benchmarks Jev as the model under test; Jev is not the grader), Zaious/JevTRPG and the texposit.com "Judge Jev" post (not reached in time), and DeepEval's ConversationalJevEval and TypeScript port (same metric family).
What was not verified
- Any eval result from these tools: none was run.
- openlayer jevals: which backend branch applies with no key when no backend is passed (a one-file read of
backends/__init__.py); its PyPI version was checked on 7 Oct (0.1.4). - DeepEval's timeout on the sync path (none is set at metric level) and its LLM "hybrid" fallback for built-in metrics (mentioned at
metrics/base_metric.pyL136–137; not read). - Which
jev.pyversion a competitor-hunter user has installed. - did-they-answer's abort on a malformed answer: traced in source, not executed.
- competitor-hunter's outcome on a malformed answer after it is marked unsure.
- The Jevals harness: its code is not public, so its error handling is Reported from its methodology page only.
- How many people use any of the four tools.
Sources and check times (4 Oct 2026, UTC)
- R7-S70, R7-S71: DeepEval at a200ece (commit 2 Oct 2026, 16:03 UTC; read 05:15–05:26). Paths under
deepeval/:evaluate/configs.pyL46;evaluate/execute/_common.pyL290–295;metrics/jev_eval/utils.pyL36–40, L51, L113–149, L194–197;metrics/jev_eval/jev_eval.pyL140–152, L226;models/system_one/limits.pyL64–81. Version 4.2.8 frompyproject.tomland PyPI JSON. - R7-S72: openlayer jevals at 0a8f895 (commit 1 Oct 2026, 16:48 UTC; read 05:15–05:26). Paths under
src/jevals/:_runner.pyL48–51, L160–166, L236–274;_gate.pyL84–90, L122–128, L156, L220–222;_eval.pyL17, L105–120;quality/_evals.pyL24–28, L211–216;backends/_wire.pyL41–42, L91–95;backends/__init__.pyL38–55. Test:tests/test_runner.pyL81. - R7-S73: Hacker News item 49780849, "Show HN: jevals" (20 Sep 2026), found through the HN search API.
- R7-S74: did-they-answer at e0ee60c (commit 4 Oct 2026, 02:15 UTC; read 05:15–05:26):
jev.pyL28, L31–64, L76–85, L93;daily.pyL54–60, L69;compact.pyL27–31. - R7-S75: competitor-hunter at 40613ea (commit 3 Oct 2026, 22:59 UTC; read 05:15–05:26):
engine/hunt.pyL175, L178–181, L527–533, L563, L576–583, L778, L806;jev/competitor-hunt.jsonL66–74. - R7-S76: claude-x-jev at c8662b0 (commit 25 Sep 2026):
skill/scripts/jev.pyL51–58, L158–184. - R7-S77: jevals.com/methodology (read 05:15–05:26). Reported
- R7-S78: Jevals repositories: Jevals/Jevals 15c0628 (README only) and jevals-data 21bb47b (35 tree entries, no code files); the Jevals account's 5 repositories, jevals.com home, /methodology and /changelog rechecked once.
- R7-S11: search suggestions from Google (
client=firefox), Bing (osjson), DuckDuckGo (ac) and YouTube at 05:21:29. "jevals" completes to "jewels" everywhere, so it has no support; YouTube offered nothing relevant. - Evidence rows B02, B06 and B08: copied from the benchmarks ledger (checked 30 Sep 2026, 20:53–21:00 UTC).