jev-1.13.0Jev benchmarks (TypeSafe AI): is Jev accurate? A graded ledger of published studies
In the three weeks after launch, TypeSafe, gateways, vendors, researchers and individuals published tests of Jev. They use different tasks, baselines and materials, so their headline numbers cannot be compared directly. This page grades each study by how far it can be checked, says who ran it and what they had at stake, and shows where the studies agree and where they conflict.
Why this title: Bing and DuckDuckGo complete "jev b" and "jev bank" to "jev benchmark"; Google completes "jev 1.13" to "jev 1.13 benchmarks" (rechecked ). A search suggestion is a demand signal, not a search volume.
Published tests disagree because they test different tasks against different baselines. Jev is usually ahead of cheap LLM setups that generate their answer as text; one study found open-weight models scored through their token probabilities about level with it overall (B35, grade B). It is close to frontier LLM judges on verdicts that can be read off the text once a threshold is tuned. It is behind the best LLMs on harder or derived judgments (math, code, logic) and on 14 of 15 social-science annotation tasks. Against trained classifiers on fixed label sets the evidence is mixed: in one Red Hat guardrail study, Jev was behind a prompt-injection classifier and ahead of a small content-safety classifier (B28, grade B; Red Hat publishes both classifiers); the rest is grade D.
This site has run no benchmark of its own and made no Jev call. Every result below is Reported: the named operator's own measurement. The site does not average results across studies and does not rank studies by headline score.
Before you quote "68%"
The figure quoted as "68% on TypeSafe's own benchmark" is TypeSafe's workflow benchmark (evals.typesafe.ai), where Jev scores 67.8%. That number is agreement with reference answers averaged from GPT-6 Astra and Claude Fable 5.1 on four workflows built by TypeSafe's own team. It is not accuracy against human labels or ground truth. GPT-5.6 workflows on the same test score 66.8–74.1%. See row B01. Reported
Can you publish Jev benchmark results? What TypeSafe's terms say
Current texts: no clause found (rechecked ). TypeSafe's current Terms of Use (last updated 19 Sep 2026), Acceptable Use Policy (23 Sep 2026), Master Customer Agreement, or MCA (23 Sep 2026), Data Processing Addendum (24 Apr 2026) and Privacy Policy (19 Nov 2025) contain no "benchmark", "performance information" or "comparison" clause. A search of each text for those words found no match; the only near hit is "comparable proceeding" in the MCA's insolvency clause, which is unrelated.
Docs legal page. docs.typesafe.ai/legal.md holds no terms text. Compared with the Internet Archive copy of 22 Sep 2026, it changed in two places only: the contact mailbox for enterprise zero data retention (ZDR) requests moved from TypeSafe's privacy address to its sales address, and a Mintlify footer line was added. The ZDR offer itself is unchanged. Data terms by route: data privacy.
Until 19 Sep 2026 the MCA had one. The MCA version marked "Last updated Aug 27, 2026", as archived on 16 Sep 2026 (21:55 UTC), listed among the customer's licence restrictions in section 2.3: "(f) publish benchmarks or performance information about the Services". The next archived version, dated 19 Sep 2026, no longer has it; nor do later captures or the live version dated 23 Sep. TypeSafe's docs call the MCA "the general terms that apply to your TypeSafe account".
What this means for the ledger. Studies published between Jev's launch on 15 Sep and 19 Sep appeared while, as far as the archive shows, the MCA contained 2.3(f): Parallel (B04, 18 Sep) and the first Jevals release (B06, 18 Sep). LiteLLM ran its test on 18 Sep and published on 20 Sep. TypeSafe itself now publishes a benchmark and code "to reproduce the results" (B01). This site draws no conclusion about any operator's compliance.
Still in the MCA (23 Sep version, quoted, not interpreted): 2.3(b), no use of the Services or Output "to perform model distillation, train a model to imitate the output of the Services, or develop (or to facilitate the development of) a similar or competing product or service"; 2.3(c), no reverse engineering; 2.3(g), no circumventing access restrictions "or conduct any security or vulnerability test".
Documented Read 30 Sep 2026, 20:48–20:51 UTC; rechecked 7 Oct 2026, 05:19 UTC (R10-S37, R10-S38; legal.md archived 22 Sep): Terms of Use, AUP, MCA, DPA, Privacy Policy, MCA archived 16 Sep (contains 2.3(f)), MCA archived 20 Sep (does not). This is a reading of the text, not legal advice.
Other claims about the terms, and what was checked
- A Hacker News comment of 18 Sep (item 49753778) said the Terms of Use (1(v)) and the MCA (2.3(f)) both prohibited publishing "benchmarks or performance information about the Services". The MCA part matches the 27 Aug text. The Terms-of-Use part is not confirmed: the archived Terms of 15 Sep (dated 14 Sep) have no such clause. Reported
- A Hacker News question of 23 Sep, "Is benchmarking Jev still a ToS violation?" (item 49813141): by the texts above, the clause was removed in the MCA version of 19 Sep.
- OpenRouter's terms make users agree to each model's terms; its Jev page points to TypeSafe's Terms of Use, not the MCA. Cloudflare's model page links
docs.typesafe.ai/legal.md. Vercel was not checked. - jevcal's statement that the terms restrict public testing was not re-read.
How each study is graded
- A Rerunnable
- Data, code and configuration published; the operator has no commercial stake in the result. "Provisional" means the linked code repository exists but this site did not read it.
- B Partly rerunnable
- Data public but code or prompts missing, or everything published but the operator has a stake (vendor, gateway, alternative-model author).
- C Method described
- Task, sample size and baseline stated; no materials published.
- D Claim only
- A headline number without task, sample size or baseline, or a secondary summary whose primary study is not identified.
Flags: COI the operator has a stake in the result · LLM-labels reference answers come from LLMs, not people or the source · small-n fewer than 200 items · no-baseline no LLM or classifier compared · sealed part of the test set is withheld · provisional code repository not read · public-data the data may be in training sets · single-task.
33 graded rows by grade, 7 Oct 2026 (select a row to jump to it)
- A Rerunnable6 of 33 rows: 3 provisional (B13, B15, B24), 3 read at pinned commits
- B Partly rerunnable12 of 33 rows
- C Method described8 of 33 rows
- D Claim only7 of 33 rows
Grade index, 7 Oct 2026: A 6 + B 12 + C 8 + D 7 = 33 graded rows (24 before this check, 9 added: B28–B36). Bars show each grade's share of the 33. Blue-bordered rows compare Jev with an LLM or a trained classifier. Grades describe how far a study can be checked, not whether its result is right. B14 moved from A (provisional) to B on 7 Oct.
Not graded: B22 (a secondary review), B25 and B27 (leads), and three items found on 7 Oct (see leads): jev-safety-eval has no results yet; jev-zh-tw-eval is not a Jev study (it tests Plumb-4B); the DAIR.AI piece is a tutorial, not a study. The reranking studies in Reranking studies are graded for that question only and are not in the 33.
Tier 1: studies read at their primary source
Reported Results in each operator's own units. "Materials" lists data / code / prompts / raw outputs. B01–B14 checked 30 Sep 2026 and B28–B36 checked 7 Oct 2026, at the UTC time shown; B14 regraded 7 Oct 2026. Model jev-1.13.0 unless stated; it is the only Jev version so far. Operators are named as organisations, sites or repositories.
| Study, operator and interest | Task, n and baseline | Result as reported | Materials | Grade and flags |
|---|---|---|---|---|
| B01 TypeSafe workflow evals, code WorkflowEvals (29 Sep) TypeSafe, the vendor. Its launch post says the workflows "were made by individuals on our model capabilities team, so some bias could exist". | Four business workflows split into Choice, Score and Noul questions: invoices (150 cases), customer service (204), agent traces (111), security incidents (240). Baselines: Claude Haiku 4.5, Opus 5, Sonnet 5; DeepSeek V4; GPT-5.6 Luna, Sol, Terra. Reference answers: average of GPT-6 Astra and Claude Fable 5.1. | Mean agreement with the LLM reference answers: Jev 67.8% at $0.0004 and 0.4 s per case; Sol workflow 74.1%, Opus 5 73.1%, Terra 67.9%, Sonnet 5 67.8%, Luna 66.8%, Haiku 4.5 53.6%. Not accuracy against human labels. The site's arithmetic cannot reproduce the homepage's "193.6x Faster, 444.6x Cheaper" from the rounded figures. Customer-service workflow (204 cases): Jev 76.0% (read 1 Oct 2026); how real ticket-triage projects handle Jev's failures: support-ticket triage. | Data yes; code yes (Apache-2.0, commit 0ac3b8a); prompts yes; outputs partly (on Hugging Face, not opened) | B COILLM-labels 20:48–20:49 UTC |
| B02 Arize, "Jev vs LLM-as-a-Judge" (23 Sep) Arize AI, an AI observability and evaluation vendor; no Jev stake stated. | Hallucination detection (RAGTruth, 2,675 responses) and summary quality (SummEval, 1,700 summaries). Human labels. 23,325 judgments in total. Baselines: Claude Opus 5 and GPT-5.6 Terra at high reasoning effort. | RAGTruth at a 0.5 cutoff: Opus 5 83%, Jev 76%. With cutoffs tuned on half the data: Jev 87%, Opus 5 87%, Terra 80%; "the intervals overlap". SummEval rank correlation: Jev 0.74, Opus 5 0.70. $0.05 vs $14.30 per 1,000 judgments. | Data public; no study code; prompts partly; outputs no | B public-data 20:53 UTC |
| B03 NavyaAI, "We Benchmarked Jev Directly" (21 Sep) An AI cost consultancy; "TypeSafe gave us API access". | AG News, 4 topics, n = 200 (sample not identified). English. Baselines: gpt-4.1-nano with one constrained token; gpt-4.1-nano and gpt-4o-mini in JSON mode; GPT-5.6 Luna. | Jev 87.5–88.0%; nano with one token 71.0%; small LLMs in JSON mode 82–83%; Luna 82%. Per million decisions: Jev $18.57, nano one-token $16.30 (cost detail on cost per decision). | Data public, sample not identified; code and prompts partly; outputs no | C public-datasingle-task 20:53 UTC |
| B04 Parallel, "Testing out Jev" (18 Sep, 21:07 UTC) A web-search API company that runs its own rerankers and classifiers. | Search reranking, topic classification, query freshness. Private data; n not stated. Baseline: Parallel's internal models (not named). | Reranking "NDCG@10 of 0.7: comparable to internal system"; topic and freshness: "Internal 'wins'". "Specialized classifiers will still often outperform Jev on cost and speed." | None | D no n 20:53 UTC |
| B05 LiteLLM, "JEV Classifier: 5.43x as Fast as Haiku" (20 Sep; run 18 Sep) An AI gateway that ships Jev as an Auto Router option. What LiteLLM's router and five others do when Jev fails, read in source: model routing. | Routing 80 synthetic prompts into three tiers; prompts and labels by the same author. 240 calls per model. Baseline: Claude Haiku 4.5, same rubric. | Match with the authored tiers: Jev 95.00% (228/240), Haiku 73.75% (177/240). Median latency 127 ms vs 688 ms; 96% lower registry-priced cost. The site's arithmetic reproduces all four ratios. | Cases and labels not published; no benchmark script linked | C COIsmall-nsingle-task 20:53 UTC |
| B06 Jevals, release 2026-09-18 One unnamed maintainer; states no affiliation with TypeSafe, gateways or labs and pays list price. | PubMedQA (yes/no), Banking77 (77 intents), HelpSteer2 (5-level score). Human labels; 300 items × 5 runs per task. Via Vercel AI Gateway, ai@7.0.106.Baselines: Gemini 3.8 Flash, GLM-5.3, DeepSeek V4.1 Flash and others; a label-prior baseline. | PubMedQA: "tied" with Gemini 3.8 Flash (Decision Score 69.0 vs 73.0; 95% range of the difference −2.5 to +10.4); accuracy 91.3%; calibration error 5.0 points; $0.029 vs $0.80 per 1,000. Banking77: Gemini "significantly ahead". HelpSteer2: "no model clearly beats answering with the label base rates". | Data public; per-decision logs yes (CC-BY-4.0); suites yes; harness code not found | B public-data 20:55–20:58 UTC · one row recomputed |
| B07 JevBench v1.5.4 (undated) Benchmark Heaven, a benchmark site that also sells custom evaluations; no Jev stake stated. | "State and a bounded rubric in, a typed answer out"; 1,624 decisions per system, 720 of them sealed. Ranking limited to "Jev-class" systems (at most 2× Jev's cost and latency); no LLM in it. | Jev 1.13.0 first by "Capability" 80.0 (Intelligence 72.0, Calibration 88.0), $0.032 per 1,000, median 0.62 s; Winnow-12B Q8 79.3; Cygnet 79.0. Laya and Kev research-preview rows: Jev vs Laya vs Kev. | Code repository linked, not opened; half the items sealed; aggregate results only for the sealed half | B sealedno-baseline 20:53 UTC |
| B08 arXiv 2609.26550 v3, "JEV-as-a-Judge" (v3 29 Sep) Carnegie Mellon University; thanks OpenAI's Researcher Access Program for API credits. | Judging: 5,172 base judgments (RewardBench, JudgeBench, HaluEval, final-answer checks); blinded human adjudication of disputes. Baselines: 16 judges, including GPT-6 Astra and two reward models. | "within three points of GPT-6 wherever a verdict can be read off the text, at 0.36% of its fee"; behind "where the verdict must be derived, as in math, code, and logic"; 12 points behind on JudgeBench. A cascade with a threshold frozen in advance escalates 31% of 1,610 pairs and is 0.9 points more accurate than GPT-6 at 41% of its fee. | Data public; reproducibility package described, location not found | B public-data 21:00 UTC · v3 replaced earlier results |
| B09 arXiv 2609.26758 v2, "Type-Safe Is Not Error-Free" Academic (affiliation not recorded). | Robustness: swap which option name (0/1, no/yes, A/B, random strings) is bound to which rubric. 1,200 yes/no questions from LocalLLaMA/typed-decisions.Models: hosted Jev, Laya, Open-Jev; no LLM or classifier. | No/yes names flip 32.5% of Jev's decisions, 30.4 points more than 0/1; AUC falls from .8146 to .5806. Type-error rate 0% throughout. Random-string names return to the neutral regime. Laya: 76.9% flipped with yes/no names; see Jev vs Laya vs Kev. | Data public; code not linked | B public-datano-baseline 21:00 UTC |
| B10 MindStudio, "Jev vs BERT and Zero-Shot NLI" (20 Sep) An AI agent platform's marketing blog. A secondary summary of an unnamed "recent independent benchmark". | Banking77, Yelp stars, an emotion set and a phishing set; n not given. Baselines: a 22M-parameter encoder with a trained head; zero-shot NLI. | Banking77: trained encoder 93.2% (8 ms on CPU), Jev 80.1% zero-shot, NLI 48.8% and 66.7%. Yelp: Jev 67.2%, trained classifier 51.9%. | None; primary study not identified | D secondary 20:53 UTC |
| B11 Opper, "Jev vs. Kev" (25 Sep; results 24 Sep) A gateway that serves Jev and hosts Kev. | Fresh items published after 20 Sep, ground truth from the source: arXiv category (n = 160), Stack Exchange site (120), GitHub bug or feature (82). Baseline: Kev 4B, an open Jev-like model (not an LLM or classifier). | Accuracy Jev 96.9 / 97.5 / 95.1% vs Kev 95.0 / 98.3 / 93.9%; calibration error Jev lower on all three. "differences under about 5 points are within noise." In context with Laya and Kev's documented limits: Jev vs Laya vs Kev. | Code and item manifest (Apache-2.0); results | B COIsmall-nno-baseline 20:53 UTC |
| B12 Bryo AI email benchmark (17 Sep) A company evaluating Jev for its own product; no stake in the result stated. | 1,565 German and English business emails, 10 categories. Private data. Baselines: Gemini 3.5 and 3.8. | The CTO's post: "Jev lost to Gemini on our email classification benchmark. I'm still interested in putting it into production." Figures from Prefactor (secondary): Jev 96.4% vs 97.5% and 98.5%, "10 to 20 times cheaper". | None | C single-task 20:53 UTC · only non-English row with an LLM baseline |
| B13 arXiv 2609.24574 v2, decision models for text annotation New York University Abu Dhabi, academic. | 18 computational-social-science classification tasks, 7,977 items; via OpenRouter. Baselines: 19 frontier and open-weight LLMs, same zero-shot protocol; the best LLM per task was chosen after the results. | Jev "trails the per-task best LLM on 14 of 15 evaluation tasks, with a median deficit of 11.6 macro-F1 points, at a median 44 times lower measured cost". Confidence better calibrated than the stated confidence of 16 of 19 LLMs; high confidence at near-chance accuracy on one empathy task. Routing low-confidence items to an LLM "matches or exceeds the LLM alone at a quarter to half of its cost". | Data public; code repository (MIT) exists, not opened | A provisionalpublic-data 21:01–21:02 UTC |
| B14 Elastic Search Labs, "Using Jev as a search reranker" (30 Sep) A search vendor; not a Jev reseller. | E-commerce reranking on public ESCI (US English): questions frozen on 50 queries, then 250 test queries (4,754 query–product pairs). Baselines: BM25, hybrid search, random order, a Jina reranker (row not extracted). | nDCG@10 for four Jev strategies 0.9533–0.9565 vs hybrid 0.9351 (+0.018 to +0.021, confidence intervals above zero); BM25 0.9234; random 0.8595 (high because most judged results are exact matches). | Data public (ESCI); notebook at elastic/elasticsearch-labs commit 3e3d688, opened 7 Oct: it runs a 10-query sample only (data/esci-sample-10-queries.json); the 250-query test sample, raw judgments and bootstrap code were not found | B public-datasingle-task 20:57 UTC (30 Sep) Regraded 7 Oct 2026: A (provisional) → B. The published notebook runs a 10-query sample, not the reported 250-query test, so the reported evaluation cannot be rebuilt from published materials. Prompts and code are published. Read 05:21 UTC. |
| B28 Red Hat Developer, "Benchmarking AI decision models against traditional guardrails" (2 Oct) Red Hat. It publishes two of the classifier baselines ( RedHatAI/deberta-v3-base-prompt-injection-v2, RedHatAI/granite-guardian-hap-125m) and one LLM baseline build, and runs the evaluation through its EvalHub and NeMo Guardrails work, so it has a stake. No Jev stake stated. | Guardrails: prompt injection, and content safety, toxicity and profanity, from EvalHub's NeMo Guardrails benchmark library, "class-balanced between risky and safe prompts". n not stated. "Jev-1.13.0, via TypeSafe's API", through an experimental NeMo Guardrails branch; a prompt was blocked "If any question returned a noul greater than 0.5".Baselines: trained classifiers deberta-v3-base-prompt-injection-v2 and granite-guardian-hap-125m; LLM judges Shieldstral-1.0-3B, Nemotron-3.5-Content-Safety (default and custom), Qwen3.6-35B-A3B-FP8, DiffusionGemma-26B; zero-shot or decision models bart-large-mnli and Laya. Classifier baselines in context: Jev vs an LLM vs a classifier; content-safety use: content moderation. | Prompt injection (accuracy): Qwen3.6-35B 89.31%, deberta-v3-base classifier 89.01%, DiffusionGemma 87.72%, Jev 86.35% (4th of 9). Content safety: Jev 86.20% (1st), DiffusionGemma 85.53%, Qwen3.6-35B 85.47%, granite-guardian-hap-125m classifier 80.27%. Median latency: Jev 348.1 ms vs deberta 54.1 ms; Jev 360.4 ms vs granite-guardian 33.2 ms. The article concludes that "decision models like Jev do not reliably outperform LLM-as-a-judge, pre-trained predictive models, or open source decision models in speed or accuracy." | Data: named public benchmark sets (n and revision not stated); code linked (NeMo Guardrails branch feat/jev-rail, eval-hub/eval-hub-contrib), not opened; prompts in the appendix; raw outputs not found | B COIprovisionalpublic-data n not stated · 7 Oct, 05:18 UTC; Tables 1–2 rechecked on a rendered page 05:26 UTC: match |
| B29 Synthpop, "Does a decision model (like Jev) know when it is guessing?" (Post 1) (1 Oct) Synthpop, a company with a commercial product; no Jev stake stated. Its repository says "This project is not affiliated with or endorsed by TypeSafe." | Calibration and self-knowledge on 15 public datasets and six generated task families. Main test: Banking77 and CLINC150 intents behind opaque codes, under nine knowledge settings; also TriviaQA, MMLU-Redux and MMLU-CF as four-option multiple choice. "About 575,000 API calls"; requests used jev-latest, "every response reported jev-1.13.0".Comparators: GLiNER2.5-Decide (an open decision model) and five fine-tuned copies of a 22M-parameter encoder (a trained-classifier ensemble). This is a confidence comparison, not a head-to-head on a fixed label set. No LLM. Using calibration in practice: confidence thresholds. | TriviaQA: "Keep only the answers Jev was at least 90% sure of, and you keep 82% of the questions, with 99.75% of those answers right". Smooth ECE "0.016 on MMLU-Redux, 0.031 on TriviaQA". Beyond the observed knowledge boundary, accuracy was "51% ... yet its confidence rose to 82%". Banking77 with one example per category: "83% confident but only 68% accurate". | Data public (pinned revisions); code yes (Apache-2.0, commit 1867760; metrics/calibration.py opened); prompts rebuilt by the code; raw outputs not published | A public-data 7 Oct, 05:18–05:20 UTC |
| B30 Maximum Effort, Minimum Reward (maximumeffort.substack.com), "How Accurately Calibrated is Jev?" (2 Oct) An individual's blog; no stake stated. | Calibration on physics distributions: 10 distributions × 5 prompt templates × 20 variations, "1,000 settings"; Jev chooses among about 49–50 binned ranges. Also distribution choice (100 settings per family) and a ladder of arithmetic tests. Prompts written by GPT-6 Astra and Claude Opus 5.5. Jev version not stated. Baseline: none; the reference is the analytic distribution, with a flat guess as the floor. Calibration in practice: confidence thresholds. | Total variation from the true distribution: "If Jev were to completely punt on the answer ... it would score a mean TV of 0.546. But Jev scores 0.518. For Uniform it is 0.77 vs. 0.39 for a flat guess, and for Poisson 0.65 vs. 0.64." Gamma problems were labelled "exponential" in "40/100 settings". | None located: no code, data or raw outputs; example prompts as images only | C no-baselinesingle-task 7 Oct, 05:18 UTC |
| B31 DoubtBench (Show HN, 4 Oct); the author's write-up "Jev isn't a better judge than Claude" (6 Oct) has the same numbers The benchmark's author (GitHub deeplearningguy), an individual; no Jev stake stated. The benchmark invites leaderboard submissions. | NVIDIA HelpSteer2 (CC BY 4.0, pinned revision) turned into Score, Choice and Noul questions: attribute scores, acceptability, pairwise preference (asked both ways) and preference strength. Labels are human annotator votes. Validation split, 7,455 questions; Jev answered 7,455 with 0 errors. Baselines: Claude Haiku 4.5, Sonnet 5.5 and Opus 5.5 with "verbalized" probabilities in JSON; Sonnet and Opus at low effort, and their refusals (30 and 36) count as wrong. A human-replay ceiling and a uniform floor. Judge use: Jev as a judge; calibration: confidence thresholds. | DoubtBench score (mean of accuracy, 1 − ECE and 1 − JS divergence) / accuracy: Jev 1.13 67.8 / 48.2%; Claude Haiku 4.5 69.2 / 45.8%; Opus 5.5 67.8 / 44.0%; Sonnet 5.5 67.4 / 42.5%; human ceiling 90.4 / 82.6%. On accuracy: "paired McNemar test, p < 0.001 against each". The author's caveat: "One dataset, one task". | Data yes (data/doubtbench.jsonl, manifest hashes); code yes (MIT, commit d5345ee; adapters/jev.py, metrics/core.py opened); prompts yes; raw outputs yes | A public-data 7 Oct, 05:19 UTC; leaderboard rechecked 05:26 UTC: match |
| B32 system-one-security (GitHub repository, created 4 Oct; runs 3 Oct) An individual; no stake stated. Coordinated disclosure with a vendor is pending: "some experiments, claims and results are withheld". | Security experiments on a synthetic citation-checker knowledge base: truncation, instruction injection, fact poisoning, prepend eviction and a tripwire question. 12 to 24 claims per experiment; experiment 02 ran 3 times on Jev, the rest once. jev-1.13.0, hosted TypeSafe API.Comparators: Cloudflare Clef and Clef-flash (decision models); no LLM or classifier baseline. Security limits in context: limitations. | "The instruction-style injections tested never made a wrong claim pass on the choice answer ... 0 of 12 on every injection". But "a 'labels are swapped' note at the end of the state pushed nearly every correct claim below 0.5 on the yes/no answer". Truncation: "Jev 24/24". The repository's 2,048-token truncation finding concerns Clef, not Jev. | Data yes (synthetic, kb/); code yes (MIT, commit ecdda36; experiments/02-instruction-injection.ts opened); prompts yes; raw outputs yes, per run | A small-nno-baselinesealedsingle-task 7 Oct, 05:20 UTC |
| B33 Vals AI, "An Independent Evaluation of TypeSafe's Jev" (6 Oct) Vals AI, an AI evaluation company that publishes benchmarks and reports; no Jev stake stated. | (1) Claim verification: 400 items built from the latest 10-Ks of 48 large US companies, 5 categories × 80. (2) A preregistered 12-subtask slice of LegalBench: 396 questions (33 per subtask). Jev 1.13.0 through the provider API. Baselines: 11 LLMs, including GPT-6 Astra, GPT-6.1 Sol, GPT-6 Luna, Claude Opus 5.5, Sonnet 5.5, Haiku 4.5, Gemini 4 Argon and DeepSeek V4.1 Flash, each asked for a label and a probability. Judge use: Jev as a judge; error budgets: confidence thresholds. | "On claim verification, all twelve systems score 0.953 to 0.995. Jev ... scores 0.975 at $0.02 per 1,000 cases and ties with GPT-6 Astra, GPT-6.1 Sol, and GPT-6 Luna." "On LegalBench, Jev finishes last of twelve on raw accuracy and on class-balanced accuracy." On the 1% error-budget test (held-out split of 265 items): "Jev was one of them, automating 95% of cases at 1.6% error". ECE "0.011 on claim verification, the lowest of the twelve, but 0.108 on LegalBench, the highest". | Claim-verification data not published; LegalBench public, but the 396-question sample not identified; no code; prompts and raw outputs not published | C public-data 7 Oct, 05:25 UTC |
| B34 Oodle AI blog, "Is Jev the right model for your agents?" (22 Sep; on Hacker News 6 Oct) Oodle AI, an observability and AI-evaluation platform vendor; the post promotes its own playground. No Jev stake stated; the run was self-funded. | Two yes/no agent decisions, 200 items each (100 yes, 100 no): support escalation (public Hugging Face set with human sentiment labels) and model routing (labels from GPT-4's grades of a Mixtral answer). jev-latest through OpenRouter (resolved version not stated); P(yes) above 0.5 counts as yes.Baselines: GPT-5.4-mini, Claude Haiku 4.5, Gemini 3.5 Flash, same prompt, no tuning. Scored by three LLM judges, averaged. Routing in practice: model routing. | Escalation: GPT-5.4-mini 84.5%, Jev 69.5%, Claude Haiku 4.5 66.0%, Gemini 3.5 Flash 66.0%; Jev p50/p95 163/264 ms. The author says a 0.3 threshold "would score about 82%"; that threshold was chosen after the fact from the first 100 items. Routing: Gemini 3.5 Flash 65.5%, GPT-5.4-mini 62.5%, Jev 59.5%, Haiku 4.5 56.5%. The author's caveat: "Treat this one as a method demonstration, not a verdict." | Source datasets public; the 200-item samples not identified; task and judge prompts published; no code; raw outputs not located (promised on the vendor's playground, which needs an account) | C small-nLLM-labelspublic-data 7 Oct, 05:28 UTC |
| B35 ivnle.github.io, "Your Language Model Is Not-So-Secretly a System 1 Model" (4 Oct; on Hacker News 6 Oct) Two academic researchers on a personal blog; no commercial or Jev stake stated. The post argues a thesis ("matching Jev's overall performance ... does not require a new kind of model"); "A preprint with full details will follow". | 25 datasets in three parts, 19,505 decisions: classification (13,053: Banking77 plus 20 BTZSC sets), agent monitoring (5,800: ABCD, AgentRewardBench) and triage (652: two TypeSafe published examples; the authors "replaced 25 of the 202 published labels" in the guardrail set). Separately, the 231 public JevBench questions. jev-1.13.0 through TypeSafe's Python SDK.Baselines: 23 ranked models, mostly open-weight LLMs read from next-token probabilities in one forward pass, no training, temperature-scaled on a 30% split; closed comparators GPT-6 Luna (read from the letter it writes) and Gemini 2.5 Flash-Lite. No trained task-specific classifier. In context: Jev vs an LLM vs a classifier. | Overall macro-F1, the three parts weighted equally: "Jev scores 0.787 [0.776, 0.797]"; gemma-4-31B-it 0.794 (difference +0.006 [−0.001, +0.013]); GPT-6 Luna 0.782, "at about twice Jev's cost"; Gemini 2.5 Flash-Lite 0.735. "Seven of the 23 open models we ranked match Jev" (pre-set 0.03 margin). Agent monitoring: "six of the seven beat Jev, by 0.025 to 0.052"; triage: "Jev is ahead: six of the seven fall below it". JevBench: "Jev answers 85.7% [81.8%, 89.6%] correctly. gemma-4-31B-it beats it with 91.8%". The authors say equal weighting of the parts "favours Jev". | Data public, plus TypeSafe's published examples (the 25 relabelled items not located); prompt template shown; no code repository found (latency points only); raw outputs not located | B public-data LLM-labels in part (excluded from its calibration figures) · 7 Oct, 05:28 UTC; rechecked 05:32 UTC: match |
| B36 SREGym blog, "Jev-Driven SRE Diagnosis: What Worked and What Failed" (6 Oct; on Hacker News 7 Oct) The SREGym team, which maintains the SREGym SRE-agent benchmark and built the pipeline it tests. No Jev stake stated. | Kubernetes incident root-cause diagnosis: code collects cluster evidence, Jev picks the component, origin or victim role, cause category and key evidence item, and code assembles the diagnosis. "the 21 fault scenarios in the September 4 SREGym-Lite cohort" × 5 runs = 105 diagnoses. jev-1.13.0; 252 Jev calls.Baseline: none. Scored by gpt-6-astra at high reasoning effort against SREGym's nine-question rubric, pass at 0.70. | "Judged diagnosis passes 80/105 (76.2%)". All five attempts passed on 16 of 21 faults and failed on 5 of 21: "For every fault, either all five attempts passed or all five failed." Median diagnosis time 14.6 s; estimated Jev cost "$0.15". | Code yes at the linked commit 1a668ee (it has diverged from main); pipeline and judge oracle opened; judge.py, the rubric file and the 4 Sep cohort list not verified; per-fault scores on the page; raw Jev and judge outputs not located | B LLM-labelssmall-nno-baselinesingle-task Graded B by the research runner (the grading helper proposed A): rerun materials not all located, and the operator built the benchmark and pipeline. May move to A once those files are read. 7 Oct, 05:28–05:29 UTC |
Using Jev as a judge: rows B02, B06 and B08 are summarised, unchanged, on Jev as a judge, next to what four judge tools record when a Jev call fails (4 Oct 2026). Judge-style studies added on 7 Oct: B31 (grade A) and B33 (grade C); see Jev as a judge.
Tier 2: shorter reads
Reported About a minute per item on 30 Sep 2026 (around 20:57 UTC); "claim only" unless materials were seen. B26 is carried from earlier research.
| Study and operator | Task, n and baseline | Result as reported | Grade and flags |
|---|---|---|---|
B15 punk2898 benchmark in awesome-jev-verifiedAwesome-list maintainer; no stake stated. | 2,390 labelled questions (true/false, multiple choice, consistency, injection). Baselines: GPT-4.1-mini, GPT-5.6 Sol. | True/false 92.5% vs 87.8% (4.1-mini) and 92.1% (Sol); multiple choice 77.3% vs 74.3% and 82.5%; yes/no calibration error 0.048 (Sol 0.043); multiple choice "overconfident, same as Sol"; "445× cheaper did not reproduce" (88× vs Sol). | A provisional |
| B16 Alex Molas, "Jev can't be calibrated" (23 Sep) Individual. | Argument with one small example; no baseline. | "Jev says a fair coin lands heads with probability 0.92"; yes/no questions better calibrated than Choice in a linked experiment. | D |
| B17 AnthusAI/Jev-Calibration (19 Sep) An AI firm; no Jev stake stated. | Held-out calibration of a yes/no question on a sentiment-style set; recalibration methods compared. | Calibration error 0.117 raw, 0.052 after Platt scaling, 0.008 after isotonic; accuracy and AUC "barely move (0.872 → 0.874)". | B no-baseline |
| B18 Rene-1 31B model card (27 Sep) Alternative-model author. | "Decision Index 0.2, 37 of 40 benchmarks, evaluated locally"; compared with Jev. | Balanced skill 64.18, marked verified: false; "+9 over Jev" per the pool's alternatives notes. | D COI |
| B19 Ollaya README Local-runtime author. | LocalLLaMA/typed-decisions, n not stated; compared with Jev. | winnow:e4b "0.722 accuracy on typed decisions (TypeSafe's Jev: 0.738)". | D COI |
| B20 Jeff (firelex) README (created 28 Sep) Alternative-model author. | Several benchmarks; compared with Jev. | "On benchmarks they approach, and sometimes beat, Jev" (figures not extracted). | D COI |
| B21 "A DIY Jev" (kgluszczyk, DEV.to) (28 Sep) An engineer building an in-house alternative. | 6,762 public examples, 13 datasets. Baselines: prompted JSON and a logprob adapter on gpt-6-luna. | Average accuracy: JSON 70.1%, DIY logprob 73.8%, Jev 76.5%; calibration error 0.175 / 0.211 raw (≤0.10 calibrated) / 0.110; DIY "nearly matched Jev on most category tasks". | C public-data |
| B22 "Jev After Eight Days of Independent Tests" (DEV.to) (24 Sep) Individual; a secondary review with its own grades. | Synthesis of other studies, not a study. | "Jev sits level with mid-price LLMs and 6.5 to 11.5 points behind the frontier in the cleanest comparison". Led this site to B13. | Not graded |
| B23 Casco CVSS benchmark (28 Sep) A security product company. | 2,449 security findings scored with CVSS 3.1; 8 models including GPT-6 Astra and Claude. | Mean absolute error vs Casco's scores: Astra 1.92 (first), Jev 3.40 (seventh of eight), Haiku 3.94. Jev scored 537 of 655 informational findings above zero. | C |
| B24 jev-scifact-eval (29 Sep, "pre-registered") Individual. | SciFact claim verification. Baselines: keyword rule; VeriSci (2020); no LLM. | Label+Rationale F1 0.625 [0.567, 0.684] vs 0.303 and 0.485. | A provisionalpublic-datano-baseline |
B26 Vercel fx command safety (CEO post, 16 Sep)Vercel, a gateway selling Jev. Production claim on Products. | Not published. Baseline: GPT Luna. | "up to 18x faster (p95) and more accurate". | D COI |
Leads, not graded (treated as D until read): calfram-bench (B25, calibration audit, not read); Vercel engineer's Banking77 sample, n = 385, 80.3% as quoted by Jevals (B27, not opened); LanceDB reranker comparison (now in Reranking studies); Halv's SWE-rebench cost claim; the CLM-8B "9x faster than Jev" headline; PostHog's Jeeves; the YouTube video "Jev's Founder Explains: Why You Can't Trust AI Benchmarks"; creator tests on the videos page (M37, Nate Herk, Giuseppe Funicello). Unverified
Found 7 Oct 2026, not graded, with reasons: jev-safety-eval (Aegis 1.0 content safety, baselines planned) has no results yet: "Results will follow"; to recheck next run for content moderation. jev-zh-tw-eval is not a Jev study: it tests Plumb-4B locally through a "Jev-compatible" interface, so its numbers are Plumb-4B's. The DAIR.AI piece "Jev-as-a-Judge for Agent Evaluations" is a tutorial, not a study: one refund example, no measurement on the open page. Checked 7 Oct 2026, 05:19–05:28 UTC.
Where the studies agree and conflict, by task type
No averaging: each line names the rows it rests on. Reported
- Judging and evaluation
Agree on directionJev can match a frontier LLM judge on verdicts that can be read off the text once a threshold is chosen on labelled data (B02: 87% vs Opus 5's 87% after tuning, 76% vs 83% at 0.5). It is behind where the verdict must be derived: math, code, logic, JudgeBench −12 points (B08). B15's true/false row (92.5% vs Sol 92.1%) is consistent. Added 7 Oct: B31 (grade A; Claude models with verbalized probabilities, two at low effort) reports Jev's accuracy at 48.2% vs 42.5–45.8% and a composite of 67.8 vs 67.4–69.2; one dataset, one task. B33 (grade C) reports a tie on claim verification and Jev last of twelve on LegalBench. The size of the gap depends on the task and the threshold.
- Topic and intent classification
Depends on the baselineAhead of cheap LLM setups: B03 (87.5% vs 71–83%), B21 (76.5% vs 70.1–73.8%), B15 multiple choice vs GPT-4.1-mini. Behind stronger or best-per-task LLMs: B06 (Gemini 3.8 Flash on Banking77), B13 (14 of 15 tasks, median 11.6 macro-F1), B15 (77.3% vs Sol 82.5%). Against trained classifiers on a fixed label set, mixed: one grade B study, B28 (behind a prompt-injection classifier, 86.35% vs 89.01%; ahead of a small content-safety classifier, 86.20% vs 80.27%; Red Hat publishes both classifiers), and grade D evidence, B10 (93.2% vs 80.1%) and B04. Open-weight LLMs scored through token probabilities: about level overall in one grade B study (B35, 0.787 for Jev; open models ahead on agent monitoring, Jev ahead on triage). More on choosing between them: Jev vs an LLM vs a classifier. Against an open Jev-like model on fresh data, no difference beyond noise (B11).
- Email and ticket routing
Point different waysB12 (Jev 1.1–2.1 points below Gemini on 1,565 bilingual emails; secondary figures) and B05 (Jev above Haiku 4.5 on 80 authored prompts; the operator has a stake) use different baselines. Neither publishes materials; both are grade C. Added 7 Oct: B34 (grade C, 200 items per task) reports Jev behind GPT-5.4-mini on support escalation (69.5% vs 84.5%) and ahead of Haiku 4.5 and Gemini 3.5 Flash (66.0%).
- Reranking
AgreeB14 (+0.018 to +0.021 nDCG@10 over hybrid search; grade B since 7 Oct) and B04 ("comparable to internal system") agree that Jev can improve on or match a first-stage ranking. No extracted result compares it with a strong trained cross-encoder. Three more reranking or relevance studies: Reranking studies.
- Multi-question workflows
Vendor only - Numeric scoring
Agree: weak - Calibration
Usable, with exceptionsB06 (5.0 points on PubMedQA), B15 (yes/no level with Sol), B13 (better than 16 of 19 LLMs' stated confidence, worse than three frontier models) and B11 find the probabilities usable as a gate. B15 (multiple choice), B10 and B16 report overconfidence; B17 shows recalibration on your own data cuts the error sharply without changing accuracy. Added 7 Oct: B29 (grade A) reports 99.75% accuracy on the 82% of TriviaQA answers given at 90% confidence or more, but rising confidence beyond the knowledge boundary; B30 (grade C) reports a mean TV of 0.518 against 0.546 for a flat guess on physics distributions; B33 (grade C) reports low ECE on claim verification and high ECE on LegalBench. How to set a cutoff: confidence thresholds.
- Cascades (Jev first, then an LLM)
AgreeB08, B13 and B02's escalation section agree that sending low-confidence items to an LLM can match the LLM's accuracy at a fraction of its cost. See the hybrid pattern.
Reranking studies
Reported Checked 7 Oct 2026, 05:18–05:22 UTC. Graded on the same A–D scale for the reranking question only; these have no B number and are not in the 33 graded rows. Nothing was rerun.
| Study, operator and interest | Task, n and baseline | Result as reported | Materials | Grade and flags |
|---|---|---|---|---|
| iwhalen.com, "Pseudo-relevances with Jev" (5 Oct) An individual's blog; no stake stated. | Relevance judgments on the TREC DL 2023 judged subset (22,327 judged pairs), compared with human grades; system ranking of TREC runs. Baseline: gpt-6-luna. | Cohen's kappa against human grades "0.1907" (Jev) vs "0.2391" (GPT-6-luna); system ranking: "Jev and Luna are essentially tied" (tau values are in graphics, not extracted). Cost $0.6783 vs $1.3283. | Data on Hugging Face; code at 82d9a0f (main.py opened); metrics defined in code | A public-datasingle-task |
| Hugging Face community blog, "Introducing jev-reranker" (Sep; date not confirmed) The library's author, writing about the library. | NanoBEIR-en NanoHotpotQA via HAKARI-Bench at a pinned revision; 50 queries. Baseline: hybrid search only. | nDCG@10: hybrid 0.833; rerank 0.969; relevance_rerank (threshold 0.2) "0.975", keeping 7.62 documents per query (92.38% removed). | Code at d58594b (examples/eval.py, docs/eval.md opened); data public | B COIsmall-nno-baselinepublic-datasingle-task |
| LanceDB, "How Jev Compares to Other Rerankers" (date not confirmed) A vector-database vendor that ships a TypeSafeReranker. | 19 reranker configurations on GooAQ (20,000 queries, seed 0), BEIR NQ, HotpotQA, FiQA and SciDocs. Baselines: other rerankers and no reranker. | Example: GooAQ jev-relevance Hybrid@10 92.59, p50 169 ms. The README says the default prompt "drops below no reranker on HotpotQA". | Results and READMEs at 4616259 read; benchmark script not opened | B COIprovisionalpublic-data |
| B14 Elastic Search Labs (30 Sep) A search vendor. | ESCI reranking, 250 test queries reported. | See row B14. Does not count here: the published notebook covers 10 queries, not the reported 250-query test. | Notebook at 3e3d688 opened (10-query sample) | B regraded 7 Oct |
Result: 2 of 2 met beyond Parallel (B04) and the earlier MindStudio lead. Two reproducible reranking or relevance studies were read at pinned commits: iwhalen (grade A) and Hugging Face jev-reranker (grade B, COI). LanceDB is a third (grade B, provisional). Elastic B14 does not count.
Caveat. iwhalen measures relevance judgments (agreement with TREC grades and system ranking), not a reranking pipeline. jev-reranker is its author's own library. Two other items are not studies: the Spring blog's Modular RAG post is an implementation demo with no measurement, and refix.ai's "Jev for RAG" is an explainer.
A reranking-and-RAG page is planned for the next run, if confirmed.
What nobody has measured yet (as far as this run found)
- Non-English work with a baseline and published materials. Only B12 (German and English, grade C); one Italian repository lead. TypeSafe's docs say English is "where accuracy is currently best".
- Calibration of the Score question type.
- Long state near the 32k limit (TypeSafe documents that large irrelevant state hurts).
- A head-to-head against a fine-tuned classifier (BERT, SetFit, GLiClass) or zero-shot NLI with published materials that this site has read. B28 (grade B) links its code, but the code was not opened and n is not stated.
- Behaviour across model versions: only
jev-1.13.0exists so far. - Dates and arithmetic as separate tasks.
- Throughput and error rates under load or near the rate limits.
- Outcomes in production: no row measures the downstream cost of errors or drift.
- Images: Jev accepts text only.
How this site checked one study
The only Tested item on this page is offline arithmetic, not a Jev call. The site downloaded Jevals' published per-decision log for Jev on PubMedQA (1,500 records: 300 items × 5 runs; commit 21bb47b), recomputed the summary figures with its own script, and deleted the file afterwards.
| Figure | Published | Recomputed | Outcome |
|---|---|---|---|
| Accuracy | 91.3% | 91.33% | match |
| Median latency | 438 ms | 0.4377 s | match |
| 95th percentile latency | 653 ms | 0.6528 s | match |
| Mean cost per 1,000 | $0.029 | $0.0285 | match |
| Calibration error (10 bins) | 5.0 points | 5.10 points | near match binning may differ |
The Decision Score (69.0) was not recomputed. The check confirms the published summary matches the published log; it does not confirm the log itself.
What was not verified
- No study was rerun and no Jev call was made. The linked code repositories behind the three provisional A grades (B13, B15, B24) and JevBench's code were not opened.
- B08's reproducibility package was not located; if it is found, the row may move to grade A.
- B10's primary study is not identified (a related lead: a 26 Sep Show HN gist claiming "94% on Banking77").
- B14's Jina reranker row was not extracted; its 250-query test sample, raw judgments and bootstrap code were not found (7 Oct).
- Added 7 Oct 2026:
- B28 (Red Hat): n per benchmark set, and the code at a pinned commit, were not read. The article returned HTTP 403 to a plain fetch; its tables were confirmed on a browser render.
- B29: the arXiv paper (2610.01006) and Part 2 were not read.
- B30: the Jev version is not stated.
- B33: the claim-verification data and the LegalBench 396-question sample were not located.
- B34: which 200 items were sampled; the playground experiments (need an account); OpenRouter's resolved Jev version.
- B35: no code repository found; the 25 relabelled guardrail items and the preprint not located.
- B36: the 4 Sep SREGym-Lite cohort list,
judge.pyand the rubric file were not read; SREGym's earlier Jev post was not read. - Reranking: the LanceDB benchmark script was not opened; the Hugging Face and LanceDB post dates were not confirmed on the pages; iwhalen's Kendall tau values are in images only.
- A 240-example measurement in a Hacker News comment on the Red Hat thread (item 49981222) was not read.
- Who operates the Hugging Face account
LocalLLaMA, whosetyped-decisionssubsets carry the same four workflow names as TypeSafe's WorkflowEvals, is unknown. - The WorkflowEvals Hugging Face datasets and published runs were not opened.
Sources for the 7 Oct 2026 update (UTC)
40 source IDs (R10-S…) in 22 entries, with pinned commits and opened files
- R10-S11: search suggestions from Google (
client=firefox), Bing (osjson) and DuckDuckGo (ac), 05:18:59. - R10-S37: Terms, AUP, MCA, DPA, Privacy Policy, 05:19:15; 0 matches for the three terms.
- R10-S38: docs legal.md and its 22 Sep archive copy, 05:19–05:20.
- R10-S40, R10-S201, R10-S202: Red Hat Developer article (B28), 05:18:45; browser render 05:26:40; HN 49933476.
- R10-S203, R10-S204, R10-S205: Synthpop post (B29), 05:18:39; code 1867760,
src/beyond_answer_confidence/metrics/calibration.py, 05:20:05. - R10-S206, R10-S207: maximumeffort.substack.com post (B30), 05:18:39.
- R10-S208, R10-S209: DoubtBench at d5345ee (B31): README,
adapters/jev.py,metrics/core.py,results/jev-1.13.0/run.json, 05:19:22; rechecked 05:26:47. - R10-S225: DoubtBench write-up (B31), 05:25:16.
- R10-S210: jev-safety-eval at 4bb2819 (no results), 05:19:09.
- R10-S211: jev-zh-tw-eval at fe49a4a,
eval/run_fixed.py(not a Jev study), 05:20:25. - R10-S212: system-one-security at ecdda36 (B32): README,
docs/CLAIMS.md,experiments/02-instruction-injection.ts, 05:20:25. - R10-S213, R10-S214: Elastic post (B14), 05:18:40; notebook directory at 3e3d688, 05:21:25.
- R10-S215, R10-S216, R10-S223: iwhalen.com post, 05:18:39; code 82d9a0f,
main.py, 05:21:38. - R10-S217, R10-S218: Hugging Face blog post, 05:22:04; code d58594b,
examples/eval.py,docs/eval.md, 05:22:24. - R10-S219, R10-S220: LanceDB post, 05:22:03; results README at 4616259, 05:22:14.
- R10-S221, R10-S222: Spring blog and refix.ai (not studies), 05:22.
- R10-S224: Vals AI evaluation (B33), 05:25:16.
- R10-S401, R10-S402: Oodle AI post (B34), 05:28:06; HN 49982760.
- R10-S403, R10-S404: ivnle.github.io post (B35), 05:28:06; rechecked 05:32:03; HN 49980496.
- R10-S405, R10-S406: DAIR.AI tutorial (not a study), 05:28:07.
- R10-S407, R10-S408: SREGym post (B36), 05:28:07; HN 49986765.
- R10-S409, R10-S410, R10-S411: SREGym at 1a668ee (B36):
clients/jev_diag/config.py,driver.py,sregym/conductor/oracles/llm_as_a_judge/llm_as_a_judge_oracle.py;mainHEAD 28db972 bygit ls-remote, 05:28:45–05:29:20.