jev-1.13.0 · docs checked Setting Jev (TypeSafe AI) confidence thresholds: what the numbers mean and how to test them on your data
Jev returns probabilities for Choice and Score, a confidence number for both, and a single yes-probability for Noul. None of these is the probability that the answer is correct on your task. This page explains each field, shows how TypeSafe's own examples fit a formula, lists independent calibration tests, and gives a seven-step procedure for choosing a cutoff.
confidence summarises how peaked the probability distribution is. For a Choice with n options it matches (n·p_max − 1)/(n − 1) in TypeSafe's documented examples, which is 2·p_top − 1 when there are two options. It is not the probability that the answer is correct. Noul has no confidence field; its number is P(yes). Pick a cutoff per question on your own labelled data, and recalibrate if a reliability check shows the probabilities are off. This site recommends no threshold.
What each field means
| Type | Fields | Official meaning | What the pool found | Label |
|---|---|---|---|---|
| Choice | choice, probabilities, confidence | choice is the highest-probability option; probabilities sum to 1; confidence is "derived from" how the probabilities are spread (flat is low, one peak is high). | Fits (n·p_max − 1)/(n − 1) in 7 of 7 documented examples, within ±0.01. It adds no information beyond the probabilities. | Documented; fit Tested (offline arithmetic) |
| Score | score, legend, probabilities, confidence | score is a probability-weighted position that "can land between levels". The jaggedness page says score levels are "weak in numerical calibration". | The same formula fits the 3-level examples but not the documented 4-level and 5-level examples. The Score confidence formula is undocumented. | Documented; Tested (offline arithmetic) |
| Noul | noul | Probability that the answer is yes. "There is no separate confidence value for a Noul." | Near 0.5 means undecided, not "medium". A Noul and its negation need not add up to 1 (official example: 0.72 + 0.47 = 1.19). | Documented |
| All | Rounding | Documented examples show two decimals. | Probabilities are reported to be rounded to two decimals after the winner is picked, which can flip near-ties (SDK #15). Treat differences of 0.01 or less as ties. | Reported |
TypeSafe's documented Choice example: probabilities and confidence
department Choice (documentation values, not a pool measurement). With three options, (3 × 0.88 − 1) ÷ 2 = 0.82, within 0.01 of the returned 0.81. The two-option shortcut 2 × 0.88 − 1 = 0.76 does not fit. A confidence of 0.81 does not mean the answer is right 81% of the time.Formula check on all 12 documented examples
Tested (offline arithmetic, no Jev call) on TypeSafe's documented example responses, labelled jev-1.13.0. TypeSafe's docs call the three-option formula an approximation.
| Documented example | n | p_max | Returned confidence | (n·p−1)/(n−1) | 2p−1 |
|---|---|---|---|---|---|
API Choice department | 3 | 0.88 | 0.81 | 0.82 | 0.76 |
| Quickstart Choice | 3 | 0.85 | 0.78 | 0.775 | 0.70 |
Choice department (primitives page) | 3 | 0.61 | 0.42 | 0.415 | 0.22 |
Choice shipping_issue | 5 | 0.74 | 0.67 | 0.675 | 0.48 |
Choice requested_resolution | 4 | 0.40 | 0.20 | 0.20 | −0.20 |
Choice tone | 3 | 0.84 | 0.76 | 0.76 | 0.68 |
Choice return_reason | 5 | 1.00 | 1.00 | 1.00 | 1.00 |
API Score frustration | 3 | 0.95 | 0.92 | 0.925 | 0.90 |
| Score, 3 levels (0, 0.57, 0.43) | 3 | 0.57 | 0.35 | 0.355 | 0.14 |
| Score, 3 levels (0, 0.74, 0.26) | 3 | 0.74 | 0.61 | 0.61 | 0.48 |
| Score, 4 levels (0, 0, 0.48, 0.52) | 4 | 0.52 | 0.52 | 0.36 | 0.04 |
| Score, 5 levels (0, 0.14, 0.86, 0, 0) | 5 | 0.86 | 0.89 | 0.825 | 0.72 |
Green: the formula fits within ±0.01. Red: it does not. A MarkTechPost guide (23 Sep) presents the formula as "TypeSafe's published statistic"; the official text only uses it to approximate three-option confidence in a demo.
What TypeSafe says (vendor claims)
- Jev "is trained with RLCD … to return calibrated decisions". No calibration metric (ECE, Brier score, reliability plot) is published on the pages read.
- "The correct threshold values depend on your domain and the performance of the model for your use case." The 0.5 and 0.9 values on the Confidence page appear in illustrative code, not as recommendations.
- For Noul: use 0.5 when yes and no are equally actionable, raise the cutoff when a false yes is costly, lower it when a missed yes is costly, and send middle values to a person.
- "Low confidence: Do not act. Route to a human, request clarification, or fall back to a different system."
- "Don't carry a threshold tuned on a Noul over to a Choice." Pin the versioned model ID once thresholds are tuned.
Sources: Confidence page, Models page, Noul page, jaggedness page, checked 28 Sep 2026. Documented as TypeSafe's own claims.
Seven steps to choose a cutoff
What real implementations do with answers below their cutoff (held, acts anyway, fallback and others, counted across 64 implementations) is on when Jev fails. How to escalate unsure verdicts when Jev is used as an eval judge is on Jev as a judge.
- Freeze the question. Fix the question text, criteria,
model="jev-1.13.0"(never an alias) and how you serialise the request. - Label real cases. Collect a few hundred per question, including out-of-scope, ambiguous and high-impact cases. Split them into development and held-out test sets that share no source conversation.
- Record the raw output. For each development case, store the probabilities and, for Choice and Score, the confidence. You may compute your own statistic from the probabilities.
- Check calibration. Plot a reliability diagram. If it is off, fit a recalibrator (isotonic or Platt) on development data only.
- Sweep cutoffs on development data. For each cutoff, count the cost of wrongly accepted answers and the review load. Choose against a written maximum error rate and your review capacity. Use a separate cutoff for each question and each question type.
- Evaluate once on the held-out set. Report the counts and the review load.
- Re-run after every change to the model version, criteria or question wording, and check for drift in production with logged outcomes.
This is the maintained version of the pool's 25 September workload-evaluation protocol, which remains as a dated report.
Thresholds that projects chose (examples, not recommendations)
| Project | Setting | How it is used |
|---|---|---|
| QuantDinger | JEV_MIN_CONFIDENCE=0.55 | Risk and execution checks on a trade entry gate |
| Jevmail | Top-two probability gap below 0.15 | A low-confidence flag, not a review gate |
| WordPress Jev Comment Triage | Spam 0.75, scam 0.65, hold 0.45 | Default site thresholds |
From the 28 Sep source-path audit. They show what three projects picked; none was validated on your data. More examples read in source: five ticket-triage projects use 0.5 to 0.9 (support-ticket triage); four of six model routers gate on confidence at 0.3 to 0.75 and two only log it (model routing); moderation projects delete from 0.70 to 0.81, and only one sends its middle band to a person (content moderation). All read 1–2 Oct 2026.
Independent calibration tests and tools
| Item | Date | Method | Reported result | Licence |
|---|---|---|---|---|
| AnthusAI/Jev-Calibration | 19 Sep | 8,801 constructed sentiment examples on jev-1.13.0; 60/40 split; Platt vs isotonic | Expected calibration error 0.117 raw, 0.052 after Platt scaling, 0.008 after isotonic recalibration | None stated |
| Adilmp/does-jev-confidence-mean-anything | 19 Sep | 8,000 judgments against human labels (civil_comments), 4 wordings, 2 base rates | Error 0.157–0.209 raw, 0.006–0.023 after recalibration; ranking quality (AUC) about 0.91, unchanged | None stated |
| Alex Molas, "Jev can't be calibrated" | 23 Sep | Argument with one small example: "Jev says a fair coin lands heads with probability 0.92"; no dataset. HN discussion 65 points, 61 comments (thread not read) Corrected 30 Sep 2026: this row said "no experiment". Grade D on the benchmarks ledger (B16) either way. | Calibration depends on your data distribution; treat outputs as scores and recalibrate on a few hundred labels | — |
| Prefactor article | 25 Sep | Opinion from an agent-observability vendor; no data | Notes that TypeSafe has not published its calibration method | — |
| abhixhek/jevcal | 18 Sep | Command-line tool: threshold per question for a target accuracy; drift check; demo uses a simulator | Publishes no Jev numbers | MIT |
| Adilmp/jevcal (a different project with the same name) | 19 Sep | Single-file recalibration | — | MIT |
| jevkit-calibrate 0.1.0 | 21 Sep | Reliability diagrams, expected calibration error, Brier score, thresholds from labelled outcomes | — | MIT |
Reported Each result comes from one dataset and one model version, measured by its author. The pool did not rerun them. Repositories without a licence cannot be reused freely.
Calibration results in larger studies (added 30 Sep 2026). The graded benchmarks ledger has more calibration rows, each with operator, task, n and grade. B06 Jevals reports an expected calibration error of 5.0 points on PubMedQA (the pool recomputed 5.10 from the published log). B13 (NYU Abu Dhabi) finds Jev's confidence better calibrated than the stated confidence of 16 of 19 LLMs, but reports one task with high confidence at near-chance accuracy. B15 (punk2898) finds yes/no calibration level with GPT-5.6 Sol and multiple-choice overconfidence. B17 (AnthusAI) is the recalibration study in the table above. All Reported. The procedure on this page stays the same: choose the cutoff on your own labelled data.
Added 7 Oct 2026: new calibration rows on the ledger, B29 Synthpop (grade A), B30 (grade C) and B31 DoubtBench (grade A), each Reported by its operator.
What was not verified
- How Score confidence is computed for 4 or more levels.
- Whether TypeSafe's calibration claim holds on any given task; the independent tests each cover one dataset.
- abhixhek/jevcal's statement that TypeSafe's customer agreement restricts publishing benchmark results was not re-read. The agreement itself was read on 30 Sep 2026: the current MCA has no such clause, but the version dated 27 Aug 2026 had one (section 2.3(f)) until the version dated 19 Sep 2026. See the terms box; this is a reading of the text, not legal advice.
- The pool made no Jev call; the formula check uses documentation examples only.