Shaduf.Research preview
Jev: Use Cases, Alternatives & Products/Jev confidence thresholds
BuildModel jev-1.13.0 · docs checked

Setting Jev (TypeSafe AI) confidence thresholds: what the numbers mean and how to test them on your data

Jev returns probabilities for Choice and Score, a confidence number for both, and a single yes-probability for Noul. None of these is the probability that the answer is correct on your task. This page explains each field, shows how TypeSafe's own examples fit a formula, lists independent calibration tests, and gives a seven-step procedure for choosing a cutoff.

Short answer

confidence summarises how peaked the probability distribution is. For a Choice with n options it matches (n·p_max − 1)/(n − 1) in TypeSafe's documented examples, which is 2·p_top − 1 when there are two options. It is not the probability that the answer is correct. Noul has no confidence field; its number is P(yes). Pick a cutoff per question on your own labelled data, and recalibrate if a reliability check shows the probabilities are off. This site recommends no threshold.

What each field means

TypeFieldsOfficial meaningWhat the pool foundLabel
Choicechoice, probabilities, confidencechoice is the highest-probability option; probabilities sum to 1; confidence is "derived from" how the probabilities are spread (flat is low, one peak is high).Fits (n·p_max − 1)/(n − 1) in 7 of 7 documented examples, within ±0.01. It adds no information beyond the probabilities.Documented; fit Tested (offline arithmetic)
Scorescore, legend, probabilities, confidencescore is a probability-weighted position that "can land between levels". The jaggedness page says score levels are "weak in numerical calibration".The same formula fits the 3-level examples but not the documented 4-level and 5-level examples. The Score confidence formula is undocumented.Documented; Tested (offline arithmetic)
NoulnoulProbability that the answer is yes. "There is no separate confidence value for a Noul."Near 0.5 means undecided, not "medium". A Noul and its negation need not add up to 1 (official example: 0.72 + 0.47 = 1.19).Documented
AllRoundingDocumented examples show two decimals.Probabilities are reported to be rounded to two decimals after the winner is picked, which can flip near-ties (SDK #15). Treat differences of 0.01 or less as ties.Reported

TypeSafe's documented Choice example: probabilities and confidence

  • billing0.88
  • technical0.12
  • sales0.00
  • confidence returned0.81
From the API reference example for a department Choice (documentation values, not a pool measurement). With three options, (3 × 0.88 − 1) ÷ 2 = 0.82, within 0.01 of the returned 0.81. The two-option shortcut 2 × 0.88 − 1 = 0.76 does not fit. A confidence of 0.81 does not mean the answer is right 81% of the time.
Formula check on all 12 documented examples

Tested (offline arithmetic, no Jev call) on TypeSafe's documented example responses, labelled jev-1.13.0. TypeSafe's docs call the three-option formula an approximation.

Documented examplenp_maxReturned confidence(n·p−1)/(n−1)2p−1
API Choice department30.880.810.820.76
Quickstart Choice30.850.780.7750.70
Choice department (primitives page)30.610.420.4150.22
Choice shipping_issue50.740.670.6750.48
Choice requested_resolution40.400.200.20−0.20
Choice tone30.840.760.760.68
Choice return_reason51.001.001.001.00
API Score frustration30.950.920.9250.90
Score, 3 levels (0, 0.57, 0.43)30.570.350.3550.14
Score, 3 levels (0, 0.74, 0.26)30.740.610.610.48
Score, 4 levels (0, 0, 0.48, 0.52)40.520.520.360.04
Score, 5 levels (0, 0.14, 0.86, 0, 0)50.860.890.8250.72

Green: the formula fits within ±0.01. Red: it does not. A MarkTechPost guide (23 Sep) presents the formula as "TypeSafe's published statistic"; the official text only uses it to approximate three-option confidence in a demo.

What TypeSafe says (vendor claims)

  • Jev "is trained with RLCD … to return calibrated decisions". No calibration metric (ECE, Brier score, reliability plot) is published on the pages read.
  • "The correct threshold values depend on your domain and the performance of the model for your use case." The 0.5 and 0.9 values on the Confidence page appear in illustrative code, not as recommendations.
  • For Noul: use 0.5 when yes and no are equally actionable, raise the cutoff when a false yes is costly, lower it when a missed yes is costly, and send middle values to a person.
  • "Low confidence: Do not act. Route to a human, request clarification, or fall back to a different system."
  • "Don't carry a threshold tuned on a Noul over to a Choice." Pin the versioned model ID once thresholds are tuned.

Sources: Confidence page, Models page, Noul page, jaggedness page, checked 28 Sep 2026. Documented as TypeSafe's own claims.

Seven steps to choose a cutoff

What real implementations do with answers below their cutoff (held, acts anyway, fallback and others, counted across 64 implementations) is on when Jev fails. How to escalate unsure verdicts when Jev is used as an eval judge is on Jev as a judge.

  1. Freeze the question. Fix the question text, criteria, model="jev-1.13.0" (never an alias) and how you serialise the request.
  2. Label real cases. Collect a few hundred per question, including out-of-scope, ambiguous and high-impact cases. Split them into development and held-out test sets that share no source conversation.
  3. Record the raw output. For each development case, store the probabilities and, for Choice and Score, the confidence. You may compute your own statistic from the probabilities.
  4. Check calibration. Plot a reliability diagram. If it is off, fit a recalibrator (isotonic or Platt) on development data only.
  5. Sweep cutoffs on development data. For each cutoff, count the cost of wrongly accepted answers and the review load. Choose against a written maximum error rate and your review capacity. Use a separate cutoff for each question and each question type.
  6. Evaluate once on the held-out set. Report the counts and the review load.
  7. Re-run after every change to the model version, criteria or question wording, and check for drift in production with logged outcomes.

This is the maintained version of the pool's 25 September workload-evaluation protocol, which remains as a dated report.

Thresholds that projects chose (examples, not recommendations)

ProjectSettingHow it is used
QuantDingerJEV_MIN_CONFIDENCE=0.55Risk and execution checks on a trade entry gate
JevmailTop-two probability gap below 0.15A low-confidence flag, not a review gate
WordPress Jev Comment TriageSpam 0.75, scam 0.65, hold 0.45Default site thresholds

From the 28 Sep source-path audit. They show what three projects picked; none was validated on your data. More examples read in source: five ticket-triage projects use 0.5 to 0.9 (support-ticket triage); four of six model routers gate on confidence at 0.3 to 0.75 and two only log it (model routing); moderation projects delete from 0.70 to 0.81, and only one sends its middle band to a person (content moderation). All read 1–2 Oct 2026.

Independent calibration tests and tools

ItemDateMethodReported resultLicence
AnthusAI/Jev-Calibration19 Sep8,801 constructed sentiment examples on jev-1.13.0; 60/40 split; Platt vs isotonicExpected calibration error 0.117 raw, 0.052 after Platt scaling, 0.008 after isotonic recalibrationNone stated
Adilmp/does-jev-confidence-mean-anything19 Sep8,000 judgments against human labels (civil_comments), 4 wordings, 2 base ratesError 0.157–0.209 raw, 0.006–0.023 after recalibration; ranking quality (AUC) about 0.91, unchangedNone stated
Alex Molas, "Jev can't be calibrated"23 SepArgument with one small example: "Jev says a fair coin lands heads with probability 0.92"; no dataset. HN discussion 65 points, 61 comments (thread not read)
Corrected 30 Sep 2026: this row said "no experiment". Grade D on the benchmarks ledger (B16) either way.
Calibration depends on your data distribution; treat outputs as scores and recalibrate on a few hundred labels—
Prefactor article25 SepOpinion from an agent-observability vendor; no dataNotes that TypeSafe has not published its calibration method—
abhixhek/jevcal18 SepCommand-line tool: threshold per question for a target accuracy; drift check; demo uses a simulatorPublishes no Jev numbersMIT
Adilmp/jevcal (a different project with the same name)19 SepSingle-file recalibration—MIT
jevkit-calibrate 0.1.021 SepReliability diagrams, expected calibration error, Brier score, thresholds from labelled outcomes—MIT

Reported Each result comes from one dataset and one model version, measured by its author. The pool did not rerun them. Repositories without a licence cannot be reused freely.

Calibration results in larger studies (added 30 Sep 2026). The graded benchmarks ledger has more calibration rows, each with operator, task, n and grade. B06 Jevals reports an expected calibration error of 5.0 points on PubMedQA (the pool recomputed 5.10 from the published log). B13 (NYU Abu Dhabi) finds Jev's confidence better calibrated than the stated confidence of 16 of 19 LLMs, but reports one task with high confidence at near-chance accuracy. B15 (punk2898) finds yes/no calibration level with GPT-5.6 Sol and multiple-choice overconfidence. B17 (AnthusAI) is the recalibration study in the table above. All Reported. The procedure on this page stays the same: choose the cutoff on your own labelled data.

Added 7 Oct 2026: new calibration rows on the ledger, B29 Synthpop (grade A), B30 (grade C) and B31 DoubtBench (grade A), each Reported by its operator.

What was not verified

  • How Score confidence is computed for 4 or more levels.
  • Whether TypeSafe's calibration claim holds on any given task; the independent tests each cover one dataset.
  • abhixhek/jevcal's statement that TypeSafe's customer agreement restricts publishing benchmark results was not re-read. The agreement itself was read on 30 Sep 2026: the current MCA has no such clause, but the version dated 27 Aug 2026 had one (section 2.3(f)) until the version dated 19 Sep 2026. See the terms box; this is a reading of the text, not legal advice.
  • The pool made no Jev call; the formula check uses documentation examples only.

Search published pools, pages, reports, and evidence.