Shaduf.Research preview
CompareLedger rows and docs checked

Jev (TypeSafe AI) vs an LLM with JSON output vs a trained classifier: when to use which

For a classification, routing, gating or scoring task you can call Jev, ask a general LLM for structured output, train a classifier on your labels, or use a zero-shot NLI model. This page gives ten criteria for the choice, the published head-to-head results behind them, and a combined pattern several studies support.

Short answer

Use Jev when you have few labels or the labels change, and you want a probability you can threshold. Use an LLM when you need text, long inputs, arithmetic or the best accuracy on hard judgments. Use a trained classifier (or an open local model) when you have thousands of labels for a fixed set, need CPU or on-device inference, or cannot send data out. Use Jev first, then an LLM for low-confidence items when you want LLM accuracy at lower cost.

All benchmark evidence here is Reported by the operators named on the benchmarks ledger; this site ran no test. There are no prices on this page: compare costs on cost per decision and with the worksheet there.

How strong is the classifier comparison?

Until 7 Oct 2026 the comparison between Jev and trained classifiers or zero-shot NLI rested only on grade D rows: B04 (Parallel: "Internal 'wins'" on topic and freshness, no n) and B10 (MindStudio's secondary summary of an unnamed study). No graded study with published materials compares Jev with a fine-tuned classifier such as BERT, SetFit or GLiClass, or with zero-shot NLI (B28 below uses its publisher's own guardrail classifiers; its code was not read). Criterion 1's "thousands of labels" advice is therefore a reasoned expectation plus weak evidence, not a measured result.

Added 7 Oct 2026: the ledger now has one grade B study with trained-classifier baselines, B28 (Red Hat, which publishes both classifiers; code not read, n not stated). Its result is mixed: Jev was behind a prompt-injection classifier and ahead of a small content-safety classifier.

Ten selection criteria

Each row says which approach it favours, why, and the evidence with its label. Rows are ordered by how early the question usually comes up, not by weight. Documented rows cite TypeSafe's docs (read 30 Sep 2026, 20:47–20:48 UTC); Reported rows cite ledger studies.

CriterionFavoursWhyEvidenceStill unknown
1. Labelled data: none, some or thousandsNone or some: Jevor LLMThousands, fixed labels: classifierJev and LLMs need no training. A few hundred labels make a threshold reliable. With thousands of labels for a fixed set, a trained model learns the label boundary directly and is likely to win on accuracy, speed and cost.B02 thresholds tuned on half the labels; B17 recalibration Reported. B10 encoder 93.2% vs Jev 80.1% and B04, both grade D ReportedNo graded study with materials against a fine-tuned classifier
2. Labels or options change at run timeJevor LLMOptions and criteria are sent with each request; a trained classifier must be retrained. Caution: renaming options can change Jev's answers.TypeSafe primitives docs Documented; B04 "specify the question and labels at request time"; B09: no/yes names flip 32.5% of decisions vs 0/1 ReportedHow much accuracy moves when options are edited in practice
3. You need a free-text explanation or open outputLLMJev returns no text. TypeSafe lists "Generation: Use a generative model" among Jev 1.13's failure modes.Jaggedness page, failure mode 9 Documented; limitations—
4. Latency budget; CPU-only or on-deviceTight hosted budget: JevCPU or on-device: classifier, NLI or open modelStudies report Jev medians of about 0.13–0.6 s per call, several times faster than LLM calls. Jev has no public weights, so it cannot run on-device.B05, B06, B08, B15 Reported; B10 encoder 8 ms on CPU (grade D); no weights DocumentedLatency under load and near the rate limits
5. Data residency or offline requirementClassifier, NLI or open local modelJev runs only as a hosted API, from TypeSafe or gateways that route to TypeSafe. TypeSafe offers zero data retention directly "for enterprise customers"; some gateways label the TypeSafe provider zero-retention. Data still leaves your system on every route.TypeSafe legal index; route terms quoted on data handling by route (1 Oct) DocumentedTypeSafe subprocessors; EU residency not offered at the TypeSafe layer
6. You need a probability you can thresholdJevor LLM logprobs calibrated on your dataJev returns a probability per option. Most studies find it better calibrated than LLMs' stated confidence, but not everywhere.B06, B13, B15; B21 calibrated logprob adapter ≤0.10; B13 empathy task and B16 overconfidence Reported; see confidence thresholdsCalibration of Score; drift across versions
7. Input length and contextLong inputs: LLMJev takes 32k tokens for state plus the longest question and 64k per request. TypeSafe says large state full of irrelevant detail hurts accuracy.Models page and jaggedness page Documented; limitationsNo independent long-state measurement
8. Arithmetic, dates, other languages, imagesCode plus an LLMnot Jev aloneTypeSafe documents weakness on math, counting and date comparison; English is best; no image input.Jaggedness and models pages Documented; B08 math/code/logic, B23 numeric severity, B12 the only bilingual row ReportedGraded non-English and date studies
9. Volume and costHigh volume: Jev over a frontier LLMcheapest LLM setups or self-hosted classifiers can be cheaperPer-call cost differences are large against frontier LLMs, small against the cheapest LLM setups, and self-hosting changes the picture at scale.B03, B04 Reported. Figures: cost per decisionYour own token counts; Jev's fixed per-request token overhead is on the cost page
10. Vendor dependence and access riskClassifier or open modelOne vendor. Signups were paused 22–27 Sep; new-user credit is disabled; rate limits "can change without notice" and changed this week; every gateway routes to TypeSafe.Access status (30 Sep) Documented. A hosted "decision API" from OpenAI was reported on 29 Sep (The New Stack headline) ReportedThe OpenAI report was not checked at OpenAI's own source

A combined option: Jev first, then an LLM

Cascade pattern supported by three studies

  1. 1. Every item goes to JevJev returns an answer and a probability for each option.
  2. 2. Compare with a thresholdChosen in advance on your own labelled data. Above it, your code acts on Jev's answer.
  3. 3. Low confidence goes to an LLMOr to a person. Only this share pays LLM prices.
The threshold must be chosen on your own labelled data before use; confidence thresholds gives a procedure.
  • B08 (CMU): with a threshold frozen in advance, the cascade escalated 31% of 1,610 held-out pairs and was 0.9 points more accurate than GPT-6 at 41% of its fee.
  • B13 (NYU Abu Dhabi): routing low-confidence items to an LLM "matches or exceeds the LLM alone at a quarter to half of its cost".
  • B02 (Arize): an escalation recipe in its judging study.

Reported Each figure is the operator's own; the escalation share and savings depend on the task and threshold.

Head-to-head results from the ledger

Rows where Jev is compared with an LLM or a classifier. Each row shows who won on the operator's own metric. This is not a ranking, and the rows must not be summed. Reported

Row and taskWho won, on which metricOperatorGrade
B01 Four multi-question business workflowsOpus 5 and Sol above Jev on agreement with LLM reference answers; Jev level with Sonnet 5 and Terra; Jev lowest cost and time per caseTypeSafe (vendor)B COILLM-labels
B02 Hallucination detection; summary qualityTie with Opus 5 after threshold tuning (Opus ahead at 0.5); Jev slightly ahead on SummEval correlation; Terra behindArizeB
B03 AG News topics (n = 200)Jev ahead of gpt-4.1-nano, gpt-4o-mini and Luna on accuracy; nano with one token cheaper per decisionNavyaAIC
B04 Reranking; topic; freshnessTie with an internal reranker (NDCG@10); internal classifiers win on topic and freshnessParallelD
B05 Routing tiers (80 authored prompts)Jev ahead of Haiku 4.5 on match rate, latency and costLiteLLM (gateway)C COI
B06 PubMedQA, Banking77, HelpSteer2PubMedQA tie with Gemini 3.8 Flash; Banking77 Gemini ahead; HelpSteer2 nobody above base rateJevals (independent)B
B08 LLM-judge benchmarksWithin 3 points of GPT-6 on text-readable verdicts; GPT-6 ahead on math, code, logic; cascade ahead of GPT-6CMU (academic)B
B10 Banking77, Yelp and two moreTrained encoder ahead on Banking77; Jev ahead on Yelp; zero-shot NLI behind JevMindStudio (secondary)D
B12 Business email, 10 categories, German and EnglishGemini 3.5 and 3.8 ahead on accuracy (secondary figures)Bryo AIC
B13 18 social-science annotation tasksBest LLM per task ahead on 14 of 15 (macro-F1); Jev's confidence better than most LLMs' stated confidenceNYU Abu Dhabi (academic)A provisional
B15 True/false, multiple choice, injectionTrue/false: Jev level with Sol, ahead of GPT-4.1-mini. Multiple choice: Sol aheadpunk2898A provisional
B21 13 public datasetsJev ahead of a DIY logprob adapter and prompted JSON on average; DIY close on category taskskgluszczykC
B23 CVSS severity scoringGPT-6 Astra first; Jev seventh of eight by errorCascoC
Carried: general LLM with structured outputNo score. OpenAI's Structured Outputs guarantees schema conformance, not decision semantics or probabilitiesOpenAI docs (22 Sep report)Documented
Carried: GLiClassNo score. A local zero-shot or multi-label classifier; no Jev head-to-head recordedAlternatives hubNo score

Named open Jev-like models (Kev, Laya, CLM and others) are on the alternatives hub; hosted Jev, Laya and Kev are compared on the same rows on Jev vs Laya vs Kev; the one operator comparison of Jev and Kev is B11.

What was not verified

  • No approach was tested by this site on any task. The criteria come from TypeSafe's documentation and from other operators' published results.
  • No graded comparison with a fine-tuned classifier or zero-shot NLI exists with published materials (see the caveat above).
  • OpenAI's reported decision API was not checked at OpenAI's own source.
  • Per-route data terms are quoted on data handling by route (1 Oct); this page does not repeat them.

Search published pools, pages, reports, and evidence.