Jev (TypeSafe AI) vs an LLM with JSON output vs a trained classifier: when to use which
For a classification, routing, gating or scoring task you can call Jev, ask a general LLM for structured output, train a classifier on your labels, or use a zero-shot NLI model. This page gives ten criteria for the choice, the published head-to-head results behind them, and a combined pattern several studies support.
Use Jev when you have few labels or the labels change, and you want a probability you can threshold. Use an LLM when you need text, long inputs, arithmetic or the best accuracy on hard judgments. Use a trained classifier (or an open local model) when you have thousands of labels for a fixed set, need CPU or on-device inference, or cannot send data out. Use Jev first, then an LLM for low-confidence items when you want LLM accuracy at lower cost.
All benchmark evidence here is Reported by the operators named on the benchmarks ledger; this site ran no test. There are no prices on this page: compare costs on cost per decision and with the worksheet there.
How strong is the classifier comparison?
Until 7 Oct 2026 the comparison between Jev and trained classifiers or zero-shot NLI rested only on grade D rows: B04 (Parallel: "Internal 'wins'" on topic and freshness, no n) and B10 (MindStudio's secondary summary of an unnamed study). No graded study with published materials compares Jev with a fine-tuned classifier such as BERT, SetFit or GLiClass, or with zero-shot NLI (B28 below uses its publisher's own guardrail classifiers; its code was not read). Criterion 1's "thousands of labels" advice is therefore a reasoned expectation plus weak evidence, not a measured result.
Added 7 Oct 2026: the ledger now has one grade B study with trained-classifier baselines, B28 (Red Hat, which publishes both classifiers; code not read, n not stated). Its result is mixed: Jev was behind a prompt-injection classifier and ahead of a small content-safety classifier.
Ten selection criteria
Each row says which approach it favours, why, and the evidence with its label. Rows are ordered by how early the question usually comes up, not by weight. Documented rows cite TypeSafe's docs (read 30 Sep 2026, 20:47–20:48 UTC); Reported rows cite ledger studies.
| Criterion | Favours | Why | Evidence | Still unknown |
|---|---|---|---|---|
| 1. Labelled data: none, some or thousands | None or some: Jevor LLMThousands, fixed labels: classifier | Jev and LLMs need no training. A few hundred labels make a threshold reliable. With thousands of labels for a fixed set, a trained model learns the label boundary directly and is likely to win on accuracy, speed and cost. | B02 thresholds tuned on half the labels; B17 recalibration Reported. B10 encoder 93.2% vs Jev 80.1% and B04, both grade D Reported | No graded study with materials against a fine-tuned classifier |
| 2. Labels or options change at run time | Jevor LLM | Options and criteria are sent with each request; a trained classifier must be retrained. Caution: renaming options can change Jev's answers. | TypeSafe primitives docs Documented; B04 "specify the question and labels at request time"; B09: no/yes names flip 32.5% of decisions vs 0/1 Reported | How much accuracy moves when options are edited in practice |
| 3. You need a free-text explanation or open output | LLM | Jev returns no text. TypeSafe lists "Generation: Use a generative model" among Jev 1.13's failure modes. | Jaggedness page, failure mode 9 Documented; limitations | — |
| 4. Latency budget; CPU-only or on-device | Tight hosted budget: JevCPU or on-device: classifier, NLI or open model | Studies report Jev medians of about 0.13–0.6 s per call, several times faster than LLM calls. Jev has no public weights, so it cannot run on-device. | B05, B06, B08, B15 Reported; B10 encoder 8 ms on CPU (grade D); no weights Documented | Latency under load and near the rate limits |
| 5. Data residency or offline requirement | Classifier, NLI or open local model | Jev runs only as a hosted API, from TypeSafe or gateways that route to TypeSafe. TypeSafe offers zero data retention directly "for enterprise customers"; some gateways label the TypeSafe provider zero-retention. Data still leaves your system on every route. | TypeSafe legal index; route terms quoted on data handling by route (1 Oct) Documented | TypeSafe subprocessors; EU residency not offered at the TypeSafe layer |
| 6. You need a probability you can threshold | Jevor LLM logprobs calibrated on your data | Jev returns a probability per option. Most studies find it better calibrated than LLMs' stated confidence, but not everywhere. | B06, B13, B15; B21 calibrated logprob adapter ≤0.10; B13 empathy task and B16 overconfidence Reported; see confidence thresholds | Calibration of Score; drift across versions |
| 7. Input length and context | Long inputs: LLM | Jev takes 32k tokens for state plus the longest question and 64k per request. TypeSafe says large state full of irrelevant detail hurts accuracy. | Models page and jaggedness page Documented; limitations | No independent long-state measurement |
| 8. Arithmetic, dates, other languages, images | Code plus an LLMnot Jev alone | TypeSafe documents weakness on math, counting and date comparison; English is best; no image input. | Jaggedness and models pages Documented; B08 math/code/logic, B23 numeric severity, B12 the only bilingual row Reported | Graded non-English and date studies |
| 9. Volume and cost | High volume: Jev over a frontier LLMcheapest LLM setups or self-hosted classifiers can be cheaper | Per-call cost differences are large against frontier LLMs, small against the cheapest LLM setups, and self-hosting changes the picture at scale. | B03, B04 Reported. Figures: cost per decision | Your own token counts; Jev's fixed per-request token overhead is on the cost page |
| 10. Vendor dependence and access risk | Classifier or open model | One vendor. Signups were paused 22–27 Sep; new-user credit is disabled; rate limits "can change without notice" and changed this week; every gateway routes to TypeSafe. | Access status (30 Sep) Documented. A hosted "decision API" from OpenAI was reported on 29 Sep (The New Stack headline) Reported | The OpenAI report was not checked at OpenAI's own source |
A combined option: Jev first, then an LLM
Cascade pattern supported by three studies
- 1. Every item goes to JevJev returns an answer and a probability for each option.
- 2. Compare with a thresholdChosen in advance on your own labelled data. Above it, your code acts on Jev's answer.
- 3. Low confidence goes to an LLMOr to a person. Only this share pays LLM prices.
- B08 (CMU): with a threshold frozen in advance, the cascade escalated 31% of 1,610 held-out pairs and was 0.9 points more accurate than GPT-6 at 41% of its fee.
- B13 (NYU Abu Dhabi): routing low-confidence items to an LLM "matches or exceeds the LLM alone at a quarter to half of its cost".
- B02 (Arize): an escalation recipe in its judging study.
Reported Each figure is the operator's own; the escalation share and savings depend on the task and threshold.
Head-to-head results from the ledger
Rows where Jev is compared with an LLM or a classifier. Each row shows who won on the operator's own metric. This is not a ranking, and the rows must not be summed. Reported
| Row and task | Who won, on which metric | Operator | Grade |
|---|---|---|---|
| B01 Four multi-question business workflows | Opus 5 and Sol above Jev on agreement with LLM reference answers; Jev level with Sonnet 5 and Terra; Jev lowest cost and time per case | TypeSafe (vendor) | B COILLM-labels |
| B02 Hallucination detection; summary quality | Tie with Opus 5 after threshold tuning (Opus ahead at 0.5); Jev slightly ahead on SummEval correlation; Terra behind | Arize | B |
| B03 AG News topics (n = 200) | Jev ahead of gpt-4.1-nano, gpt-4o-mini and Luna on accuracy; nano with one token cheaper per decision | NavyaAI | C |
| B04 Reranking; topic; freshness | Tie with an internal reranker (NDCG@10); internal classifiers win on topic and freshness | Parallel | D |
| B05 Routing tiers (80 authored prompts) | Jev ahead of Haiku 4.5 on match rate, latency and cost | LiteLLM (gateway) | C COI |
| B06 PubMedQA, Banking77, HelpSteer2 | PubMedQA tie with Gemini 3.8 Flash; Banking77 Gemini ahead; HelpSteer2 nobody above base rate | Jevals (independent) | B |
| B08 LLM-judge benchmarks | Within 3 points of GPT-6 on text-readable verdicts; GPT-6 ahead on math, code, logic; cascade ahead of GPT-6 | CMU (academic) | B |
| B10 Banking77, Yelp and two more | Trained encoder ahead on Banking77; Jev ahead on Yelp; zero-shot NLI behind Jev | MindStudio (secondary) | D |
| B12 Business email, 10 categories, German and English | Gemini 3.5 and 3.8 ahead on accuracy (secondary figures) | Bryo AI | C |
| B13 18 social-science annotation tasks | Best LLM per task ahead on 14 of 15 (macro-F1); Jev's confidence better than most LLMs' stated confidence | NYU Abu Dhabi (academic) | A provisional |
| B15 True/false, multiple choice, injection | True/false: Jev level with Sol, ahead of GPT-4.1-mini. Multiple choice: Sol ahead | punk2898 | A provisional |
| B21 13 public datasets | Jev ahead of a DIY logprob adapter and prompted JSON on average; DIY close on category tasks | kgluszczyk | C |
| B23 CVSS severity scoring | GPT-6 Astra first; Jev seventh of eight by error | Casco | C |
| Carried: general LLM with structured output | No score. OpenAI's Structured Outputs guarantees schema conformance, not decision semantics or probabilities | OpenAI docs (22 Sep report) | Documented |
| Carried: GLiClass | No score. A local zero-shot or multi-label classifier; no Jev head-to-head recorded | Alternatives hub | No score |
Named open Jev-like models (Kev, Laya, CLM and others) are on the alternatives hub; hosted Jev, Laya and Kev are compared on the same rows on Jev vs Laya vs Kev; the one operator comparison of Jev and Kev is B11.
What was not verified
- No approach was tested by this site on any task. The criteria come from TypeSafe's documentation and from other operators' published results.
- No graded comparison with a fine-tuned classifier or zero-shot NLI exists with published materials (see the caveat above).
- OpenAI's reported decision API was not checked at OpenAI's own source.
- Per-route data terms are quoted on data handling by route (1 Oct); this page does not repeat them.