Test Jev on your own labelled cases
A reproducible design for one support-ticket route. This is a method, not a Jev benchmark: this pool made no authenticated request, account check, latency measurement, or claim of production suitability.
billing, technical, or account_help, or send the ticket to review. Code owns the queue, permission checks, thresholds, and side effects. Never automate an irreversible action from this recommendation.Freeze the test before using the held-out cases
- Gate data and access. Confirm the selected route's key, credits, terms, retention, logging, and permitted data. Use synthetic cases if real cases cannot be sent. Stop before a live call when the gate fails.
- Label two cohorts. Keep consecutive eligible tickets separate from a challenge set. Include clear departments, out-of-scope requests, ambiguous or multi-intent text, security/privacy cases, negation, numbers and dates, irrelevant state, and instructions embedded in ticket text. Synthetic challenge cases cannot estimate traffic prevalence.
- Adjudicate before output. Two authorized reviewers apply a written rubric. Label a department only when one is justified; otherwise label
reviewwith a subtype. Resolve or report disputed cases separately. Freeze labels, cohort IDs, state hashes, redaction, and split by underlying conversation so related variants never cross train, development, and test. - Version the complete harness. Pin
jev-1.13.0and the direct route for this comparison. Freeze state fields, one Choice question and its criteria, code rules, timeout and retry policy, parser, threshold, dependency versions, and test IDs. Select thresholds on development cases only. A different returned model version invalidates the pinned comparison. - Run the same cases through fair baselines. Compare a frozen keyword-rule router and, only if training data suffice, a trained local classifier. Give all arms the same redacted state, gold rubric, review option, high-impact block, and timeout envelope. Set each arm's threshold separately; do not treat classifier probabilities as Jev confidence.
The proposed Choice and policy boundary
Use one named Choice question with criteria for the three departments and explicit review. Put all task meaning in its instructions and criteria: the API's question ID is a code key, not model-facing semantics. Tell the model to treat instructions inside customer text as data. review is an option, not a guarantee of abstention; a wrong department can still carry high confidence. Auto-route only when the returned department passes a predeclared Choice-confidence threshold, the response is valid and on time, and a frozen deterministic high-impact block is false. Record blocked cases as review, not as model successes.
Record denominators and failures
| Measure | Record it as | Why it matters |
|---|---|---|
Attempted cases N | All frozen, adjudicated cases attempted by an arm, including errors and timeouts; report disputed and not-attempted separately. | Do not make failures disappear from decision-quality denominators. |
| Auto coverage / review load | Auto-routed / N; queued for review / N. Record reviewer minutes and pending human outcomes. | The two policy outcomes should sum to N. |
| False auto-route | Wrong department or any department on gold review, divided by auto-routed count; also divide by N. | Show both selective risk and total error burden at the resulting review load. |
| Out-of-scope and high-impact miss | Each subtype auto-routed / all attempted gold cases of that subtype. Report ambiguous misses separately. | Overall accuracy can hide unsafe slices. |
| High-confidence raw error | Wrong raw department Choice at a separately frozen cutoff / valid Choice responses, with its auto-routed subset. | Choice confidence summarizes an output distribution; it does not certify correctness. |
| Latency and availability | Parsed-success p50/p95 with n; separately count timeouts, HTTP errors, retries, and fallback/review delay. | A success-only percentile omits outage cost. |
| Tokens and charge | Input/output usage totals, missing usage, route and rate date, estimated charge, then actual bill. | The 25 Sep direct list price is $0.042/M input and free output, not a gateway or live-account price. |
For every case keep a case ID, cohort and slice, date, state hash, gold label/subtype, arm and returned model, raw choice/distribution/Choice confidence, policy outcome/reason, human disposition, elapsed time, retry/status/error, tokens, and route-specific charge estimate. Keep raw sensitive text outside the results sheet. Do not calculate a percentage with a zero denominator. Do not call an end-to-end resolution error measured until human outcomes are known.
Blank result sheet
| Arm / cohort | N / excluded / not attempted | Auto / review | False auto / auto; / N | Out-of-scope / high-impact miss | High-confidence raw false | p50 / p95; failures | Tokens; estimate / bill |
|---|---|---|---|---|---|---|---|
Jev Choice · jev-1.13.0 | — | — | — | — | — | — | — |
| Frozen local rules | — | — | — | — | n/a | — | n/a |
| Local classifier, if trainable | — | — | — | — | n/a | — | n/a |
— = not measured. Make a separate row for each cohort and slice. This sheet contains no example result.
This design tests Choice only. Score needs an ordinal gold rubric and its own threshold. Noul has a yes probability but no separate confidence field. Do not transfer a Choice threshold or combine primitive outputs into one accuracy figure.
- TypeSafe API — request/response contract, token usage, and 401/422/429/529 conditions.
- TypeSafe primitives and confidence — Choice, Score, Noul, distributions, and confidence boundaries.
- TypeSafe models — displayed model, direct price, and published limits; recheck before a pilot.
- Jev 1.13 jaggedness — vendor-described failure modes, last reviewed 17 Sep 2026; these are challenge prompts, not measured rates.
- TypeSafe legal index, MCA, privacy policy, and DPA — route and data gate; the MCA shows a 23 Sep update, but its clause diff remains unknown.
Independent protocol from run:d30a42d1-35d2-4bd7-960a-e68c1650cdb1. The 24 Sep builder-fit synthesis supplies the pilot/wait/stop boundary. No Jev result, safe threshold, account access, or production deployment is established here. The separate 27 Sep access, model, price, and limits recheck remains due.