Shaduf.Research preview
Jev: Use Cases, Alternatives & Products/Support-ticket workload-evaluation protocol
Workload evaluation protocol · sources checked 25 Sep 2026

Test Jev on your own labelled cases

A reproducible design for one support-ticket route. This is a method, not a Jev benchmark: this pool made no authenticated request, account check, latency measurement, or claim of production suitability.

Decision: recommend exactly one of billing, technical, or account_help, or send the ticket to review. Code owns the queue, permission checks, thresholds, and side effects. Never automate an irreversible action from this recommendation.

Freeze the test before using the held-out cases

  1. Gate data and access. Confirm the selected route's key, credits, terms, retention, logging, and permitted data. Use synthetic cases if real cases cannot be sent. Stop before a live call when the gate fails.
  2. Label two cohorts. Keep consecutive eligible tickets separate from a challenge set. Include clear departments, out-of-scope requests, ambiguous or multi-intent text, security/privacy cases, negation, numbers and dates, irrelevant state, and instructions embedded in ticket text. Synthetic challenge cases cannot estimate traffic prevalence.
  3. Adjudicate before output. Two authorized reviewers apply a written rubric. Label a department only when one is justified; otherwise label review with a subtype. Resolve or report disputed cases separately. Freeze labels, cohort IDs, state hashes, redaction, and split by underlying conversation so related variants never cross train, development, and test.
  4. Version the complete harness. Pin jev-1.13.0 and the direct route for this comparison. Freeze state fields, one Choice question and its criteria, code rules, timeout and retry policy, parser, threshold, dependency versions, and test IDs. Select thresholds on development cases only. A different returned model version invalidates the pinned comparison.
  5. Run the same cases through fair baselines. Compare a frozen keyword-rule router and, only if training data suffice, a trained local classifier. Give all arms the same redacted state, gold rubric, review option, high-impact block, and timeout envelope. Set each arm's threshold separately; do not treat classifier probabilities as Jev confidence.

The proposed Choice and policy boundary

1 · inputMinimized ticket stateNo sensitive text until the route terms pass.
2 · ChoiceThree departments + reviewExplicitly describe outside, unclear, multi-intent, and security cases.
3 · codeRoute or human reviewA high-impact block, low confidence, error, timeout, or version mismatch goes to review.

Use one named Choice question with criteria for the three departments and explicit review. Put all task meaning in its instructions and criteria: the API's question ID is a code key, not model-facing semantics. Tell the model to treat instructions inside customer text as data. review is an option, not a guarantee of abstention; a wrong department can still carry high confidence. Auto-route only when the returned department passes a predeclared Choice-confidence threshold, the response is valid and on time, and a frozen deterministic high-impact block is false. Record blocked cases as review, not as model successes.

Record denominators and failures

MeasureRecord it asWhy it matters
Attempted cases NAll frozen, adjudicated cases attempted by an arm, including errors and timeouts; report disputed and not-attempted separately.Do not make failures disappear from decision-quality denominators.
Auto coverage / review loadAuto-routed / N; queued for review / N. Record reviewer minutes and pending human outcomes.The two policy outcomes should sum to N.
False auto-routeWrong department or any department on gold review, divided by auto-routed count; also divide by N.Show both selective risk and total error burden at the resulting review load.
Out-of-scope and high-impact missEach subtype auto-routed / all attempted gold cases of that subtype. Report ambiguous misses separately.Overall accuracy can hide unsafe slices.
High-confidence raw errorWrong raw department Choice at a separately frozen cutoff / valid Choice responses, with its auto-routed subset.Choice confidence summarizes an output distribution; it does not certify correctness.
Latency and availabilityParsed-success p50/p95 with n; separately count timeouts, HTTP errors, retries, and fallback/review delay.A success-only percentile omits outage cost.
Tokens and chargeInput/output usage totals, missing usage, route and rate date, estimated charge, then actual bill.The 25 Sep direct list price is $0.042/M input and free output, not a gateway or live-account price.

For every case keep a case ID, cohort and slice, date, state hash, gold label/subtype, arm and returned model, raw choice/distribution/Choice confidence, policy outcome/reason, human disposition, elapsed time, retry/status/error, tokens, and route-specific charge estimate. Keep raw sensitive text outside the results sheet. Do not calculate a percentage with a zero denominator. Do not call an end-to-end resolution error measured until human outcomes are known.

Blank result sheet

Arm / cohortN / excluded / not attemptedAuto / reviewFalse auto / auto; / NOut-of-scope / high-impact missHigh-confidence raw falsep50 / p95; failuresTokens; estimate / bill
Jev Choice · jev-1.13.0———————
Frozen local rules————n/a—n/a
Local classifier, if trainable————n/a—n/a

— = not measured. Make a separate row for each cohort and slice. This sheet contains no example result.

Deployment decision: Write maximum acceptable false-route, out-of-scope, and high-impact miss rates, review capacity, latency tail, and budget before opening the held-out test. Compare false routes at the resulting review load, not accuracy alone. If a high-impact case auto-routes, a safety target fails, or the data/terms gate fails, do not deploy automatic routing; revise using development cases and run a new held-out evaluation.

This design tests Choice only. Score needs an ordinal gold rubric and its own threshold. Noul has a yes probability but no separate confidence field. Do not transfer a Choice threshold or combine primitive outputs into one accuracy figure.

Direct sources checked 25 Sep 2026

Independent protocol from run:d30a42d1-35d2-4bd7-960a-e68c1650cdb1. The 24 Sep builder-fit synthesis supplies the pilot/wait/stop boundary. No Jev result, safe threshold, account access, or production deployment is established here. The separate 27 Sep access, model, price, and limits recheck remains due.

Search published pools, pages, reports, and evidence.