Shaduf.Research preview
Jev: Use Cases, Alternatives & Products/Jev limitations and fit check
Start hereJev 1.13 · limits checked

What Jev (TypeSafe AI) cannot do: documented limits and a task fit check

Jev answers typed questions from a fixed answer set. Tasks that need free text, raw images, exact arithmetic or an answer outside the list you gave it are a poor fit. This page lists the 17 limits TypeSafe documents, with its workaround for each, and has a three-question check you can run on your task.

Short answer

TypeSafe documents that Jev 1.13 is weak at literal-versus-intended reading, arithmetic and counting, date comparison, multi-step questions, long irrelevant state and adversarial text. It cannot generate text or take images, audio or video, and it works best in English. Keep computation, permissions and actions in your code, and use Jev to pick among options you define, score against a rubric, or estimate a yes probability.

Fit check: does your task fit?

What the three questions test

QuestionWhy it mattersTypeSafe's own guidance
1. Can you list the possible answers in advance?Choice takes up to 255 named options, Score an ordered rubric of at least 2 and at most 10 levels, Noul a yes/no question. Jev cannot return an answer you did not offer. Arithmetic and date comparison belong in code, not in the question."System One models do not write replies, produce code, or generate explanations of their reasoning. You define the possible answers through primitives." System One concepts
2. Will your own code decide what happens after the answer?A valid answer can still be wrong. In the audited projects the action policy lives in the calling code and differs a lot between projects."code owns the workflow and AI handles narrow, structured decisions … Keep control flow, deterministic rules, and side effects in code." How to build with System One
3. What happens when Jev is unsure, wrong or unavailable?Errors happen (401, 422, 429, 529 are documented). One audited trading gate lets an entry through when Jev fails; a moderation plugin holds the comment instead. Of six moderation projects read on 2 Oct, four leave the post up when Jev fails (content moderation). For a Noul, the middle band of probabilities is the "unsure" zone."Low confidence: Do not act. Route to a human, request clarification, or fall back to a different system." Confidence page

Documented Quotes checked 28 Sep 2026. The check runs entirely in your browser: it sends nothing, calls no model and stores nothing. It gives a starting recommendation, not an accuracy estimate. To measure accuracy on your own cases, see Confidence thresholds; for real examples, see the failure matrix.

Documented limits and workarounds (Jev 1.13)

Rows 1–10 come from TypeSafe's model jaggedness page ("Applies to jev-1.13. Last reviewed 2026-09-17"; section numbers shown). Rows 11–17 come from the Models page, the API reference and the primitives page. All were checked on 28 Sep 2026. Documented These are failure modes TypeSafe documents, not error rates anyone measured.

#LimitTypeSafe saysConsequence for youDocumented workaroundSource
1Literal reading"answers the question you wrote, not the one you meant"Scoping words and negations are taken at face valueState the exact condition; put boundary cases in the criteria; split an interpretive question into two literal ones and combine them in code§1
2Arithmetic and counting"Jev is not a calculator"; "does not count reliably"Sums, counts and numeric comparisons are unreliableDo arithmetic in code; ask one question per item and add up in code; turn numbers into named ranges§2
3A Score is not a number"score levels are weak in numerical calibration"You cannot rebuild a magnitude from scoreCompare score with a threshold; do not interpolate between levels§2
4Dates and times"reads dates as text, not as ordered quantities"Ordering, durations and date windows are unreliableExtract date parts with a Choice (with a "not stated" option) and compare in code; see the date-extraction cookbook§3
5Indirect questions"multiple hops of reasoning costs accuracy"Double negatives and property-of-a-property questions get worse answersAsk directly and name the relevant field of the state§4
6Large, irrelevant state"Accuracy falls as the state grows with content unrelated to the decision"Long context lowers accuracy and makes errors harder to traceRetrieve and filter in code first, or filter passages with a Noul; see the RAG passage cookbook§5
7Adversarial content"does not treat it as hostile by default"Injected or self-arguing text in the state can move the answerWrite explicit criteria and test edge cases before deployment§6
8Contradictory instructions"might get confused"Criteria that invert the instruction (for example in a Noul) perform worseKeep criteria consistent with the instruction§7
9No structural guaranteesA question and its negation need not agree (0.72 + 0.47 = 1.19)Arithmetic across questions and moving a threshold between question types are invalidAsk each decision one way; enforce identities in code§8
10No text generation"not trained to generate text"No free-form extraction, summaries, replies or codeProduce candidates with a regex or an LLM, then let Jev pick one with a Choice§9
11Text-only input"No image, audio, or video input."Images, audio and video need a separate conversion stepConvert to text or structured fields first, and check that step separatelyModels page
12Context limits64k tokens per request; 32k for state plus the longest questionLong documents must be cut or filteredFilter the state; ask several questions about one state in one requestModels page
13Languages"English is the primary training language… Other languages, including CJK scripts, are handled but not equally well"No multilingual accuracy figures are published"test on your own content"; watch confidence when routing non-English text (language handling in moderation projects: content moderation)Models page
14Closed answer setA Choice picks only from the options you give (at most 255)It cannot propose a new class, and it always picks oneAdd an "other" or "none of the above" optionAPI reference
15Questions are independent"one answer does not become context for another question"No chaining within one requestMake a second request from your code when one answer depends on anotherPrimitives page
16Rate limits are not fixedLimits "can change without notice"Capacity planning is uncertain. They changed twice: between 29 Sep, 19:24 UTC and 30 Sep, 20:47 UTC from 250,000 tokens/s and 1,200 requests/min to 100K tokens/s and 40 requests/s, and between 2 Oct, 05:14 UTC and 3 Oct, 05:16 UTC to 100K tokens/s and 80 requests/s (rechecked unchanged 7 Oct 2026, 05:19 UTC; current limits)Back off and retry; enterprise plans for guaranteed limits; see errors and rate limitsModels page
17No per-customer tuning"not fine-tuned or LoRA-adapted with customer data"Domain adaptation happens only through state, instructions and criteriaPut rules in the criteria; split decisions; train your own model on Jev outputs (AutoResearch cookbook)Models page

Reported, not documented: text in state that contains a shell curl … https:// command has been blocked by a Cloudflare HTML 403 before reaching the API (reported 23 Sep, still open). This matters for coding-agent guardrails; see errors and rate limits. Reported

When to use something else

  • You need generated text or code: an LLM, optionally with structured output.
  • You have thousands of labelled examples and a stable label set: a trained classifier may be cheaper and local; see alternatives.
  • You need the model to run on your own hardware: independent open-weight models, with their own licences and accuracy; see alternatives.
  • The rule can be written exactly (a regex, a threshold, a lookup): write the rule.

For a full comparison with 10 criteria (labelled data, changing labels, text output, latency, residency, probabilities, input length, arithmetic and language, volume, vendor risk) and the published head-to-head results, see Jev vs an LLM vs a classifier: when to use which. What published tests found for each task type is on the benchmarks ledger.

What was not verified

  • This pool did not test any limit by calling Jev; every row is TypeSafe's own documentation.
  • How often each weakness occurs on a real task is not measured here. The only published non-English test with a baseline is Bryo's German and English email benchmark (grade C, B12). TypeSafe's jaggedness page still said "Last reviewed 2026-09-17" on 30 Sep 2026.
  • Gateway-specific limits beyond the listed 32k context were not checked.

Search published pools, pages, reports, and evidence.