Jev vs Laya vs Kev: same-criteria comparison of Jev (TypeSafe AI) and two open models
Laya and Kev are two open-weights models often compared with hosted Jev: "jev vs laya" is a YouTube search suggestion, and several posts and videos compare them. This page sets the three side by side on the same rows: licences, base model, hardware, context, API, calibration, languages and maintenance. It then lists the only published tests that ran at least two of them, and ends with a "pick … if" table.
Only one same-items test finds an open model level with hosted Jev: Opper's 362 fresh items put Kev-4B within noise (Opper hosts Kev). JevBench scores Laya and earlier Kev previews well below Jev.
The main limits of each are Documented: Jev is hosted only and paid per input token. Laya's English checkpoint reads about 512 tokens. Kev is English-only and needs a GPU or a 32 GB Mac for the 4B model. This does not mean "Kev matches Jev": Opper says differences under about 5 points are within noise at its sample sizes, and Opper both serves Jev and hosts Kev.
How much text each one reads
Context per request, as each author states it (log scale)
- Hosted Jev32k for state plus the longest question; 64k per request
- Kev 0.8B, 4B, 9BTrained mostly on states of up to 384 tokens; the server accepts 65,536
- Kev-27BTrained on states of up to 32,768; the server accepts 65,536
- Laya multilingual1,024 by default; "up to 8,192" with
max_len=8192 - Laya typed-decisions1,024
- Laya (English)512
- Stated working range
- Range the small Kev models were mostly trained on
- Accepted, but the author warns accuracy may drop
90512f1. Documented 1 Oct 2026.Same rows for all three
Every cell is the model owner's own documentation, card metadata or repository record, read on 1 Oct 2026 between 05:16 and 05:24 UTC unless another date is shown. Documented "Not stated" and "not checked" are recorded as such. The pool ran none of the three models.
| Row | Hosted Jev | Laya | Kev |
|---|---|---|---|
| What it is, who makes it | Hosted API model jev-1.13.0 by TypeSafe AI; no weights published | Open-weights encoder decision models by Convai Innovations; code at NandhaKishorM/laya (card and code link to each other) | Open-weights decision models (0.8B, 4B, 9B, 27B) by Jared Palmer; code at jaredpalmer/kev |
| Code licence | No model code published; the client SDKs are MIT | Apache-2.0 | Apache-2.0 |
| Weights licence | No weights. Use is governed by TypeSafe's Terms and MCA; MCA 2.3(b) bars using Output to "perform model distillation" or build competing products | Apache-2.0 (card metadata of laya and laya-multilingual) | Apache-2.0 (card metadata of kev-4b, kev-9b, kev-27b) |
| Base model and its licence | Not disclosed | The card metadata has no base-model field. The card's checkpoint table names ModernBERT-large (English and typed-decisions; Apache-2.0) and mmBERT-base (multilingual; MIT) | Kev-4B: Qwen3.5-4B-Base (LoRA adapter); Kev-27B: Qwen3.8-27B (full fine-tune). Both Apache-2.0 in their Hugging Face metadata. Kev-9B's base licence: not checked |
| Size and hardware (author's statement) | Hosted; nothing to run | 421M (English) and 322M (multilingual) parameters. CPU or GPU; the card's timings are on one T4 GPU | Kev-4B: "~9 GB of GPU memory for serving", or a 32 GB Mac. Kev-27B: one 80 GB-class GPU (B200, H200 or H100); 51 GB of weights |
| Context | 64k per request; 32k for state plus the longest question | 512 (English); 1,024 by default, "up to 8,192" (multilingual); 1,024 (typed-decisions) | Server accepts 65,536 plus 8,192 per question; the small models trained mostly on 384 tokens or fewer, Kev-27B on up to 32,768 |
| API shape | POST /v1/systemone with Choice, Score and Noul | The same request and response shape through laya-serve. It listens on all interfaces with no authentication unless LAYA_API_KEY is set | The same shape, plus /permute, /separate and GET /v1/models; the README example uses TypeSafe's Python SDK |
| Local runtimes | None (hosted only) | Listed by Ollaya. Not in Ollama's official decision-model list (Ollama names nimble and tev1) | Listed by Ollaya; Kev's own server (CUDA, ROCm, MLX) and a Modal template. Not in Ollama's official list |
| Calibrated output | Yes, stated: trained "to return calibrated decisions" | English card: "mathematically calibrated probabilities". Multilingual card: "Ships uncalibrated" (temperature 1.0) | Yes, stated: each checkpoint ships a fitted temperature, applied by default |
| Languages | Not stated in the docs searched | 100+ through the multilingual checkpoint. The card warns the English checkpoint "collapses on non-Latin scripts" | English only |
| Latest dated change | jev-1.13.0; jev-latest and jev-preview point to it; no newer model | Release v0.3.22 (29 Sep, 17:48 UTC); cards last changed 24 Sep | Repository push 1 Oct, 03:58 UTC; Kev-9B and Kev-27B cards changed 30 Sep; no GitHub release since 20 Sep |
Kev's README states "No Jev outputs were used for training" (author statement, Reported). Licences for code, weights and base model are tracked separately because they differ: check the exact checkpoint you deploy.
Published tests that ran at least two of the three
All results are the operators' own. Reported Grades and flags come from the benchmarks ledger, where the grading method is explained; this page does not regrade, average or rank them.
Three graded rows, each on its operator's own scale
| Row | Operator and its interest | What it ran | Result as published | Grade |
|---|---|---|---|---|
| B11 Opper, "Jev vs. Kev" (25 Sep) | A gateway that serves Jev and hosts Kev | Jev and Kev-4B on the same route; 362 items published after 20 Sep, ground truth from the source | Within noise on accuracy; Jev's calibration error lower on all three tasks | B COIsmall-n |
| B07 JevBench v1.5.4 | Benchmark Heaven, which sells custom evaluations; no Jev stake stated | 1,624 decisions per system (720 sealed): Jev, three Laya checkpoints and Kev research previews | Jev 80.0; Kev previews 53.8–59.7; Laya 23.0–41.7 | B sealedno-baseline |
| B09 arXiv 2609.26758 v2 | Academic | 1,200 yes/no questions with option names swapped: hosted Jev, Laya, Open-Jev (no Kev) | Both models read option names, not only definitions; Laya more so (76.9% vs 32.5% flipped) | B public-datano-baseline |
| Laya card, "Laya (with routing) vs TypeSafe Jev" | Laya's author | typed-decisions (2,000 decisions), AG News, Emotion, Banking77 | Laya ahead on typed-decisions (with a checkpoint fine-tuned on that benchmark), AG News and Emotion; behind on Banking77 (0.425 vs 0.870). Its Jev figures are "never measured here" | D COI |
| Kev README and cards | Kev's author | Author suites on new sources | "Kev-27B is within a point of Jev (0.851 vs 0.857)"; the author adds that this "isn't a controlled comparison" | D COI |
No neutral operator has run Laya and Kev on the same items except JevBench (B07), where both sit well below Jev and the Kev rows are earlier previews. Seen but not graded (titles only, not watched or read in full): the videos "I Thought Jev vs Laya Would Be a Fair Fight. I Was Wrong." (Zahiruddin Tavargere), "Open-Source Jev? Kev vs Laya — We Tested Both" (RUNTIME., which does not run Jev) and "Jev vs Laya: Which One Should You Actually Use?" (AndyDevTech); and an Increscotech article that summarises the authors' own figures without a test of its own. Unverified
Pick … if
Each line rests on the rows above. It is a reading of documented limits and the graded evidence, not a measurement on your task.
| If you need… | Pick | Because | Main caveat |
|---|---|---|---|
| To run offline or on your own hardware | Laya (CPU) or Kev (GPU, or a 32 GB Mac for 4B) | Size and hardware, local runtimes | Only B11 shows near-parity with Jev, and only for Kev-4B, measured by a party that hosts Kev |
| Inputs over about 500 tokens | Jev or Kev; Laya only multilingual with max_len=8192 | Context | The small Kev models and Laya lose accuracy on long inputs, by their authors' own statements |
| A weights licence you can ship | Kev or Laya | Weights and base model rows | Check the base of the checkpoint you deploy; Laya's metadata names none. Jev has no weights, and MCA 2.3(b) bars distilling its outputs |
| A drop-in for existing System One calls | Any of the three | API shape | Secure Laya's server: no authentication unless LAYA_API_KEY is set |
| Languages other than English | Laya multilingual (its router sends non-English text there) | Languages | Kev is English only; Jev's languages are not stated; Laya's English checkpoint is confidently wrong on non-Latin scripts; the multilingual checkpoint ships uncalibrated |
| More than 20 options per question | Jev | Laya card: Banking77 0.425 vs Jev 0.870 (grade D) | Laya's option budget is shared across options |
| Robust answers when option names change | None proven | B09 | Avoid yes/no option names with Laya and Jev; test with neutral names (A/B, 0/1) |
Copies and identities
- Laya card and code belong together. The Hugging Face card links
NandhaKishorM/laya, and that repository's homepage field points back to the card. Documented - he-jev/laya is a copy of the Hugging Face model repository
convaiinnovations/layaat commitc5d78730(19 Sep 2026), not of the code repository. Its only branch head has the same commit id, date and message as the Hugging Face commit, and that id is not inNandhaKishorM/laya. It is 11 commits behind the Hugging Face repository. GitHub detects no licence on the copy. Who runshe-jevis not established. Documented - JevAny and AnyJev: whether they are the same project is still unresolved. Unverified
What was not verified
- The pool ran none of the three models and made no Jev call. Hardware figures are the authors' own.
- Whether JevAny and AnyJev are the same project, and who operates
he-jev. - Which Kev checkpoints JevBench ran.
- The base-model licence of Kev-9B.
- The "jev vs laya" videos (titles only) and the parsecailabs, binubabu and astgl posts (not opened).
- Ollaya's per-model typed-decisions scores were not attributed model by model, because the listing's layout is ambiguous.
Sources and check times
- Jev: TypeSafe models page (1 Oct 2026, 05:16 UTC); MCA dated 23 Sep 2026, date rechecked 1 Oct, 05:16 UTC.
- Laya: model card and multilingual card (05:20 UTC); card commits (05:21 UTC); code repository and releases (05:21 UTC).
- Kev: Kev-4B, Kev-9B and Kev-27B cards (05:20 UTC); README at commit 90512f1 (05:21 UTC).
- Base models: Hugging Face metadata of Qwen3.5-4B-Base, Qwen3.8-27B, ModernBERT-large and mmBERT-base (05:22 UTC).
- Copies and runtimes: he-jev/laya (05:21 UTC); Ollama 0.35 post and library search (05:20 UTC); Ollaya README (05:22 UTC).
- Tests: JevBench and arXiv 2609.26758 v2 (1 Oct, 05:20 UTC); Opper's article and results (graded on 30 Sep, 20:53 UTC; see B11).