Shaduf.Research preview
CompareModel cards, repositories and docs checked

Jev vs Laya vs Kev: same-criteria comparison of Jev (TypeSafe AI) and two open models

Laya and Kev are two open-weights models often compared with hosted Jev: "jev vs laya" is a YouTube search suggestion, and several posts and videos compare them. This page sets the three side by side on the same rows: licences, base model, hardware, context, API, calibration, languages and maintenance. It then lists the only published tests that ran at least two of them, and ends with a "pick … if" table.

Key finding

Only one same-items test finds an open model level with hosted Jev: Opper's 362 fresh items put Kev-4B within noise (Opper hosts Kev). JevBench scores Laya and earlier Kev previews well below Jev.

ReportedOpper (B11) and JevBench (B07), checked 1 Oct 2026

The main limits of each are Documented: Jev is hosted only and paid per input token. Laya's English checkpoint reads about 512 tokens. Kev is English-only and needs a GPU or a 32 GB Mac for the 4B model. This does not mean "Kev matches Jev": Opper says differences under about 5 points are within noise at its sample sizes, and Opper both serves Jev and hosts Kev.

How much text each one reads

Context per request, as each author states it (log scale)

  • Hosted Jev32k for state plus the longest question; 64k per request
  • Kev 0.8B, 4B, 9BTrained mostly on states of up to 384 tokens; the server accepts 65,536
  • Kev-27BTrained on states of up to 32,768; the server accepts 65,536
  • Laya multilingual1,024 by default; "up to 8,192" with max_len=8192
  • Laya typed-decisions1,024
  • Laya (English)512
  • Stated working range
  • Range the small Kev models were mostly trained on
  • Accepted, but the author warns accuracy may drop
Each step on the scale is four times the previous one. Laya's own table shows accuracy varying beyond about 4,000 tokens; Kev's README says "the smaller models lose accuracy on long documents". Sources: TypeSafe models page, Laya cards, Kev README at commit 90512f1. Documented 1 Oct 2026.

Same rows for all three

Every cell is the model owner's own documentation, card metadata or repository record, read on 1 Oct 2026 between 05:16 and 05:24 UTC unless another date is shown. Documented "Not stated" and "not checked" are recorded as such. The pool ran none of the three models.

Show
RowHosted JevLayaKev
What it is, who makes itHosted API model jev-1.13.0 by TypeSafe AI; no weights publishedOpen-weights encoder decision models by Convai Innovations; code at NandhaKishorM/laya (card and code link to each other)Open-weights decision models (0.8B, 4B, 9B, 27B) by Jared Palmer; code at jaredpalmer/kev
Code licenceNo model code published; the client SDKs are MITApache-2.0Apache-2.0
Weights licenceNo weights. Use is governed by TypeSafe's Terms and MCA; MCA 2.3(b) bars using Output to "perform model distillation" or build competing productsApache-2.0 (card metadata of laya and laya-multilingual)Apache-2.0 (card metadata of kev-4b, kev-9b, kev-27b)
Base model and its licenceNot disclosedThe card metadata has no base-model field. The card's checkpoint table names ModernBERT-large (English and typed-decisions; Apache-2.0) and mmBERT-base (multilingual; MIT)Kev-4B: Qwen3.5-4B-Base (LoRA adapter); Kev-27B: Qwen3.8-27B (full fine-tune). Both Apache-2.0 in their Hugging Face metadata. Kev-9B's base licence: not checked
Size and hardware (author's statement)Hosted; nothing to run421M (English) and 322M (multilingual) parameters. CPU or GPU; the card's timings are on one T4 GPUKev-4B: "~9 GB of GPU memory for serving", or a 32 GB Mac. Kev-27B: one 80 GB-class GPU (B200, H200 or H100); 51 GB of weights
Context64k per request; 32k for state plus the longest question512 (English); 1,024 by default, "up to 8,192" (multilingual); 1,024 (typed-decisions)Server accepts 65,536 plus 8,192 per question; the small models trained mostly on 384 tokens or fewer, Kev-27B on up to 32,768
API shapePOST /v1/systemone with Choice, Score and NoulThe same request and response shape through laya-serve. It listens on all interfaces with no authentication unless LAYA_API_KEY is setThe same shape, plus /permute, /separate and GET /v1/models; the README example uses TypeSafe's Python SDK
Local runtimesNone (hosted only)Listed by Ollaya. Not in Ollama's official decision-model list (Ollama names nimble and tev1)Listed by Ollaya; Kev's own server (CUDA, ROCm, MLX) and a Modal template. Not in Ollama's official list
Calibrated outputYes, stated: trained "to return calibrated decisions"English card: "mathematically calibrated probabilities". Multilingual card: "Ships uncalibrated" (temperature 1.0)Yes, stated: each checkpoint ships a fitted temperature, applied by default
LanguagesNot stated in the docs searched100+ through the multilingual checkpoint. The card warns the English checkpoint "collapses on non-Latin scripts"English only
Latest dated changejev-1.13.0; jev-latest and jev-preview point to it; no newer modelRelease v0.3.22 (29 Sep, 17:48 UTC); cards last changed 24 SepRepository push 1 Oct, 03:58 UTC; Kev-9B and Kev-27B cards changed 30 Sep; no GitHub release since 20 Sep

Kev's README states "No Jev outputs were used for training" (author statement, Reported). Licences for code, weights and base model are tracked separately because they differ: check the exact checkpoint you deploy.

Published tests that ran at least two of the three

All results are the operators' own. Reported Grades and flags come from the benchmarks ledger, where the grading method is explained; this page does not regrade, average or rank them.

Three graded rows, each on its operator's own scale

B07 JevBench: "Capability" score (0–100)

  • Jev 1.13.080.0
  • kev 8B (research preview)59.7
  • kev 4B (research preview)53.8
  • Laya typed-decisions41.7
  • Laya36.9
  • Laya multilingual23.0

B09: decisions flipped when options are named yes/no

  • Laya, yes/no names76.9%
  • Jev, yes/no names32.5%
  • Laya, 0/1 names6.5%

B11 Opper: accuracy on 362 fresh items (%)

  • arXiv category, n = 160: Jev96.9
  • arXiv category: Kev-4B95.0
  • Stack Exchange site, n = 120: Jev97.5
  • Stack Exchange site: Kev-4B98.3
  • GitHub bug or feature, n = 82: Jev95.1
  • GitHub bug or feature: Kev-4B93.9

Calibration error (lower is better): Jev 0.032 / 0.027 / 0.049; Kev-4B 0.043 / 0.125 / 0.138. Opper: "differences under about 5 points are within noise."

Blue: Jev. Green: Kev. Amber: Laya. Gray: a reference condition. The three charts use different tasks and metrics and must not be compared with each other. JevBench's Kev rows are earlier "research preview" checkpoints (0.6B, 4B, 8B), not the current Qwen3.5 and Qwen3.8 family; which checkpoint it ran was not checked.
RowOperator and its interestWhat it ranResult as publishedGrade
B11 Opper, "Jev vs. Kev" (25 Sep)A gateway that serves Jev and hosts KevJev and Kev-4B on the same route; 362 items published after 20 Sep, ground truth from the sourceWithin noise on accuracy; Jev's calibration error lower on all three tasksB COIsmall-n
B07 JevBench v1.5.4Benchmark Heaven, which sells custom evaluations; no Jev stake stated1,624 decisions per system (720 sealed): Jev, three Laya checkpoints and Kev research previewsJev 80.0; Kev previews 53.8–59.7; Laya 23.0–41.7B sealedno-baseline
B09 arXiv 2609.26758 v2Academic1,200 yes/no questions with option names swapped: hosted Jev, Laya, Open-Jev (no Kev)Both models read option names, not only definitions; Laya more so (76.9% vs 32.5% flipped)B public-datano-baseline
Laya card, "Laya (with routing) vs TypeSafe Jev"Laya's authortyped-decisions (2,000 decisions), AG News, Emotion, Banking77Laya ahead on typed-decisions (with a checkpoint fine-tuned on that benchmark), AG News and Emotion; behind on Banking77 (0.425 vs 0.870). Its Jev figures are "never measured here"D COI
Kev README and cardsKev's authorAuthor suites on new sources"Kev-27B is within a point of Jev (0.851 vs 0.857)"; the author adds that this "isn't a controlled comparison"D COI

No neutral operator has run Laya and Kev on the same items except JevBench (B07), where both sit well below Jev and the Kev rows are earlier previews. Seen but not graded (titles only, not watched or read in full): the videos "I Thought Jev vs Laya Would Be a Fair Fight. I Was Wrong." (Zahiruddin Tavargere), "Open-Source Jev? Kev vs Laya — We Tested Both" (RUNTIME., which does not run Jev) and "Jev vs Laya: Which One Should You Actually Use?" (AndyDevTech); and an Increscotech article that summarises the authors' own figures without a test of its own. Unverified

Pick … if

Each line rests on the rows above. It is a reading of documented limits and the graded evidence, not a measurement on your task.

If you need…PickBecauseMain caveat
To run offline or on your own hardwareLaya (CPU) or Kev (GPU, or a 32 GB Mac for 4B)Size and hardware, local runtimesOnly B11 shows near-parity with Jev, and only for Kev-4B, measured by a party that hosts Kev
Inputs over about 500 tokensJev or Kev; Laya only multilingual with max_len=8192ContextThe small Kev models and Laya lose accuracy on long inputs, by their authors' own statements
A weights licence you can shipKev or LayaWeights and base model rowsCheck the base of the checkpoint you deploy; Laya's metadata names none. Jev has no weights, and MCA 2.3(b) bars distilling its outputs
A drop-in for existing System One callsAny of the threeAPI shapeSecure Laya's server: no authentication unless LAYA_API_KEY is set
Languages other than EnglishLaya multilingual (its router sends non-English text there)LanguagesKev is English only; Jev's languages are not stated; Laya's English checkpoint is confidently wrong on non-Latin scripts; the multilingual checkpoint ships uncalibrated
More than 20 options per questionJevLaya card: Banking77 0.425 vs Jev 0.870 (grade D)Laya's option budget is shared across options
Robust answers when option names changeNone provenB09Avoid yes/no option names with Laya and Jev; test with neutral names (A/B, 0/1)

Copies and identities

  • Laya card and code belong together. The Hugging Face card links NandhaKishorM/laya, and that repository's homepage field points back to the card. Documented
  • he-jev/laya is a copy of the Hugging Face model repository convaiinnovations/laya at commit c5d78730 (19 Sep 2026), not of the code repository. Its only branch head has the same commit id, date and message as the Hugging Face commit, and that id is not in NandhaKishorM/laya. It is 11 commits behind the Hugging Face repository. GitHub detects no licence on the copy. Who runs he-jev is not established. Documented
  • JevAny and AnyJev: whether they are the same project is still unresolved. Unverified

What was not verified

  • The pool ran none of the three models and made no Jev call. Hardware figures are the authors' own.
  • Whether JevAny and AnyJev are the same project, and who operates he-jev.
  • Which Kev checkpoints JevBench ran.
  • The base-model licence of Kev-9B.
  • The "jev vs laya" videos (titles only) and the parsecailabs, binubabu and astgl posts (not opened).
  • Ollaya's per-model typed-decisions scores were not attributed model by model, because the listing's layout is ambiguous.

Sources and check times

Search published pools, pages, reports, and evidence.