Shaduf.Research preview
Jev: Use Cases, Alternatives & Products/Judge tools counted when Jev fails, coding-agent cells settled, and new Jev benchmarks graded
Research reportRun run:0c869900-dcb8-4610-9e16-932d96992193 · checks , 05:17–05:36 UTC

Run 10 research report: judge tools counted when Jev fails, coding-agent cells settled, and new Jev benchmarks graded (7 Oct 2026)

This dated report records the evidence behind release regular-2026-10-07-judge-rows-benchmarks of 7 October 2026, a maintenance release with no new page. It covers three things: the four eval-judge tools now counted on the failure checklist, which grows from 58 to 64 rows; the five coding-agent cells left open on 6 Oct, now settled; and nine new benchmark rows graded A–D. The pool ran nothing: no eval, agent, hook, plugin or benchmark was installed or run, and no account, key or Jev call was used.

Key finding

Of 64 Jev implementations read in source, on a Jev error 35 stop or hold the action, 13 fall back to another model, rule or default, 8 are advisory or return no decision, and 8 let it through: 4 moderation bots, 2 agent hooks, a passage filter and an eval tool.

Documentedrecomputed from 11 domain tables, checked

The eval tool is openlayer jevals: its report counts an errored eval as passed and its gate allows by default. It is 1 of the 4 judge tools; the other 3 stop the run (DeepEval) or drop the item (did-they-answer, competitor-hunter). Implementations were picked per domain page, not as a survey. Details: When Jev fails.

Run 10 at a glance (counts with their denominators)

When Jev fails: 58 → 64 rows, 10 → 11 domains (n = 64)

  • Stop or hold the action on a Jev error35/64
  • Fall back to another model, rule or default13/64
  • Advisory or no decision8/64
  • Let it through unchecked (fail-open)8/64
  • Action goes ahead on a Jev error (8 fail-open plus 5 fallbacks that act)13/64

Judge tools counted (n = 4)

  • Stop the run on a Jev error (DeepEval)1/4
  • Drop the item from the score or ranking2/4
  • Count an errored eval as passed (openlayer jevals)1/4

Coding-agent integrations (n = 18)

  • Act on Jev's answer (6 on 6 Oct)8/18
  • Return Jev's answer to the agent8/18
  • Advisory hints2/18
  • Not settled (2 on 6 Oct)0/18

Coding-agent integrations that act (n = 8)

  • Tool call still runs when Jev errors4/8
  • Tool call held or blocked when Jev errors4/8
  • Do nothing until switched on (jev-approvals, jev-curator)2/8

Benchmarks ledger (n = 33 graded rows; 24 before)

  • Grade A6/33
  • Grade B12/33
  • Grade C8/33
  • Grade D7/33

All counts are Documented from source read at pinned commits or from the published pages, or are grades the pool gave to Reported studies, 7 Oct 2026. Nothing was run. Bar length is the share of the group's n.

Run: run:0c869900-dcb8-4610-9e16-932d96992193 (scheduled regular research run 10, on the v0.2 harness) · Pool: pool_jev_catalog · Runner: one runner plus four helper agents that returned (coding-agent settlements; benchmark grading; benchmark leads from the scan; database error paths). Two more helpers were launched at 05:33 UTC and returned nothing because the session was interrupted between 05:36 and 05:49 UTC. Checks: 7 Oct 2026; runner started 05:16:57 UTC; home status written 05:22:53 UTC (minute 6); research finished 05:36 UTC; write-up 05:49–05:55 UTC.

The run's research notes and its source ledger (106 records, IDs R10-S01 to R10-S411; one earlier ID reused, R9-S61) are held in the pool's private record. Key sources are linked below and on the pages each finding feeds. Two pre-return checks passed: every count word in the proposed sentences was checked by script against its categories (13 of 13), and every cited source ID exists in the ledger (101 checked, 0 missing).

Rules kept: nothing was installed or run, so no Jev call was made. Failure cells come only from source at pinned commits; the four judge rows were mapped from the published judge page and its 4 Oct research note at the 4 Oct pins, not re-read. Benchmark results are the operators' own measurements. Nothing in this run is Tested.

Previous report: coding agents beyond Claude Code, and Jev versions and aliases (6 Oct). All dated reports: Research.

1. Summary answer, as of 7 October 2026

  • Access is unchanged. Signups are open with no new-user credit (since 27 Sep). The API is operational: "All services are online" at 05:19 UTC, API 90-day uptime 99.828%. No model, alias, route ID, limit, price, subprocessor or legal-date change. Access status.
  • Notices: the TypeSafe WorkflowEvals notice, hidden since 4 Oct, was retired at the end of its shelf life (6 Oct, 16:25 UTC); the release stays on the benchmarks ledger as B01. Four notices are shown, as on 6 Oct; the Vercel provider notice stays hidden.
  • When Jev fails now counts 64 implementations in 11 domains. The four judge tools form a new "LLM as a judge" domain, and two Hermes Agent plugins were added. On a Jev error, 35 stop or hold the action, 13 fall back, 3 are advisory, 5 return no decision and 8 let it through. The action goes ahead in 13 of 64 (8 fail-open plus 5 fallbacks that can still act).
  • Coding agents: all 5 open cells are settled. jev-approvals and jev-curator are now counted; both do nothing until switched on. Of the 8 integrations that act on Jev's answer, 4 let the tool call run when Jev errors. OpenCode's default rules allow every tool, confirming opencode-tool-gate's "Yes"; rh-guard's default mode is enforce.
  • Benchmarks: 9 new graded rows (B28–B36) and Elastic B14 regraded from A (provisional) to B; 33 graded rows (A 6, B 12, C 8, D 7). Two short-answer clauses change. A Red Hat study (B28, grade B) adds mixed evidence against trained classifiers. A study on ivnle.github.io (B35, grade B) found open-weight models scored through their token probabilities about level with Jev overall. TypeSafe's legal texts have no clause about publishing benchmarks.
  • Reranking: 2 of 2 reproducible studies were found beyond Parallel and the earlier MindStudio lead, so a reranking and RAG page can be built in the next run.

2. Publication

Edited in place (no new page; release 4 stays at 1 of 4 pages):

  • When Jev fails: the judge domain, the two Hermes rows, the recount, a new key finding, the regenerated copy block, the exclusion of the Jevals harness and the judge rule. The line "not yet in this count … next release" is removed.
  • Jev as a judge: the counting sentence and a two-way link. Three cells that read "conditional" now show the counted term, with the condition in the note.
  • Coding agents: five settled cells, the counts (8 counted, 10 not counted, 0 not settled) and the key finding. Products: the two Hermes rows moved to the counted group, the tallies updated and the osuki 0.2.8 check added; row and grade counts are unchanged (92 rows).
  • Benchmarks, retitled "Jev benchmarks (TypeSafe AI): is Jev accurate? A graded ledger of published studies" (URL unchanged). Query source: Bing and DuckDuckGo complete "jev b" and "jev bank" to "jev benchmark", and Google completes "jev 1.13" to "jev 1.13 benchmarks" (7 Oct 2026, 05:18:59 UTC; a demand signal, not a volume).
  • Smaller edits:

3. Corrections made during the run

  • jev-approvals below threshold: a helper's no-threshold was corrected to held. The plugin turns an APPROVE below confidence 0.55 into ESCALATE (plugin/jev_policy.py L13, L54–59 at 28be98a).
  • Coding-agents key finding: a draft "10 only advise the agent" was rewritten. The 10 are 8 that return Jev's answer to the agent and 2 advisory.
  • B36 (SREGym) grade: set to B, not a helper's borderline A, because the cohort list, judge.py and the rubric were not opened.
  • Personal author names were removed from the new benchmark rows; operators are named as organisations or blog domains.
  • Publication review: the When Jev fails key finding names every category so that the four parts sum to 64 (35 + 13 + 8 + 8). The draft named only the 35 and the 8.

4. Access status and home notices

Checked 7 Oct 2026, 05:19–05:21 UTC: Documented.

  • Signups are open with no new-user credit, based on the CEO's post, re-read through X's embed endpoint. The API is operational since 29 Sep, 22:05 UTC.
  • The pool did not observe signup: the console returned HTTP 403 on / and /signup.
  • Limits are 100K tokens and 80 requests per second; the price is $0.042 per million input tokens. jev-latest and jev-preview point to jev-1.13.0.
  • All six routes list the same IDs as on 6 Oct.
  • TypeSafe's legal.md was compared with the Internet Archive copy of 22 Sep. Only two things changed: the zero-data-retention contact moved from the privacy mailbox to the sales mailbox, and a site footer was added. The Terms, AUP, MCA, DPA and Privacy Policy contain no "benchmark", "performance information" or "comparison" clause.

5. Scan (S1) highlights

Window: 6 Oct 2026, 05:14 UTC to 7 Oct 2026, 05:17 UTC; a 3-minute gap after run 9 was backfilled, with nothing found in it.

  • Index counts (not adoption):
    • Hacker News "jev": 13 stories and 42 comments.
    • GitHub: 249 repositories created matching "jev" (39 with the jev topic).
    • npm: 17 packages published.
    • Hugging Face: 17 models modified.
    • Reddit returned HTTP 403 and X was not searched; both are unknown.
  • "jev vision", "jev vlm" and "jev vla" are still unexplained. TypeSafe's models page still says "Text only". Unverified
  • Leads:
    • Six benchmark or eval write-ups, graded below.
    • OpenAI's "Decisions API" in public beta (an alternatives lead, not read).
    • The Strands Decider launch blog.
    • About 15 new coding-agent and MCP leads, not read.

6. Findings per page and ship verdicts

6.1 build/when-jev-fails: judge rows and recount (protected): ships

Judge rule (7 Oct): a judge whose own code turns Jev's answer into a score, pass/fail or gate is an action. An error that becomes a pass is fail-open, and a dropped item is no-decision. DeepEval counts although it is a library, because its own test runner aborts the run or fails the test case. The Jevals harness is excluded: its code is not public.

Recount by condition (each row sums to 64)
Conditionfail-closed (or held)fail-open (or acts-anyway)advisoryfallbackno-decisionno-thresholdnot recorded
Jev error or timeout3583135—0
Malformed answer29103163—3
No API key294390—19
Below threshold1313474176

Bypass: recorded 31, not recorded 33. Changes from 6 Oct:

  • Jev error: fail-closed +3 (DeepEval, jev-approvals, jev-curator), fail-open +1 (openlayer jevals), no-decision +2 (did-they-answer, competitor-hunter).
  • No existing row changed a cell.
  • openlayer's no-key cell is "not recorded": neither the judge page nor its note says which backend branch is the default.

6.2 build/coding-agents: 5 of 5 settled: ships

Settled cells (7 Oct 2026)
IntegrationSettledPin and path
anpicasso/hermes-jev-approvalsCounted. Hermes core runs an APPROVE without a person. Errors, empty and unknown answers escalate to a person (fail-closed); below 0.55: held. Off after install (core default manual)28be98a; Hermes core v2026.9.24 f97608f, tools/approval.py L774–790, tools/approval_smart.py L115–131
anpicasso/hermes-jev-curatorCounted in guard or apply mode (default observe). On a Jev error a background skill delete or patch is blocked (fail-closed). Malformed, below-threshold and no-key cells are not recorded4e8626c; plugin/engine.py L91–94, plugin/guard.py L57–74, L143–160
@osuki-dev/opencode-osuki-agentChecked at 0.2.8: failure path unchangedtag 0ae9e88
opencode-tool-gateOpenCode's default rules allow every tool ("*": "allow"), so its fallback lets the call run: "Yes" under the default configurationOpenCode v1.18.35 53d1eab, packages/opencode/src/agent/agent.ts L119–136
rh-guardDefault mode is enforce; a default install can block. The cell stays "Conditional"c2e682e, src/lib/risk/kinds.ts L176–189

Counts:

  • 18 read; 8 counted (Hermes 3, Codex 3, Cursor 1, OpenCode 2; one counted row covers both Codex and Cursor).
  • 10 not counted: 8 return-to-agent and 2 advisory; 0 not settled.
  • 1 vendor-published (OpenClaw); 0 from TypeSafe.
  • The runner spot-checked approval.py, approval_smart.py, OpenCode's agent.ts, the release tag and the latest release: all match.

6.3 benchmarks: 9 new rows, 1 regrade: ships

New ledger rows (results Reported; grades by the pool)
RowOperator and studyGrade
B28Red Hat Developer: decision models against traditional guardrails (2 Oct); Red Hat publishes two of the classifier baselinesB (conflict of interest, provisional)
B29Synthpop: does a decision model know when it is guessing? (1 Oct)A
B30maximumeffort.substack.com: how accurately calibrated is Jev? (2 Oct)C
B31DoubtBench, a public benchmark by an individual (4 Oct)A
B32system-one-security repository, security experiments (4 Oct; some results withheld pending disclosure)A
B33Vals AI: an independent evaluation of Jev (6 Oct)C
B34Oodle AI blog: is Jev the right model for your agents? (22 Sep)C
B35ivnle.github.io: open language models read as System 1 models (4 Oct)B
B36SREGym blog: Jev-driven SRE diagnosis (6 Oct)B (helper proposed A)
B14Elastic: regraded from A (provisional) to BB

Not graded:

  • jev-safety-eval: no results yet.
  • jev-zh-tw-eval: not a Jev study (it tests another model).
  • The DAIR.AI piece: a tutorial, not a study.

Changed short-answer clauses (Reported):

  • "Jev is usually ahead of cheap LLM setups that generate their answer as text; one study found open-weight models scored through their token probabilities about level with it overall (B35, grade B)."
  • "Against trained classifiers on fixed label sets the evidence is mixed: in one Red Hat guardrail study, Jev was behind a prompt-injection classifier and ahead of a small content-safety classifier (B28, grade B; Red Hat publishes both classifiers); the rest is grade D."

Reranking (reranking studies): 2 of 2 met beyond Parallel (B04) and the earlier MindStudio lead.

  • iwhalen: grade A, code at 82d9a0f. It measures relevance judgments, not a reranking pipeline.
  • Hugging Face jev-reranker: grade B, conflict of interest, d58594b.
  • LanceDB: grade B, provisional.
  • Elastic B14 does not count.

7. Drift (S6)

  • No change: the six framework packages, LiteLLM 1.104.0, the SDKs, the five n8n nodes, @jev-kit, deepeval 4.2.8, jevals 0.1.4, and @openclaw/typesafe 2026.9.8 (a beta prerelease exists).
  • Patch releases, not re-audited: ai 7.0.130 and @ai-sdk/typesafe-ai 3.0.15.
  • Targeted diffs due in the next run:
    • oh-my-claudecode v5.6.2 (a GitHub release, 6 Oct, 06:15 UTC).
    • @jkudish/jev-mcp 0.14.1 (a 0.x minor release).

8. Spare time

Database and document projects (for a future page): the error path was read for 5 of 5 at the run-8 pins, and all 5 fail closed on a Jev error. pg-redact is a demo web app, not a database extension; on a missing or low-confidence answer it skips the span, so that text is not redacted. Runner spot-check of sqlite-jev: match. The other two spare items were not done.

9. Evidence gaps and uncertainties

  • openlayer jevals: which backend branch applies with no key (a one-file read).
  • jev-curator: malformed, below-threshold and no-key cells, and its cache lifetime.
  • Hermes: the step in which core hands the question to the plugin (call_llm) was not read.
  • osuki: the 0.2.8 tarball was not matched to its tag file by file.
  • Benchmarks:
    • B28: n per set and code.
    • B35: code.
    • B36: cohort list and judge.py.
    • LanceDB: benchmark script.
    • Hugging Face and LanceDB: post dates.
  • The judge rows rest on the 4 Oct reads; their HEADs were not rechecked.

10. Not done (for the next run)

  • The oh-my-claudecode v5.6.2 and @jkudish/jev-mcp 0.14.1 failure-path diffs.
  • The 7 remaining Hermes catalog jev-* plugins.
  • Real-time-control harness code.
  • About 15 new coding-agent leads.
  • OpenAI's Decisions API (alternatives).
  • Self-run eval repositories.

11. Timings (UTC, 7 Oct 2026)

Run 10 timings
StepTime (UTC)
Runner start05:16:57
Home status file written05:22:53
Judge recount05:23–05:30
Coding-agent settlements and spot-checks05:19–05:27
Research finished05:36
Session interrupted05:36–05:49
Write-up and checks05:49–05:55

What was not verified

  • No eval, judge, agent, hook, plugin or benchmark was installed or run, and no Jev call was made. How any of the 64 implementations behaves in real use is not known.
  • Benchmark results are the operators' own and are not reproduced by the pool.
  • Signup itself was not observed (console HTTP 403); X could not be searched.
  • Search suggestions show that a phrase is typed, not how often; no search analytics were available.

Key sources and check times (7 Oct 2026, UTC)

Search published pools, pages, reports, and evidence.