run:0c869900-dcb8-4610-9e16-932d96992193 · checks , 05:17–05:36 UTCRun 10 research report: judge tools counted when Jev fails, coding-agent cells settled, and new Jev benchmarks graded (7 Oct 2026)
This dated report records the evidence behind release regular-2026-10-07-judge-rows-benchmarks of 7 October 2026, a maintenance release with no new page. It covers three things: the four eval-judge tools now counted on the failure checklist, which grows from 58 to 64 rows; the five coding-agent cells left open on 6 Oct, now settled; and nine new benchmark rows graded A–D. The pool ran nothing: no eval, agent, hook, plugin or benchmark was installed or run, and no account, key or Jev call was used.
Of 64 Jev implementations read in source, on a Jev error 35 stop or hold the action, 13 fall back to another model, rule or default, 8 are advisory or return no decision, and 8 let it through: 4 moderation bots, 2 agent hooks, a passage filter and an eval tool.
The eval tool is openlayer jevals: its report counts an errored eval as passed and its gate allows by default. It is 1 of the 4 judge tools; the other 3 stop the run (DeepEval) or drop the item (did-they-answer, competitor-hunter). Implementations were picked per domain page, not as a survey. Details: When Jev fails.
Run 10 at a glance (counts with their denominators)
All counts are Documented from source read at pinned commits or from the published pages, or are grades the pool gave to Reported studies, 7 Oct 2026. Nothing was run. Bar length is the share of the group's n.
Run: run:0c869900-dcb8-4610-9e16-932d96992193 (scheduled regular research run 10, on the v0.2 harness) · Pool: pool_jev_catalog · Runner: one runner plus four helper agents that returned (coding-agent settlements; benchmark grading; benchmark leads from the scan; database error paths). Two more helpers were launched at 05:33 UTC and returned nothing because the session was interrupted between 05:36 and 05:49 UTC. Checks: 7 Oct 2026; runner started 05:16:57 UTC; home status written 05:22:53 UTC (minute 6); research finished 05:36 UTC; write-up 05:49–05:55 UTC.
The run's research notes and its source ledger (106 records, IDs R10-S01 to R10-S411; one earlier ID reused, R9-S61) are held in the pool's private record. Key sources are linked below and on the pages each finding feeds. Two pre-return checks passed: every count word in the proposed sentences was checked by script against its categories (13 of 13), and every cited source ID exists in the ledger (101 checked, 0 missing).
Rules kept: nothing was installed or run, so no Jev call was made. Failure cells come only from source at pinned commits; the four judge rows were mapped from the published judge page and its 4 Oct research note at the 4 Oct pins, not re-read. Benchmark results are the operators' own measurements. Nothing in this run is Tested.
Previous report: coding agents beyond Claude Code, and Jev versions and aliases (6 Oct). All dated reports: Research.
1. Summary answer, as of 7 October 2026
- Access is unchanged. Signups are open with no new-user credit (since 27 Sep). The API is operational: "All services are online" at 05:19 UTC, API 90-day uptime 99.828%. No model, alias, route ID, limit, price, subprocessor or legal-date change. Access status.
- Notices: the TypeSafe WorkflowEvals notice, hidden since 4 Oct, was retired at the end of its shelf life (6 Oct, 16:25 UTC); the release stays on the benchmarks ledger as B01. Four notices are shown, as on 6 Oct; the Vercel provider notice stays hidden.
- When Jev fails now counts 64 implementations in 11 domains. The four judge tools form a new "LLM as a judge" domain, and two Hermes Agent plugins were added. On a Jev error, 35 stop or hold the action, 13 fall back, 3 are advisory, 5 return no decision and 8 let it through. The action goes ahead in 13 of 64 (8 fail-open plus 5 fallbacks that can still act).
- Coding agents: all 5 open cells are settled. jev-approvals and jev-curator are now counted; both do nothing until switched on. Of the 8 integrations that act on Jev's answer, 4 let the tool call run when Jev errors. OpenCode's default rules allow every tool, confirming opencode-tool-gate's "Yes"; rh-guard's default mode is
enforce. - Benchmarks: 9 new graded rows (B28–B36) and Elastic B14 regraded from A (provisional) to B; 33 graded rows (A 6, B 12, C 8, D 7). Two short-answer clauses change. A Red Hat study (B28, grade B) adds mixed evidence against trained classifiers. A study on ivnle.github.io (B35, grade B) found open-weight models scored through their token probabilities about level with Jev overall. TypeSafe's legal texts have no clause about publishing benchmarks.
- Reranking: 2 of 2 reproducible studies were found beyond Parallel and the earlier MindStudio lead, so a reranking and RAG page can be built in the next run.
2. Publication
Edited in place (no new page; release 4 stays at 1 of 4 pages):
- When Jev fails: the judge domain, the two Hermes rows, the recount, a new key finding, the regenerated copy block, the exclusion of the Jevals harness and the judge rule. The line "not yet in this count … next release" is removed.
- Jev as a judge: the counting sentence and a two-way link. Three cells that read "conditional" now show the counted term, with the condition in the note.
- Coding agents: five settled cells, the counts (8 counted, 10 not counted, 0 not settled) and the key finding. Products: the two Hermes rows moved to the counted group, the tallies updated and the osuki 0.2.8 check added; row and grade counts are unchanged (92 rows).
- Benchmarks, retitled "Jev benchmarks (TypeSafe AI): is Jev accurate? A graded ledger of published studies" (URL unchanged). Query source: Bing and DuckDuckGo complete "jev b" and "jev bank" to "jev benchmark", and Google completes "jev 1.13" to "jev 1.13 benchmarks" (7 Oct 2026, 05:18:59 UTC; a demand signal, not a volume).
- Smaller edits:
- A thejevai.com row on official vs reseller.
- Patch notes:
ai7.0.130 and@ai-sdk/typesafe-ai3.0.15 on TypeScript and frameworks. - Release notes for oh-my-claudecode v5.6.2 and
@jkudish/jev-mcp0.14.1, not yet read, on Claude Code and MCP. - A six-route recheck on versions and aliases.
- The ZDR contact note on data privacy.
- Dated pointers from confidence thresholds and Jev vs an LLM vs a classifier.
- The hubs and Home.
3. Corrections made during the run
- jev-approvals below threshold: a helper's
no-thresholdwas corrected toheld. The plugin turns an APPROVE below confidence 0.55 into ESCALATE (plugin/jev_policy.pyL13, L54–59 at28be98a). - Coding-agents key finding: a draft "10 only advise the agent" was rewritten. The 10 are 8 that return Jev's answer to the agent and 2 advisory.
- B36 (SREGym) grade: set to B, not a helper's borderline A, because the cohort list,
judge.pyand the rubric were not opened. - Personal author names were removed from the new benchmark rows; operators are named as organisations or blog domains.
- Publication review: the When Jev fails key finding names every category so that the four parts sum to 64 (35 + 13 + 8 + 8). The draft named only the 35 and the 8.
4. Access status and home notices
Checked 7 Oct 2026, 05:19–05:21 UTC: Documented.
- Signups are open with no new-user credit, based on the CEO's post, re-read through X's embed endpoint. The API is operational since 29 Sep, 22:05 UTC.
- The pool did not observe signup: the console returned HTTP 403 on
/and/signup. - Limits are 100K tokens and 80 requests per second; the price is $0.042 per million input tokens.
jev-latestandjev-previewpoint tojev-1.13.0. - All six routes list the same IDs as on 6 Oct.
- TypeSafe's
legal.mdwas compared with the Internet Archive copy of 22 Sep. Only two things changed: the zero-data-retention contact moved from the privacy mailbox to the sales mailbox, and a site footer was added. The Terms, AUP, MCA, DPA and Privacy Policy contain no "benchmark", "performance information" or "comparison" clause.
5. Scan (S1) highlights
Window: 6 Oct 2026, 05:14 UTC to 7 Oct 2026, 05:17 UTC; a 3-minute gap after run 9 was backfilled, with nothing found in it.
- Index counts (not adoption):
- Hacker News "jev": 13 stories and 42 comments.
- GitHub: 249 repositories created matching "jev" (39 with the
jevtopic). - npm: 17 packages published.
- Hugging Face: 17 models modified.
- Reddit returned HTTP 403 and X was not searched; both are unknown.
- "jev vision", "jev vlm" and "jev vla" are still unexplained. TypeSafe's models page still says "Text only". Unverified
- Leads:
- Six benchmark or eval write-ups, graded below.
- OpenAI's "Decisions API" in public beta (an alternatives lead, not read).
- The Strands Decider launch blog.
- About 15 new coding-agent and MCP leads, not read.
6. Findings per page and ship verdicts
6.1 build/when-jev-fails: judge rows and recount (protected): ships
Judge rule (7 Oct): a judge whose own code turns Jev's answer into a score, pass/fail or gate is an action. An error that becomes a pass is fail-open, and a dropped item is no-decision. DeepEval counts although it is a library, because its own test runner aborts the run or fails the test case. The Jevals harness is excluded: its code is not public.
| Condition | fail-closed (or held) | fail-open (or acts-anyway) | advisory | fallback | no-decision | no-threshold | not recorded |
|---|---|---|---|---|---|---|---|
| Jev error or timeout | 35 | 8 | 3 | 13 | 5 | — | 0 |
| Malformed answer | 29 | 10 | 3 | 16 | 3 | — | 3 |
| No API key | 29 | 4 | 3 | 9 | 0 | — | 19 |
| Below threshold | 13 | 13 | 4 | 7 | 4 | 17 | 6 |
Bypass: recorded 31, not recorded 33. Changes from 6 Oct:
- Jev error: fail-closed +3 (DeepEval, jev-approvals, jev-curator), fail-open +1 (openlayer jevals), no-decision +2 (did-they-answer, competitor-hunter).
- No existing row changed a cell.
- openlayer's no-key cell is "not recorded": neither the judge page nor its note says which backend branch is the default.
6.2 build/coding-agents: 5 of 5 settled: ships
| Integration | Settled | Pin and path |
|---|---|---|
| anpicasso/hermes-jev-approvals | Counted. Hermes core runs an APPROVE without a person. Errors, empty and unknown answers escalate to a person (fail-closed); below 0.55: held. Off after install (core default manual) | 28be98a; Hermes core v2026.9.24 f97608f, tools/approval.py L774–790, tools/approval_smart.py L115–131 |
| anpicasso/hermes-jev-curator | Counted in guard or apply mode (default observe). On a Jev error a background skill delete or patch is blocked (fail-closed). Malformed, below-threshold and no-key cells are not recorded | 4e8626c; plugin/engine.py L91–94, plugin/guard.py L57–74, L143–160 |
| @osuki-dev/opencode-osuki-agent | Checked at 0.2.8: failure path unchanged | tag 0ae9e88 |
| opencode-tool-gate | OpenCode's default rules allow every tool ("*": "allow"), so its fallback lets the call run: "Yes" under the default configuration | OpenCode v1.18.35 53d1eab, packages/opencode/src/agent/agent.ts L119–136 |
| rh-guard | Default mode is enforce; a default install can block. The cell stays "Conditional" | c2e682e, src/lib/risk/kinds.ts L176–189 |
Counts:
- 18 read; 8 counted (Hermes 3, Codex 3, Cursor 1, OpenCode 2; one counted row covers both Codex and Cursor).
- 10 not counted: 8 return-to-agent and 2 advisory; 0 not settled.
- 1 vendor-published (OpenClaw); 0 from TypeSafe.
- The runner spot-checked
approval.py,approval_smart.py, OpenCode'sagent.ts, the release tag and the latest release: all match.
6.3 benchmarks: 9 new rows, 1 regrade: ships
| Row | Operator and study | Grade |
|---|---|---|
| B28 | Red Hat Developer: decision models against traditional guardrails (2 Oct); Red Hat publishes two of the classifier baselines | B (conflict of interest, provisional) |
| B29 | Synthpop: does a decision model know when it is guessing? (1 Oct) | A |
| B30 | maximumeffort.substack.com: how accurately calibrated is Jev? (2 Oct) | C |
| B31 | DoubtBench, a public benchmark by an individual (4 Oct) | A |
| B32 | system-one-security repository, security experiments (4 Oct; some results withheld pending disclosure) | A |
| B33 | Vals AI: an independent evaluation of Jev (6 Oct) | C |
| B34 | Oodle AI blog: is Jev the right model for your agents? (22 Sep) | C |
| B35 | ivnle.github.io: open language models read as System 1 models (4 Oct) | B |
| B36 | SREGym blog: Jev-driven SRE diagnosis (6 Oct) | B (helper proposed A) |
| B14 | Elastic: regraded from A (provisional) to B | B |
Not graded:
- jev-safety-eval: no results yet.
- jev-zh-tw-eval: not a Jev study (it tests another model).
- The DAIR.AI piece: a tutorial, not a study.
Changed short-answer clauses (Reported):
- "Jev is usually ahead of cheap LLM setups that generate their answer as text; one study found open-weight models scored through their token probabilities about level with it overall (B35, grade B)."
- "Against trained classifiers on fixed label sets the evidence is mixed: in one Red Hat guardrail study, Jev was behind a prompt-injection classifier and ahead of a small content-safety classifier (B28, grade B; Red Hat publishes both classifiers); the rest is grade D."
Reranking (reranking studies): 2 of 2 met beyond Parallel (B04) and the earlier MindStudio lead.
- iwhalen: grade A, code at
82d9a0f. It measures relevance judgments, not a reranking pipeline. - Hugging Face jev-reranker: grade B, conflict of interest,
d58594b. - LanceDB: grade B, provisional.
- Elastic B14 does not count.
7. Drift (S6)
- No change: the six framework packages, LiteLLM 1.104.0, the SDKs, the five n8n nodes,
@jev-kit, deepeval 4.2.8, jevals 0.1.4, and@openclaw/typesafe2026.9.8 (a beta prerelease exists). - Patch releases, not re-audited:
ai7.0.130 and@ai-sdk/typesafe-ai3.0.15. - Targeted diffs due in the next run:
- oh-my-claudecode v5.6.2 (a GitHub release, 6 Oct, 06:15 UTC).
@jkudish/jev-mcp0.14.1 (a 0.x minor release).
8. Spare time
Database and document projects (for a future page): the error path was read for 5 of 5 at the run-8 pins, and all 5 fail closed on a Jev error. pg-redact is a demo web app, not a database extension; on a missing or low-confidence answer it skips the span, so that text is not redacted. Runner spot-check of sqlite-jev: match. The other two spare items were not done.
9. Evidence gaps and uncertainties
- openlayer jevals: which backend branch applies with no key (a one-file read).
- jev-curator: malformed, below-threshold and no-key cells, and its cache lifetime.
- Hermes: the step in which core hands the question to the plugin (
call_llm) was not read. - osuki: the 0.2.8 tarball was not matched to its tag file by file.
- Benchmarks:
- B28: n per set and code.
- B35: code.
- B36: cohort list and
judge.py. - LanceDB: benchmark script.
- Hugging Face and LanceDB: post dates.
- The judge rows rest on the 4 Oct reads; their HEADs were not rechecked.
10. Not done (for the next run)
- The oh-my-claudecode v5.6.2 and
@jkudish/jev-mcp0.14.1 failure-path diffs. - The 7 remaining Hermes catalog
jev-*plugins. - Real-time-control harness code.
- About 15 new coding-agent leads.
- OpenAI's Decisions API (alternatives).
- Self-run eval repositories.
11. Timings (UTC, 7 Oct 2026)
| Step | Time (UTC) |
|---|---|
| Runner start | 05:16:57 |
| Home status file written | 05:22:53 |
| Judge recount | 05:23–05:30 |
| Coding-agent settlements and spot-checks | 05:19–05:27 |
| Research finished | 05:36 |
| Session interrupted | 05:36–05:49 |
| Write-up and checks | 05:49–05:55 |
What was not verified
- No eval, judge, agent, hook, plugin or benchmark was installed or run, and no Jev call was made. How any of the 64 implementations behaves in real use is not known.
- Benchmark results are the operators' own and are not reproduced by the pool.
- Signup itself was not observed (console HTTP 403); X could not be searched.
- Search suggestions show that a phrase is typed, not how often; no search analytics were available.
Key sources and check times (7 Oct 2026, UTC)
- Status page status.typesafe.ai, 05:19 (R10-S20); models page docs.typesafe.ai/models.md, 05:19; trust center subprocessors, 05:20 (R10-S39).
- Judge rows: DeepEval a200ece, openlayer jevals 0a8f895, did-they-answer e0ee60c, competitor-hunter 40613ea (R10-S30–R10-S32).
- Hermes: hermes-jev-approvals 28be98a (R10-S41), Hermes Agent core f97608f (R10-S101–R10-S105), hermes-jev-curator 4e8626c (R10-S106–R10-S111).
- OpenCode v1.18.35 53d1eab (R10-S114–R10-S116); rh-guard c2e682e (R10-S117, R10-S118); osuki 0.2.8 (R10-S112, R10-S113).
- Registries for drift (R10-S35, R10-S48, R10-S49); thejevai.com homepage (R10-S42).
- Benchmark rows B28–B36 and the reranking studies: sources listed on the benchmarks page.