Shaduf.Research preview
Jev: Use Cases, Alternatives & Products/Jev benchmark grades and coding-agent settlement
Research reportRun run:4804d142-c2f9-40b1-8aa1-418259bd9a1a · checks , 05:15–05:40 UTC

Jev benchmark grades and coding-agent settlement: research of 11 October 2026

This dated report records the evidence behind release regular-2026-10-11-benchmark-grades-coding-agents of 11 October 2026. It refreshes benchmarks (three comparisons graded from opened files: 34 → 37 graded rows), Jev alternatives (Microsoft-Decision-1 added: 30 → 31 substitutes), coding agents (24 integrations read in source) and When Jev fails (71 → 74 rows). No new pages. The pool ran nothing: no package, plugin, model or benchmark was installed or run, and no Jev or OpenAI call was made.

Key findings

Benchmarks: Two independent grade-A comparisons of Jev with OpenAI's Decisions API differ by task: Decisions API ahead on RPG move choices (B38); Jev ahead on exact judge score in Portuguese, level on routing and injection (B39).

Reportedresults in the operators' own unitsDocumentedmetric and results files opened at pinned commits, checked , 05:22–05:26 UTC

Coding agents: Of 24 integrations read in source, 13 act on Jev's answer, 10 do not, 1 is unsettled. On a Jev error the 13: 5 let the call run, 3 hold, 5 fall back.

Documentedsource at pinned commits, checked , 05:17–05:40 UTC

When Jev fails: Of 74 Jev implementations read in source, on a Jev error 38 stop or hold the action, 20 fall back, 8 let it through, 5 return no decision and 3 are advisory.

Documentedrecounted by script from the published cells plus 3 new rows, checked , 05:30–05:40 UTC

No page sums or ranks the comparisons into an overall verdict: each result belongs to its task. The action still goes ahead on a Jev error in 15 of 74 implementations (8 fail-open plus 7 fallbacks that can still end in the action). Coding agents: every coding-agent integration the pool has read in source (24 on 11 Oct 2026), picked from the pool's scans and catalog reads; not a survey.

Run 13 at a glance (counts with their denominators)

Benchmarks ledger by grade (n = 37 graded rows; 34 on 10 Oct)

  • A (7 before; B38, B39, B40 added)10/37
  • B (unchanged)12/37
  • C (unchanged)8/37
  • D (unchanged)7/37

Alternatives by layer (n = 31; 30 on 10 Oct)

  • Model weights12/31
  • Local runtimes or research code11/31
  • Request (API) adapters4/31
  • Classifier or LLM approaches2/31
  • Hosted API from another provider (OpenAI Decisions API; Microsoft-Decision-1, new)2/31

Coding-agent integrations read in source (n = 24; 22 on 10 Oct)

  • Act on Jev's answer (counted)13/24
  • Do not act on Jev's answer (not counted)10/24
  • Not settled (Hoppin-Sorter/jev-router)1/24

On a Jev error, the 13 counted (n = 13)

  • Let the call run5/13
  • Hold the call3/13
  • Fall back to another model or the host's original path5/13

When Jev fails, on a Jev error (n = 74; 71 on 10 Oct)

  • Stop or hold the action (fail-closed or held)38/74
  • Fall back (fallback-<what>)20/74
  • Let it through (fail-open)8/74
  • Return no decision5/74
  • Advisory3/74
  • Action goes ahead (8 fail-open plus 7 fallbacks that can still act)15/74

All counts are Documented from source read at pinned commits, official pages or the published pages, 11 Oct 2026. Nothing was run. Bar length is the share of the group's n. The grade bar counts rows by grade; it is not a score for Jev.

Run: run:4804d142-c2f9-40b1-8aa1-418259bd9a1a (scheduled regular research run 13) · Pool: pool_jev_catalog · Runner: one runner plus two research helpers (A: real-time-control harnesses and drift; B: coding agents) and a spare-time helper. Checks: 11 Oct 2026; runner started 05:15:58 UTC; home status written 05:19:56 UTC; runner done 05:34:59 UTC; manager review 05:40 UTC.

The run's research notes and its source ledger (154 records, R13-S… IDs) are held in the pool's private record. Key sources are linked below and on the pages each finding feeds. Pre-return checks passed: 173 cited IDs in the research files plus 51 in the report, 0 missing; a recheck of 74 IDs after the manager's decisions, 0 missing. 14 count-word sentences were checked: every category list adds up to its total, and every category with 3 or more items (or 10% of the total) is named.

Rules kept: nothing was installed or run, so no Jev or OpenAI call was made. Failure cells come only from source at pinned commits or from published package files read in memory. Study results are the operators' own measurements. Nothing in this run is Tested.

Previous report: 10 Oct research report. All dated reports: Research.

1. Summary answer, as of 11 October 2026

  • Access is unchanged (11 Oct 2026, 05:17 UTC). Status open_with_conditions since 27 Sep (signups open, no new-user credit); the API is operational. Access status.
  • Benchmarks. The three comparisons listed as leads on 10 Oct are now graded A from their metric and results files opened at pinned commits (B38, B39, B40). The ledger has 37 graded rows: A 10, B 12, C 8, D 7. The two comparisons with OpenAI's Decisions API point in different directions by task. B40 is a reranking-pipeline study in which Jev and Cohere Rerank 4 Pro are within noise.
  • Alternatives. Microsoft-Decision-1, a hosted decision model published by Microsoft on 9 Oct, is added to the "Hosted API from another provider" layer. The page now lists 31 substitutes: 12 weights, 11 runtimes, 4 adapters, 2 classifier or LLM approaches and 2 hosted APIs.
  • Coding agents. Three rows are now counted: jev-triage (OpenClaw), the jev-for-all launcher, and hermes-jev-mana (Hermes Agent), which is inert after its documented install. Hoppin-Sorter/jev-router is read and not settled.
  • When Jev fails. The three new rows are routing fallbacks: n = 74, routing fallbacks 5 → 8, action goes ahead 15 of 74 (unchanged number).
  • Fast control loops (research for run 14). The condition is met: two third-party harnesses that call Jev were read at pinned commits with every failure cell recorded. Nothing is published on use-cases/real-time-control this run (section 6).

2. New rule of 11 Oct 2026: grades rest on opened files

Rule: a comparison is graded only from its metric file and its results file, opened at a pinned commit. A README number counts only when the opened results file contains it. Until then the comparison is shown as "lead, not graded", with no grade and no grade colour.

Reason: on 10 Oct three published comparisons could only be listed as leads, because their metric and results files had not been opened. The alternatives page needed grades to compare Jev with another provider's hosted API.

Standing check from 11 Oct: a graded comparison carries the same grade, flags, operator, task and date on benchmarks, alternatives, Laya vs Kev and reranking for RAG. No page sums, averages or ranks the comparisons into an overall verdict. A sentence that summarises them names each task and who came out ahead in it, as the operator reported.

3. What this release changed

No new pages. Edited in place:

  • Benchmarks: rows B38, B39 and B40 (grade A, with flags); new key-finding clause; "Added 11 Oct" paragraph; grade counts 37 = A 10, B 12, C 8, D 7; leads for classifier-bakeoff, decisionbench and 10 new comparison repositories.
  • Jev alternatives: Decisions API evidence cell rebuilt from graded rows; new Microsoft-Decision-1 row; counts 30 → 31.
  • Coding agents: new key finding; jev-triage, jev-for-all and hermes-jev-mana counted; Hoppin-Sorter/jev-router not settled; wording fixes for jev-effort-router and jev-approvals; triage table extended; README-only candidate gates 8 → 6.
  • When Jev fails: 3 new rows; recount n = 71 → 74; new key finding.
  • One linked line each on Laya vs Kev (B39) and reranking for RAG (B40).

4. Corrections to published pages

  • Alternatives, Decisions API "Parity or quality evidence" cell. It said "Lead, not graded" for every comparison. Two of them are now graded A (B38, B39). The cell is rebuilt from the graded rows only, and only that row is re-dated (section 5.2).
  • Alternatives, "Hosted API from another provider: 1 entry". Microsoft published Microsoft-Decision-1 on 9 Oct (R13-S32), so the count of 1 was incomplete from 9 Oct. It is now 2, and the total is 31.
  • Coding agents, jev-effort-router row. The failure cells hold. Two wording fixes: "routing on in its defaults" is true of the plugin's own config only, since Hermes plugins are off until enabled; and with Hermes defaults (model: "", config_defaults.py L21–22; R13-S217) the plugin does nothing, because it acts only when the provider is Ollama:Cloud.
  • Coding agents, jev-approvals row. "May ask the main model" is right only when a main model is configured. Added: "(with no main model configured, it asks you)".
  • Coding agents, jev-triage cells. Below threshold was published as "acts-anyway"; the source shows the message goes to the main model untouched (index.js L45; R13-S200), so the cell is now "fallback-main model". Malformed answer was "not recorded" on 10 Oct and is now settled as "fallback-main model".

5. Findings per page

5.1 benchmarks (protected): B38–B40 graded

Each row was graded from its metric file, its results file and the data reference, opened at the pinned commit (one tree call per repository). Results are Reported in the operator's own units; nothing was recomputed or run. Grade scale as published: A = data, code and configuration are published and the operator states no commercial stake. Flags are listed separately and never averaged into a grade. Each operator is an individual with no stake stated.

New graded rows (checked 11 Oct 2026, 05:23 UTC)
RowOperator and pinTaskResult as reported (operator's units)GradeFlags
B38komzweb/jev-vs-decisions-rpg at f81c004 (9 Oct)Choosing NPC moves in a small turn-based RPG: Jev jev-1.13.0 vs OpenAI Decisions API gpt-6-luna; 100 matches; 500 answers per model on the per-move testDecisions API won 82 of 100 matches (Jev 8; 10 draws). Per-move accuracy: Decisions API 59.4%, Jev 51.0% (p = 0.007). Calibration error: Jev 0.048, Decisions API 0.161AGame task, not a production workload; task built by the operator; one run; full call logs not committed
B39mateusfonsek/openai-decisions-vs-jev-vs-laya at 160e112 (8 Oct)Portuguese, three tasks of 150 cases each: agent routing, prompt-injection detection, LLM-as-judge (1–5 score). Jev vs Decisions API vs Laya run locallyJev ahead on the exact judge score: 80.7% vs 72.0% (p = 0.011, as stated in the README). Decisions API ahead on judge scores within ±1 point: 100.0% vs 98.7%, and on injection precision: 100.0% vs 97.3% (false positives on hard negatives 0.0% vs 4.0%). Routing 88.0% (Jev) vs 86.7% and injection F1 96.6% vs 95.8%: not significant, per the operator. Laya 30.7% on routingAPortuguese only; dataset generated by one LLM (Claude); one run per case; Jev called via the jev-latest alias; source of the p-value not located
B40anessbelbati/jev-rerank-bench at fecba75 (25 Sep)Reranking-pipeline study: each system reorders a 30-passage BM25 shortlist; 8 English datasets (1,617 questions) plus BRIGHT subsets, NevIR and MIRACL FrenchnDCG@10 averaged per dataset: Jev 0.692, Cohere Rerank 4 Pro 0.691, within noise (p = 0.906). Weighting each query equally puts Cohere ahead: 0.756 vs 0.738AReranking only; datasets chosen by the operator; order depends on the averaging; per-dataset results go both ways
Who came out ahead, task by task, as each operator reported (no overall verdict)
StudyTask and metricAhead
B38RPG matches won (of 100)Decisions API (82 vs 8)
B38RPG per-move accuracyDecisions API (59.4% vs 51.0%)
B38RPG calibration error (lower is better)Jev (0.048 vs 0.161)
B39Portuguese judge, exact scoreJev (80.7% vs 72.0%)
B39Portuguese judge, within ±1 pointDecisions API (100.0% vs 98.7%)
B39Portuguese routing accuracyLevel: 88.0% vs 86.7%, not significant per the operator
B39Portuguese injection F1Level: 96.6% vs 95.8%, not significant per the operator
B39Portuguese injection precisionDecisions API (100.0% vs 97.3%)

The anchors #lead-decisions-rpg, #lead-decisions-laya and #lead-rerank-bench stay on the page and point to B38, B39 and B40.

Leads, not graded (no grade, no grade colour):

  • alpersonalwebsite/classifier-bakeoff at 4a1c4a0: report opened, metric code not opened. The report states 200 synthetic messages written by a Claude model, with Decisions API 88.0%, Sonnet 86.0% and Jev 84.0%, "all within noise" Reported (R13-S67).
  • KizitoNaanma/decisionbench: not done this run; results location not identified.
  • 10 new comparison repositories from this run's scan, README or description only, not opened (R13-S23): kartikanand73/decision-api-benchmark, etsabary/luna-decisions-benchmark, janagarajsn/jev-vs-laya-vs-vega, Mike-Animal-Counseling/laya-jev-lab, arghya05/jev-boundary-study, truongxxxx/jev-polymarket-test, Markonick/jev-bakeoff, davra02/req-classifier, gbozelli/pokemon_match, alexcpn/llm-invaders.

Of the 10 A rows, 3 are provisional in the page's own sense (B13, B15, B24) and 7 were read at pinned commits (B29, B31, B32, B37, B38, B39, B40).

5.2 alternatives (protected): evidence cell and Microsoft-Decision-1

Count basis: the 30 published rows plus Microsoft-Decision-1, counted by layer by script. A row with two layers counts once, under its primary layer.

Alternatives by layer (n = 31, 11 Oct 2026)
LayerRowsNew this run
Model weights12none
Local runtime or research code11none
Request (API) adapter4none
Classifier or LLM approach2none
Hosted API from another provider2Microsoft-Decision-1 (was 1 row: OpenAI Decisions API)
New and changed rows (checked 11 Oct 2026, 05:17–05:26 UTC)
RowWhat it isPriceParity or quality evidence
OpenAI Decisions APIHosted API from another provider (unchanged; only this cell re-dated)unchangedGraded comparisons Reported, checked 11 Oct 2026. RPG NPC move choices: the Decisions API came out ahead (82 of 100 matches; per-move accuracy 59.4% vs 51.0%) (B38, A, komzweb, 9 Oct). Portuguese routing, injection detection and judging: Jev ahead on exact judge score (80.7% vs 72.0%); the Decisions API ahead within ±1 point (100.0% vs 98.7%) and on injection precision; the operator reports no significant difference on routing or injection (B39, A, mateusfonsek, 8 Oct). Results differ by task; the pool states no overall parity. classifier-bakeoff is a lead, not graded.
Microsoft-Decision-1 (R13-S32, R13-S33)Hosted decision model from another provider, published 9 Oct. Documented by Microsoft: "available in Microsoft Foundry and through OpenRouter". Not in OpenRouter's /models list fetched at 05:17 UTC on 11 Oct, so OpenRouter availability rests on Microsoft's statement. Posted on HN as item 50024913. Data terms not read. An alternative, not a product built on Jev."Input tokens cost $0.042 USD per million tokens. Output tokens are free" (Microsoft)No graded comparison. One lead: kartikanand73/decision-api-benchmark (not opened). The pool makes no price or quality comparison.

5.3 build/coding-agents: settlement

Count basis: 22 published rows plus 2 README-only gates read this run (jev-for-all, Hoppin-Sorter/jev-router) = 24. Manager decisions of 11 Oct 2026, 05:40 UTC: hermes-jev-mana counted under the 7 Oct rule (plugins that are off after install still count, with the state stated next to the row); jev-for-all counted on its launcher only.

  • 5 let the call run + 3 hold + 5 fall back (1 may ask the main model: jev-approvals; 1 keeps its configured model: jev-effort-router; 1 passes the message to the main model: jev-triage; 1 uses its default tier: jev-for-all; 1 runs the original Hermes branch: hermes-jev-mana) = 13 counted.
  • 8 return Jev's answer to the agent + 2 give advisory hints = 10 not counted.
  • 1 not settled (Hoppin-Sorter/jev-router).
  • 13 + 10 + 1 = 24.
  • Change: read 22 → 24; counted 10 → 13; not counted 10 (unchanged); not settled 2 → 1. On a Jev error: held 3/10 → 3/13; the call runs 5/10 → 5/13; fall back to a model or the original path 2/10 → 5/13.
  • Under the 7 Oct rule, 3 of the 13 counted rows do nothing until switched on: jev-approvals, jev-curator and hermes-jev-mana.
New counted rows: failure cells (source at pinned commits, 11 Oct 2026)
IntegrationJev error or timeoutMalformedBelow thresholdNo keyTimeoutBypass
jev-triage (OpenClaw) at 6360e21; host OpenClaw at aa6008a. Skips the main model for "thanks" and "spam" at a score of 0.9 or morefallback-main modelthe host catches the failure and the message goes to the main modelfallback-main modelsettled this run (was not recorded)fallback-main modelbelow 0.9 the message goes on untouched (was acts-anyway)fallback-main modela non-ok status returns nothing5,000 msrecorded
jev-for-all start launcher (Claude Code, OpenCode) at npm 0.7.1 (gitHead 74f62e7). Counted on the launcher only; the OpenCode plugin part narrows the tool list and is not settledfallback-default modelstandard tier (Claude Sonnet), medium effortfallback-default modelfallback-default modelstandard tier unless light ≥ 0.6 or heavy ≥ 0.5fallback-default model3,000 msrecorded
hermes-jev-mana (Hermes Agent) at 3daf639. Inert after the documented install: every Jev call is refused because admit_unlabeled defaults to false; acts only if the operator sets admit_unlabeled: true.fallback-original Hermes branchon the configured modelfallback-original Hermes branchfallback-original Hermes branchnear-tie margin 0.15not recordedclient key handling not read1.0 s per call, 5.0 s per turnrecorded
  • fallback-<what>: a named substitute
  • not recorded: not read in source
  • Not settled: Hoppin-Sorter/jev-router (Claude Code, Codex, Cursor) at bdb4721 rewrites the model on every main-loop request in its default "auto" mode. How the host engine handles the rewrite cannot be read at a pinned commit, because the Claude Code engine source is not public.
  • README-only candidate gates: 8 → 6 (jev-for-all and Hoppin-Sorter/jev-router were read in source).
  • Drift: @openclaw/typesafe 2026.10.1 was released 10 Oct; the jev-triage host branch, read at 2026.10.1 up to L650, has the same effect.
Triage of untriaged scan items (README or package description only, Reported; on the page)
ItemsCandidate gateClaude CodeReturns to the agentNot a coding-agent integrationIn-host shapingNot readableTotal
Run 12's scan items74522020
Run 13's scan items33310111

Combined: 31 lines, 30 unique projects (one npm package is the same project as a GitHub repository).

5.4 build/when-jev-fails: recount 71 → 74 (protected)

A script read the published cells (71 rows) and first reproduced the published error split, 38 / 17 / 8 / 5 / 3. It then added the three rows counted on coding agents. Each is a routing fallback: when a decision fails, the work goes to the main, default or configured model by the original path.

On a Jev error (n = 74)
OutcomeCountChange since 10 Oct
Stop or hold38unchanged
Fall back2017 → 20 (jev-triage, jev-for-all, hermes-jev-mana)
Let it through8unchanged
No decision5unchanged
Advisory3unchanged
Total7471 → 74
  • Routing fallbacks: 5 → 8.
  • Action goes ahead on a Jev error: 15 of 74 (8 fail-open plus 7 non-routing fallbacks; the number is unchanged).
  • No key "not recorded": 18 → 19 (hermes-jev-mana).
  • The sentence on the 8 that let it through (4 moderation bots, 2 agent hooks, a passage filter and an eval tool) is unchanged.

6. Access status and home notices

Checked 11 Oct 2026, 05:17 UTC: Documented. Status unchanged.

  • Status open_with_conditions since 27 Sep 2026, 22:33:25 UTC: signups open, no free credit for new accounts. Service operational ("All services are online"). Limits, price and aliases unchanged since 10 Oct (R13-S05, R13-S06).
  • Shown on Home (2): new-user credit disabled (rank 1); lookalike resellers (rank 2).
  • Retired: none. New notices: none.
  • The status page also lists api.us.typesafe.ai ("served from the closest US region") beside the regional hostname it already listed. TypeSafe's docs do not document these hostnames (11 Oct 2026, 05:17 UTC; R13-S05, R13-S07). Note only. Access status.
  • Microsoft-Decision-1 is an alternative from another provider, not a change to Jev access.

7. Game and robotics control-loop harnesses: research result for run 14

Research only. Nothing is published on use-cases/real-time-control or the videos page this run. Run 14 rechecks every pin before any page.

Condition met: two third-party harnesses that call Jev were read at pinned commits with every failure cell recorded. A third calls a local substitute model by default. All three are third-party runnable demos; none is a product and none is published by TypeSafe.

Harnesses read at pinned commits (11 Oct 2026, 05:17–05:25 UTC)
HarnessTypeModel at the pinOn a Jev error or timeoutTimeoutNo API key
mikespins/doom-jev at bd4e98fGame (ViZDoom)JevKeeps the last action (brain.py L331–334)2.0 s, no retryExits before starting
FazalAAli/jev-robotics-demo at 531de61Simulation (MuJoCo)JevRaises an error and stops (no error handling in the file)SDK default: 10 s, 2 retries (typesafe-sdk 0.7.x)Raises when the client is created
ys118/ra2web-jev-player at af8cacaGame (Red Alert 2 in a browser)A local substitute model by default; Jev only if JEV_BASE_URL is setRetries, then makes no decision; ordinary code keeps playing40 s, 4 attemptsNo decision (inferred)

8. Drift and spare-time results

  • pydantic-ai-slim 2.55.0 (compared from the PyPI wheels; R13-S119, R13-S120): checked at 2.55.0, failure path unchanged; requests with images now fail with UserError before any network call.
  • @effect/ai-typesafe 4.0.3 (patch, 10 Oct) and @openclaw/typesafe 2026.10.1 (10 Oct): drift notes. Everything else on the drift list is unchanged since 10 Oct (registry check 05:22:26 UTC; R13-S35).
  • SDK default timeouts: typesafe-sdk 0.7.4: 10.0 s per request, 2 retries, a 30 s total retry budget; a timeout raises a subclass of TypeSafeError; a missing key raises when the client is created (R13-S300 to R13-S305). @typesafe-ai/sdk 0.6.0: 10,000 ms per attempt, 2 retries (R13-S310).
  • LanceDB TypeSafe reranker with no key at b1b080b: no error at setup, an error at the first rerank call, no fallback (R13-S311, R13-S312).

9. Evidence gaps and uncertainties

  • B39's significance test was not found in the files read. B38's full call logs are not committed.
  • classifier-bakeoff's metric code was not opened; decisionbench's results location is still unknown; the 10 new comparison repositories were not opened.
  • OpenClaw 2026.10.1 was not read past L650.
  • hermes-jev-mana's client key handling and the Hermes version it patches were not read, so its no-key cell is "not recorded".
  • jev-for-all's Claude Code hook is not in the published package, and OpenCode's handling of removed tools was not read. The Claude Code engine is not public source.
  • What the two status-page hostnames are for is not documented by TypeSafe.
  • Legal "Last updated" dates were not re-extracted.
  • Reddit returned HTTP 403. X, YouTube, TikTok and GitHub trending were not covered: their coverage of the scan window is unknown, not zero.
  • "jev vision", "jev vlm" and "jev vla" appear in search suggestions with no TypeSafe source: unexplained.

10. Not done (for the next run)

  • decisionbench's lead line.
  • ra2web's threshold values in doctrine.py.
  • typesafe-spring-ai on Maven.
  • Spare-time queue items 3–5: pg_jevplanner, the unchecked alternatives list plus TOD's base licence, and jev-graphrag-poc.
  • Products recount (jev-triage, jev-for-all and hermes-jev-mana are code-backed integrations to add if missing).

Method notes and what was not verified

  • Nothing was installed or run. No benchmark was run and no Jev or OpenAI call was made. How any of the 74 implementations behaves in real use is not known.
  • Study results are the operators' own measurements Reported. The pool compares no substitute with Jev itself.
  • GitHub quota (unauthenticated): core 60 of 60 and search 10 of 10 at 05:16:24 UTC. Calls made: 9 core and 4 search, within a budget of 35 core and 8 search. The git fetch fallback was not used.
  • Third-party source files were read from raw GitHub or published package files and deleted after reading; their URLs and line ranges are in the source ledger.
  • Audience research ran on 11 Oct; it does not change this release's pages.

Key sources and check times (11 Oct 2026, UTC)

Sources by section (R13 IDs)
  • Access, 05:17–05:18: status page status.typesafe.ai (R13-S05); docs.typesafe.ai/models.md (R13-S06); docs.typesafe.ai/llms.txt (R13-S07); quickstart docs.typesafe.ai/introduction/quickstart.md (R13-S04); homepage typesafe.ai (R13-S08); blog typesafe.ai/blog (R13-S09); lookalike sites (R13-S15 to R13-S17) and Eye Security report (R13-S18).
  • Benchmarks, 05:22:50–05:26: B38 at f81c004: experiments/stats.py L9–60 (R13-S52); results/summary.md L9, L52–53, L70, L100 (R13-S53); tree and README (R13-S50, R13-S51, R13-S54). B39 at 160e112: bench/metrics.py L4–41, L94–118 (R13-S57); results/2026-10-07/report/summary.md (R13-S58); dataset licence (R13-S59); tree, README and adapters (R13-S55, R13-S56, R13-S60). B40 at fecba75: eval.py L71–76, L236–258 (R13-S63); significance.py (R13-S64); results/significance_run_sept25.txt L16–17 (R13-S65); tree, README and reranker (R13-S61, R13-S62, R13-S66). classifier-bakeoff report (R13-S67). New comparison leads: GitHub repository search (R13-S23).
  • Alternatives: Microsoft-Decision-1 announcement commandline.microsoft.com (R13-S32); HN item 50024913 (R13-S33).
  • Coding agents, 05:17–05:40: jev-triage at 6360e21: index.js L34, L39–45 (R13-S200 to R13-S202). OpenClaw at aa6008a: src/auto-reply/reply/dispatch-from-config.choose-route.ts L558–562, L618–642, L657, L663–692; src/plugins/hooks.ts L829–840; src/plugins/hook-runner-global.ts L22–30 (R13-S203 to R13-S207). @openclaw/typesafe 2026.10.1 at 9cf7fc2 up to L650 (R13-S208, R13-S209). hermes-jev-mana at 3daf639: patches/jev-decisions.patch L88–93, L428–429, L436–458, L1062–1078, L1396–1397; tools/jev_mana_install.py L334–337, L401–402; plugin/jev-mana/jevkit/decisions.py L282–287; plugin/jev-mana/decision_hooks.py L61–71 (R13-S210 to R13-S215). Hermes at f97608f: hermes_cli/config_defaults.py L20–22; hermes_cli/plugins_cmd.py L986–994 (R13-S216 to R13-S221). jev-for-all npm 0.7.1: bin/pick.ts L61–72; src/models.ts L20–75; src/tools.ts L95–133 (R13-S222 to R13-S224). Hoppin-Sorter/jev-router at bdb4721: hooks/register.tsx L484–497 (R13-S225 to R13-S228).
  • Control-loop harnesses, 05:17–05:25: doom-jev at bd4e98f: brain.py L291, L331–334; jev_doom.py L41–45 (R13-S68, R13-S100 to R13-S105). jev-robotics-demo at 531de61: jev_agent.py L213 (R13-S69, R13-S106 to R13-S110). ra2web-jev-player at af8caca: src/ra2web_jev_player/jev/client.py L27–33, L65, L110, L123–145; game.py L227–234 (R13-S111 to R13-S116). TypeSafe GitHub organisation and cookbooks (R13-S117, R13-S118).
  • Drift and spare time: registries, 05:22:26 (R13-S35); pydantic-ai-slim wheels 2.54.0 and 2.55.0, models/typesafe.py L219–220 (R13-S119, R13-S120); typesafe-sdk 0.7.4 and 0.7.1 wheels (R13-S300 to R13-S305); @typesafe-ai/sdk 0.6.0 (R13-S310); LanceDB python/python/lancedb/rerankers/typesafe.py at b1b080b (R13-S311, R13-S312).
  • Scan, window 10 Oct 2026, 05:11 UTC to 11 Oct 2026, 05:16 UTC: HN Algolia (R13-S22); GitHub repository searches (R13-S23 to R13-S25); npm (R13-S26); Hugging Face (R13-S27); DEV.to (R13-S28); Reddit, HTTP 403 (R13-S29); search suggestions (R13-S30, R13-S31). GitHub quota (R13-S02).

Search published pools, pages, reports, and evidence.