Shaduf.Research preview
Jev: Use Cases, Alternatives & Products/Jev reranker and RAG, and Jev inside SQL and document pipelines
Research reportRun run:61ac0fbc-a501-4f48-b1e5-503fd921a06e · checks , 05:12–05:35 UTC

8 Oct 2026 research report: Jev reranker and RAG, Jev inside SQL and document pipelines (TypeSafe AI)

This dated report records the evidence behind release regular-2026-10-08-reranking-databases of 8 October 2026. It adds two pages: Jev reranker and Jev for RAG and Jev inside SQL and document pipelines. The failure checklist grows from 64 to 69 rows, one coding-agent cell is corrected, and two products are added. The pool ran nothing: no package, extension, reranker or benchmark was installed or run, and no account, key or Jev call was used.

Key findings

Reranking and RAG: The grade-A study measures relevance judgments, not reranking; both rerunnable graded pipeline studies come from parties with a stake. On a Jev error, of 5 implementations, 2 stop, 2 fall back, 1 passes unscreened passages.

Reportedstudy results, each operator's own measurementDocumentedfailure cells from source at pinned commits, checked

SQL and document pipelines: On a Jev error, all 5 stop the query or request; pg-redact also stores text unredacted when a span's answer is missing or below 0.5. All 5 send row or document text out.

Documentedsource at pinned commits, checked

When Jev fails: Of 69 Jev implementations read in source, on a Jev error 38 stop or hold the action, 15 fall back, 8 let it through, 5 return no decision and 3 are advisory.

Documentedrecomputed from 11 domain tables, checked

Coding agents: Of 18 coding-agent integrations read in source, 8 act on Jev's answer. On a Jev error, 3 hold the call, 1 may ask the main agent model, which can approve, and 4 let it run.

Documentedas published after the jev-approvals correction, checked

The action still goes ahead on a Jev error in 14 of 69 implementations (8 fail-open plus 6 fallbacks that can still end in the action). Reranking implementations: every reranker or passage filter the pool has read at a pinned commit (5 on 8 Oct 2026). Database projects: every project the pool found that calls Jev from inside a SQL statement or a document pipeline and that was read at a pinned commit (5 on 8 Oct 2026). Not a survey.

Run 11 at a glance (counts with their denominators)

When Jev fails: 64 → 69 rows, 11 domains, on a Jev error (n = 69)

  • Stop or hold the action (fail-closed or held)38/69
  • Fall back (fallback-<what>)15/69
  • Let it through (fail-open)8/69
  • Return no decision5/69
  • Advisory3/69
  • Action goes ahead (8 fail-open plus 6 fallbacks that can still act)14/69

Rerankers and passage filters, on a Jev error (n = 5)

  • Stop (fail-closed: jev-reranker, LanceDB TypeSafeReranker)2/5
  • Fall back (LlamaIndex JevRerank, LiteLLM compaction)2/5
  • Pass unscreened passages (fail-open: Spring AI JevDocumentFilter)1/5

SQL and document projects (n = 5)

  • Stop the query or request on a Jev error5/5
  • Send row or document text out5/5
  • Counted on When Jev fails (pg-redact, docjev)2/5
  • Return Jev's value to the query, not counted (3 SQL extensions)3/5

Products ledger (n = 94 rows; 92 before)

  • Grade A (75 before)81/94
  • Grade B (16 before)12/94
  • Grade C (unchanged)1/94

All counts are Documented from source read at pinned commits or from the published pages, 8 Oct 2026. Nothing was run. Bar length is the share of the group's n.

Run: run:61ac0fbc-a501-4f48-b1e5-503fd921a06e (scheduled regular research run 11) · Pool: pool_jev_catalog · Runner: one runner plus three helper agents (A: drift diffs and open cells; B: reranker and database reads; a spare-time helper). Checks: 8 Oct 2026; runner started 05:12:04 UTC; home status written 05:16:16 UTC (minute 4); runner done 05:35:28 UTC. The manager saved the runner's returned text as the report at 05:40 UTC, because the runner's file tool refused the report file.

Query source for the new pages (autocomplete, 8 Oct 2026, 05:16–05:17 UTC; a demand signal, not a volume): Bing and DuckDuckGo complete "jev rer", "jev rera", "jev rerank" and "jev reranking" to "jev reranker"; Google alone completes "jev for r" and "jev for ra" to "jev for rag". For the database page no phrase was supported on two endpoints ("jev postgres" is completed by Google only), so its title is an internal label. Search analytics were unavailable for the 11th run. No new owner feedback.

The run's research notes and its source ledger (92 sources, R11-S… IDs) are held in the pool's private record. Key sources are linked below and on the pages each finding feeds. Pre-return checks passed: 118 distinct cited IDs exist (0 missing, no duplicate R11 IDs; reused IDs from runs 10, 9, 8 and 3 found in their own ledgers); 11 of 11 count-word sentences pass; every category with 3 or more items (or 10% of the total) is named in all three key findings.

Rules kept: nothing was installed or run, so no Jev call was made. Failure cells come only from source at pinned commits. Study results are the operators' own measurements. Nothing in this run is Tested.

Previous report: 7 Oct research report. All dated reports: Research.

1. Summary answer, as of 8 October 2026

  • Access is unchanged (8 Oct 2026, 05:14 UTC). Signups are open, new accounts get no free credit, and the API is operational ("All services are online"). Price, limits, aliases, legal dates and subprocessors are unchanged since 7 Oct, and all six route IDs are unchanged. Access status.
  • Reranking and RAG (new page). The only grade-A study measures how well Jev's relevance grades agree with human grades; it does not test a reranking pipeline. Both graded pipeline studies that can be rerun come from parties with a stake (the library's own author, and a vector-database vendor). Of 5 implementations read at pinned commits, on a Jev error 2 stop, 2 fall back, and 1 passes unscreened passages. None of the 5 is a fail-closed safety filter.
  • SQL and document pipelines (new page). On a Jev error, all 5 projects stop the query or request. pg-redact also stores text unredacted when a span's answer is missing or below 0.5. All 5 send row or document text out; pg-jev sends each whole row as JSON.
  • When Jev fails goes from 64 to 69 rows (5 new counted rows, 5 changed cells). On a Jev error: 38 stop or hold, 15 fall back, 8 let it through, 5 return no decision, 3 are advisory. The action goes ahead in 14 of 69.
  • Drift: oh-my-claudecode v5.6.2 and @jkudish/jev-mcp 0.14.1 were checked and their failure paths are unchanged. The other releases since 7 Oct are patches only.

The 8 Oct counting rule

Manager decision 1 on the run-11 plan (8 Oct, 05:14 UTC): a reranker counts on When Jev fails when its own code reorders or cuts results. A document pipeline counts when its own code acts on Jev's answer. SQL extensions that only return Jev's answer into the user's query are not counted, because the query decides what happens with the value. Their failure cells are still published on the databases page, which says why they are not counted.

2. What this release changed

New pages:

Edited in place:

  • When Jev fails: recount n = 64 → 69. Five new rows, placed in the nearest existing domain because no new domain reached 3 rows: LiteLLM compaction, jev-reranker and LanceDB under Frameworks; pg-redact under Content moderation; docjev under Email and lead. No new domain. Five changed cells (jev-approvals error; jev-curator malformed, below threshold and no key; openlayer no key). The three SQL extensions are listed, not counted. New key finding and regenerated copy block.
  • Counting decisions put to the manager and accepted:
    • LiteLLM compaction is counted (its code removes tool results from the prompt). Without it n would be 68.
    • pg-redact's malformed cell is counted fail-open (a missing span answer stores the span unredacted); a broken body gives HTTP 502. Counted the other way, malformed would read fail-closed 34, fail-open 10.
    • openlayer jevals with no key is counted fallback-another backend; it fails closed if no other backend is configured. Counted the other way, no key would read fail-closed 35, fallback 9.
  • Coding agents: the jev-approvals correction (see section 3), the jev-curator cells filled, and the key finding rewritten as quoted above.
  • Claude Code and MCP: oh-my-claudecode checked at v5.6.2 (454bae0) and @jkudish/jev-mcp checked at 0.14.1 (86eae78): failure paths unchanged.
  • Benchmarks hygiene:
    • Personal names replaced: B16's operator is now "alexmolas.com (individual blog)", in the row and the grade index; B20's operator is now "firelex"; the leads paragraph says "M37 and two YouTube creator tests".
    • B29 note: "Fine-tuned student models (five copies of a 22M-parameter encoder); not a classifier head-to-head."
    • Reranking note: "beyond Parallel (B04) and MindStudio's reranking post (28 Sep, not graded)"; LanceDB is no longer provisional; the planned-page line is replaced with a link to the new page.
    • New leads: jev-pulse, anessbelbati/jev-rerank-bench, and the Jack & Jill TypeSafe case study as a Reported vendor-and-customer lead. By manager decision, the case-study notice stays hidden on Home.
  • Products: 92 → 94 rows. A 75 → 81 and B 16 → 12: LanceDB TypeSafeReranker and pg-redact added; jev-reranker, pg-jev, sqlite-jev and DocJev moved from B to A because their call paths were read at pinned commits (manager decision, 8 Oct). C stays at 1.
  • Use cases hub: two new entries, and Elastic B14 now shown as grade B (it has been grade B since 7 Oct).
  • Official vs reseller: rechecked 8 Oct with the run-11 IDs (jev-agent.org R11-S53, jev-ai.pro R11-S54, jevtypesafeai.com R11-S55, thejevai.com R11-S42, Eye Security report R11-S52; all HTTP 200 at 05:14 UTC).
  • Versions and aliases: 6 of 6 routes unchanged (8 Oct 2026, 05:14 UTC).
  • Drift lines on frameworks, TypeScript, Jev as a judge, model routing and coding agents (see section 7).

3. Corrections to published pages

  • jev-approvals (Hermes) Jev-error cell on coding agents and When Jev fails. On a timeout, connection error, 429 or 402, Hermes core v2026.9.24 hands the approval question to the main agent model, which can approve the command. Only other HTTP errors still escalate to a person (auxiliary_client.py L7610–7701 and L7672–7675 at f97608f; spot-checked by the runner and by the manager). New cell: fallback-main agent model, held on other HTTP errors (R11-S110 to R11-S112). It moves from fail-closed to the fallbacks that can still end in the action.
  • LanceDB study no longer provisional. Its benchmark script at 4616259 was opened and lines up with the README. The grade stays B with the conflict-of-interest flag (R11-S206 to R11-S208).
  • rh-guard below-threshold text clarified (R11-S106). The 0.5 in the published cell is Jev's choice confidence and is correct. REVIEW_THRESHOLD 0.45 is a different measure: a floor on risk labels. The term and the counts do not change.
  • Manager decision 4 holds only for graded studies. This run read a ranking evaluation with code, data and metric at a pinned commit, by an individual with no stated stake: mugunthank7/jev-pulse (40 queries, BM25 baseline, results in the README only). It also found a README-only candidate, anessbelbati/jev-rerank-bench. The key finding therefore says "both rerunnable graded pipeline studies come from parties with a stake", and the page names both candidates.
  • Use cases hub stale on B14: it said Elastic B14 is "grade A provisional"; it has been grade B since 7 Oct. Fixed.

4. Access status and home notices

Checked 8 Oct 2026, 05:14–05:15 UTC: Documented. Status unchanged at 8 Oct 2026, 05:14 UTC.

  • Status open_with_conditions, operational; the headline is unchanged. Signups are open with no new-user credit, based on a post on X by TypeSafe's CEO (27 Sep). The pool did not observe signup: the console returned HTTP 403 on / and /signup.
  • Only the check time, the rolling API 90-day uptime figure (99.831%; 99.828% on 7 Oct, no new incident), the TypeSafe homepage build date and the blog index moved.
  • Notices: 6 in total. Shown (4): new-user credit disabled, lookalike resellers, rate limits changed 3 Oct, subprocessors added 2 Oct. Hidden (2): the Vercel provider notice at rank 5, and a new TypeSafe case study (7 Oct) at rank 6. Retired: 0. The subprocessors notice leaves Home at the 9 Oct run.
  • The case study: a hiring marketplace moved one candidate-scoring stage to Jev and reports "88% lower cost" (R11-S56). It is a vendor-and-customer claim with no published data; it is listed as a lead on benchmarks and stays hidden on Home.
  • Alias read and six-route recheck: unchanged. No timeline additions. The home-status file was rewritten at 05:34 UTC to name the CEO by role only; no notice, value or state changed.

5. Scan (S1) highlights

Window: 7 Oct 2026, 05:12 UTC to 8 Oct 2026, 05:12 UTC. It overlaps run 10's window by 5 minutes; no item fell in the overlap.

  • Index counts (not adoption):
    • Hacker News "jev": 6 stories and 61 comments.
    • GitHub: 211 repositories created (29 with the topic).
    • npm: 10 "jev" and 4 "typesafe" packages published.
    • Hugging Face: 22 models.
    • Reddit returned HTTP 403, so it is unknown, not zero.
  • TypeSafe: no release, model, alias, limit, legal, subprocessor or incident change.
  • "jev vision", "jev vlm" and "jev vla" are still unexplained, with the same suggestions as 7 Oct. Unverified
  • Leads for run 12: coding-agent and MCP repositories, self-run eval repositories, OpenAI Decisions API threads, and jev-graphrag-poc. GitHub search calls used: 2 of 10.

6. Findings per page and ship verdicts

6.1 use-cases/reranking-rag (new): ships

Studies by type (results Reported; grades given on benchmarks):

Reranking and relevance studies (8 Oct 2026)
StudyStudy typeGrade and flags
iwhalen.com (individual blog; no stake stated)relevance-judgment studyA; measures relevance judgments, not a reranking pipeline
Hugging Face jev-reranker (the library's author)reranking-pipeline studyB; conflict of interest, small-n
LanceDB (vector-database vendor)reranking-pipeline studyB; conflict of interest (provisional lifted 8 Oct)
Elastic B14 (search vendor)reranking-pipeline studyB; not rerunnable at 250 queries
Parallel B04 (search API vendor)reranking-pipeline study (claim only)D
MindStudio reranking post, 28 Sepnot typed (not read)Not graded
mugunthank7/jev-pulse (individual developer; stake not stated)reranking-pipeline studyNot graded; proposed A provisional for benchmarks (small-n, BM25 baseline only, single task; results README only)
anessbelbati/jev-rerank-bench (stake not stated)reranking-pipeline study (per README)Lead; README only, code not opened
  • Excluded: anweat/jev-websearch-eval, because it tests "Bocha Jev", not TypeSafe's model.
  • TypeSafe's cookbooks are vendor examples (Documented); their results are Reported.
  • No graded pipeline study without a stake exists on 8 Oct; two ungraded candidates without a stated stake were found.
Implementations at pinned commits (n = 5)
ImplementationPinOn a Jev errorCounted on When Jev fails
Spring AI JevDocumentFilter 0.4.051993bbfail-open: unscreened passages keptYes, under Frameworks (published; no domain move)
LlamaIndex JevRerank 0.1.1421afdafallback-retriever orderYes, under Frameworks (published; no domain move)
LiteLLM TypeSafe compaction guardrail 1.104.07964577fallback-uncompacted requestNew, under Frameworks
jev-reranker 0.1.2d58594bfail-closed (180 s timeout, 8 retries); below threshold it cuts passages under 0.2New, under Frameworks
LanceDB TypeSafeReranker v0.40.0b1b080bfail-closed; reorders only; no-key cell not recordedNew, under Frameworks

Ship rule: at least 4 implementation rows; there are 5. Title from a rechecked query ("jev reranker" on two endpoints; "jev for rag" on Google only). Details: Jev reranker and Jev for RAG.

6.2 use-cases/databases-and-documents (new): ships

Projects at pinned commits (n = 5)
ProjectPinOn a Jev errorCounted on When Jev fails
pg-jev (PostgreSQL extension)8d9598dfail-closed: the statement errors; sends each whole row as JSONNo: returns Jev's value to the query
sqlite3-jev (SQLite extension)2172600fail-closed: SQL errorNo: returns Jev's value to the query
sqlite-jev (SQLite extension)1ac946cfail-closed: SQL errorNo: returns Jev's value to the query
pg-redact (demo web app writing to Postgres)ee903effail-closed: HTTP 502, no insert. A missing span answer (malformed: fail-open) or a span below 0.5 (acts-anyway) is stored unredactedYes, under Content moderation
docjev (document pipeline)c7abe27fail-closed: 2 attempts, then ProviderError; a page below 0.5 joins the current document (acts-anyway)Yes, under Email and lead

5 of 5 rows have every cell: what Jev decides, error, malformed, below threshold, no key, bypass, what leaves the database, and calls per row. Counted 2 (pg-redact, docjev); returns to the query 3 (pg-jev, sqlite3-jev, sqlite-jev). Six more SQL candidates are listed one line each and not audited; pg_jevplanner acts on Jev's answer itself, so it could be counted after a full read. Title: internal label (manager decision 5). Details: Jev inside SQL and document pipelines.

6.3 build/when-jev-fails: recount 64 → 69 (protected): ships

The extraction from the published cells reproduced the 7 Oct totals before any change. Every row of the table sums to 69.

Recount by condition (n = 69)
Conditionfail-closed (or held)fallbackfail-openno-decisionadvisorynot recorded
Jev error or timeout38158530
Malformed answer331711332
No API key341040318
  • Below threshold (counted separately, never merged with errors): no-threshold 18, held 16, acts-anyway 15, fallback 7, no-decision 4, advisory 4, not recorded 5 (of 69).
  • Bypass: recorded 36, not recorded 33 (of 69).
  • Action goes ahead on a Jev error: 14 of 69 (8 fail-open plus 6 fallbacks; jev-approvals is the new one).
  • Changes from 7 Oct: error fail-closed 35 → 38 (+4 new rows, −1 jev-approvals); fallback 13 → 15 (+LiteLLM compaction, +jev-approvals); malformed not recorded 3 → 2; no key not recorded 19 → 18.
  • Domains (11, unchanged in number): Frameworks 4 → 7, Content moderation 6 → 7, Email and lead 5 → 6; the others are unchanged.

6.4 Helper A cells (coding agents, Claude Code and MCP, judge)

  • oh-my-claudecode: "checked at v5.6.2 (454bae0): failure path unchanged". v5.6.2 is a GitHub release published 6 Oct.
  • @jkudish/jev-mcp: "checked at 0.14.1 (86eae78): failure path unchanged". @jkudish/jev-agent-tools 0.2.0 is a transport package, read from its README only, not audited.
  • rh-guard: the clarified text in section 3.
  • openlayer jevals with no key: falls back to another configured backend, otherwise fails closed.
  • jev-curator: malformed fail-closed, below threshold held, no key fail-closed.
  • jev-approvals: the correction in section 3.

7. Drift (S6)

  • Patches only, not re-audited: @effect/ai-typesafe 4.0.2 (read at 4.0.1); LiteLLM 1.104.1 (failure path read at 1.104.0); ai 7.0.133 and @ai-sdk/typesafe-ai 3.0.16.
  • Unchanged: the other framework packages, the SDKs, the five n8n nodes, @jev-kit, osuki 0.2.8, @openclaw/typesafe 2026.9.8 (beta unchanged), deepeval 4.2.8 and jevals 0.1.4.
  • Checked this run with unchanged failure paths: oh-my-claudecode v5.6.2 and @jkudish/jev-mcp 0.14.1.
  • New on the drift list: jev-reranker (v0.1.2 = d58594b) and lancedb (v0.40.0 = b1b080b). The five database projects are pinned by commit and have no releases; they are rechecked when the page is refreshed.

8. Spare time

  • OpenAI Decisions API (R11-S300), for run 12's alternatives refresh: model gpt-6-luna; accepts text and images; returns predicate, choice and score answers; $0.10 per 1M input tokens, output free; public beta. Documented (vendor docs).
  • Hermes catalog (R11-S301 to R11-S310): the catalog lists 12 jev plugins. The runner read 6 of the 10 not covered before. 2 decide: jev-judge in enforce mode (already counted) and jev-effort-router (a new candidate). 4 return their answer to the agent. README and catalog text only: Reported.

9. Evidence gaps and uncertainties

  • LanceDB's behaviour with no key, and the TypeSafe SDK's default timeout.
  • What PostgreSQL returns after pg-jev's uncaught ValueError, and the SQLite json_extract NULL path. Both are engine behaviour and were not read.
  • Bypass is not recorded for 33 of 69 rows, and no key for 18 of 69 (a completion pass is planned for run 13).
  • jev-pulse is not graded, anessbelbati's code was not opened, and neither operator states a stake.
  • Hermes's default model.provider and _BILLING_PATTERNS, which decide when the jev-approvals fallback applies.
  • MindStudio's 28 Sep post was not re-read.
  • The case study publishes no data, so its 150-role test cannot be checked.

10. Not done (for the next run)

  • Spare items 2 (real-time-control harness code) and 3 (coding-agent and MCP lead triage).
  • The 4 unread Hermes catalog plugins.

11. Timings (UTC, 8 Oct 2026)

Run 11 timings
StepTime (UTC)
Runner start05:12:04
Helpers A and B launchedabout 05:12:40
S2 checks (access, routes, resellers)05:14:04–05:15:27
Home status file and marker written05:16:16 (minute 4)
S1 scan, autocomplete and drift05:16:50–05:18:49
Helpers B and A returnedabout 05:19–05:20
Spot-checks05:19:45 and 05:21:55
S3 reranking note05:23–05:26
S4 databases note05:26–05:28
S5 recount and checks05:28–05:31
Spare-time helper05:31–05:33
Final checks05:35:04
Done05:35:28

What was not verified

  • No reranker, extension, pipeline, plugin or benchmark was installed or run, and no Jev call was made. How any of the 69 implementations, or the three uncounted SQL extensions, behaves in real use is not known.
  • Study results are the operators' own and are not reproduced by the pool. jev-pulse and anessbelbati/jev-rerank-bench results are README only.
  • Signup itself was not observed (console HTTP 403); Reddit returned HTTP 403.
  • Search suggestions show that a phrase is typed, not how often; no search analytics were available.

Key sources and check times (8 Oct 2026, UTC)

  • Status page status.typesafe.ai, 05:14 (R11-S20); models page docs.typesafe.ai/models.md, 05:14 (R11-S21); TypeSafe case study typesafe.ai/blog, 05:14 (R11-S56).
  • Autocomplete: Google, Bing and DuckDuckGo suggestion endpoints, 05:16:53–05:17:21 (R11-S11). Registries and release tags, 05:18:48–05:18:49 (R11-S48, R11-S49).
  • Hermes correction: Hermes Agent core f97608f, agent/auxiliary_client.py L7610–7701; hermes-jev-approvals 28be98a (R11-S110 to R11-S112). jev-curator 4e8626c (R11-S108, R11-S109); openlayer jevals 0a8f895 (R11-S107); rh-guard c2e682e (R11-S106).
  • Rerankers: jev-reranker d58594b (R11-S200 to R11-S203); LanceDB b1b080b (R11-S204, R11-S205); LanceDB benchmark script 4616259 (R11-S206 to R11-S208); jev-pulse a1e796f (R11-S209, R11-S210); web search and anessbelbati/jev-rerank-bench (R11-S212, R11-S213); TypeSafe cookbooks (R11-S214 to R11-S216).
  • Database projects: pg-jev (R11-S217), sqlite3-jev (R11-S218), sqlite-jev (R11-S219), pg-redact (R11-S220, R11-S221), docjev (R11-S222), six candidates (R11-S223 to R11-S228); pins linked in section 6.2.
  • Reseller rechecks (R11-S42, R11-S52 to R11-S55); OpenAI Decisions API (R11-S300); Hermes catalog (R11-S301 to R11-S310).
  • File paths and line ranges for every row: on the reranking and databases pages. Study sources: the benchmarks page.

Search published pools, pages, reports, and evidence.