Shaduf.Research preview
Use casesFive implementations read at pinned commits and studies sorted by type, checked · search suggestions checked 05:16–05:17 UTC

Jev reranker and Jev for RAG (TypeSafe AI): does Jev improve reranking, and what happens to your passages when it fails

Jev reranker evidence by study type, and what 5 reranking and passage-filter implementations do with your passages when Jev fails. Checked 8 Oct 2026. When the pool checked search suggestions on , Bing and DuckDuckGo completed "jev rer", "jev rera", "jev rerank" and "jev reranking" to "jev reranker". Google completed "jev for r" and "jev for ra" to "jev for rag" at 05:17:19 UTC; Bing and DuckDuckGo did not. A suggestion shows that people search for something; it does not show how many. This page sorts the published studies by what they measure, and reads the source code of five rerankers and passage filters to show what reaches your prompt when the Jev call fails.

Key finding

The grade-A study measures relevance judgments, not reranking; both rerunnable graded pipeline studies come from parties with a stake. On a Jev error, of 5 implementations, 2 stop, 2 fall back, 1 passes unscreened passages.

Reportedstudies: each operator's own measurementDocumentedfailure cells: source at pinned commits, checked

The only grade-A study (iwhalen.com) measures how often Jev's relevance grades agree with human grades on TREC queries; it does not test a reranking pipeline. Both rerunnable graded pipeline studies are published by parties with a stake: the library's own author (Hugging Face blog) and a vector-database vendor that ships a Jev reranker (LanceDB).

Implementations: every reranker or passage filter the pool has read at a pinned commit (5 on 8 Oct 2026): the two published framework rows, LiteLLM's compaction guardrail (read 6 Oct), and two read on 8 Oct (jev-reranker, LanceDB). Studies: every row on benchmarks that tests reranking or relevance, plus two leads read on 8 Oct. Not a survey. The pool installed no package, ran no benchmark and made no Jev call.

Which failure behaviour do you need?

Open the case that matches your pipeline; opening one closes the others. Written from the cells in the implementation table; it is not advice beyond them. Counts on a Jev error: fail-open 1, fallback 2, fail-closed 2 (1 + 2 + 2 = 5 rows). Documented checked

A screen that never lets an unscreened passage throughSafety filter · fail-open 1 of 5

Of the 5 rows, 1 is a safety screen: Spring AI JevDocumentFilter (prompt injection and relevance). It is fail-open on a Jev error and on a malformed answer, and it keeps passages below its 0.70 injection threshold. None of the 5 rows is a fail-closed safety filter. If you need one, wrap the call so an error removes the passage or stops the request. jev-reranker and LanceDB use that shape for their relevance scores.

Better order, and the request must not be lostfallback 2 of 5

2 of 5 rows keep the pipeline running on a Jev error. LlamaIndex JevRerank returns the retriever's order. LiteLLM's compaction guardrail forwards the request uncompacted. Your prompt then holds the same passages it would have held without Jev.

An error rather than a silent fallbackfail-closed 2 of 5

2 of 5 rows raise on a Jev error or a malformed answer. jev-reranker raises after up to 8 retries, with a 180 s timeout per request. LanceDB TypeSafeReranker lets the TypeSafe SDK's error through. Your code must catch it, or the query fails.

Studies: what each one measures

Studies are graded only on benchmarks. This table shows each study's type, its grade there, one result as reported and a link to its benchmarks row. Nothing was rerun. Reported results in each operator's units; checked

Study, operator and interestStudy typeTask, n and baselineResult as reportedGrade and flagsBenchmarks row
iwhalen.com, "Pseudo-relevances with Jev"An individual's blog; no stake stated.relevance-judgment studyNot a reranking pipelineTREC DL 2023 judged pairs (22,327). Jev (jev-latest) and GPT-6-luna as relevance judges, compared with human gradesCohen's kappa against human grades "0.1907" (Jev) vs "0.2391" (GPT-6-luna). System ranking: "Jev and Luna are essentially tied"A public-datasingle-taskReranking studies: iwhalen
Hugging Face community blog, "Introducing jev-reranker"The library's author, writing about the library (COI).reranking-pipeline studyNanoBEIR-en NanoHotpotQA, 50 queries. Baseline: hybrid search onlynDCG@10: hybrid 0.833; rerank 0.969; relevance_rerank (threshold 0.2) "0.975", keeping 7.62 documents per queryB COIsmall-nno-baselinepublic-datasingle-taskRerunnable: code at d58594bReranking studies: Hugging Face
LanceDB, "How Jev Compares to Other Rerankers"A vector-database vendor that ships TypeSafeReranker (COI).reranking-pipeline studyGooAQ (20,000 queries, seed 0), BEIR NQ, HotpotQA, FiQA, SciDocs; 19 reranker configurations, including no rerankerGooAQ jev-relevance Hybrid@10 92.59, p50 169 ms. The default prompt "drops below no reranker on HotpotQA (hybrid 92.67 vs 96.19)"B COIpublic-dataProvisional lifted 8 Oct 2026: the benchmark script reranking/benchmark/bench.py at 4616259 was opened and lines up with the README. Grade unchanged. Per-query score files are not committed; a rerun needs the APIReranking studies: LanceDB
Elastic Search Labs, "Using Jev as a search reranker"A search vendor.reranking-pipeline studyESCI e-commerce search; a 250-query test reportedNot restated here; see row B14B public-datasingle-taskGrade B since 7 Oct. Not rerunnable at 250 queries: the notebook at 3e3d688 runs a 10-query sample onlyB14
Parallel, "Testing out Jev"A web-search API company that runs its own rerankers.reranking-pipeline claimSearch reranking; private data; n not stated"NDCG@10 of 0.7: comparable to internal system"D no nClaim only; no materialsB04
MindStudio, "JEV as a Steerable Reranker" (28 Sep)not readType not assignedNot readNone citedNot gradedNo ledger row

Ungraded candidates without a stated stake (found 8 Oct)

So far, every rerunnable graded pipeline study comes from a party with a stake. This run also read two candidates that have no stated stake. Neither is graded. The lead adds both to the benchmarks leads. Reported README results; checked

Candidate and operatorStudy typeTask, n and baselineResult as reportedStatus
mugunthank7/jev-pulseAn individual developer's demo; no product stake seen; stake not stated.reranking-pipeline studyAmazon ESCI search: 40 queries, 653 products. Jev (typesafe/jev-1.13 via OpenRouter) vs BM25 onlynDCG@10: BM25 0.660, Jev 0.772 ("+0.112, standard error ~0.037"). README only; run outputs not committedNot graded yet. Code, data and metric at pinned commit a1e796f. The runner proposed A (provisional) for the next benchmarks refresh, with small-n, weak-baseline and single-task flags; that grade is not applied on this page
anessbelbati/jev-rerank-benchStake not stated; no affiliation statement found.reranking-pipeline study (per README)14 datasets; Jev vs Cohere Rerank 4, ZeroEntropy zerank-2 and a chat-model baseline"Equal-dataset nDCG@10: Jev rubric 0.692, Cohere Pro 0.691, ZeroEntropy zerank-2 0.682"Not graded (lead). README read at fecba75; code not opened

Excluded: anweat/jev-websearch-eval tests "Bocha Jev" (bocha-jev-v1) on Bocha's hosted endpoint. It is not established to be TypeSafe's Jev, so it is out of scope.

Search boundary: one web search on 8 Oct 2026, 05:14 UTC, for items published since 15 Sep 2026, plus two leads from 7 Oct opened at pinned commits. Listed by that search but not read: emretheus/jev-rag-benchmark, denser-org/rerank-bench-jev (a RAG vendor; probable stake) and three blog posts (hevmind.com, tryopine.com, levelup.gitconnected.com). TypeSafe's "Jack & Jill & Jev" case study scores job candidates, not retrieved passages; it is not a retrieval study and is not in this table.

Five implementations: what reaches your prompt when Jev fails

Each cell comes from source at the commit shown; README text was not used. Library details (publisher, maturity, API) stay on framework integrations. Documented checked ; Spring AI and LlamaIndex cells as published on frameworks and when Jev fails

Implementation (pin)What Jev decidesJev error or timeoutMalformed answerBelow thresholdNo keyBypass
Spring AI TypeSafe JevDocumentFilter 0.4.0Tag v0.4.0 = 51993bb. Safety screen. Counted under FrameworksPer retrieved passage: injection and relevance screenfail-openUnscreened passages kept; 10 s per attempt, 2 retries, 30 s budgetfail-openPassage passes throughacts-anywayInjection < 0.70: passage keptfail-closedBlank key: startup exceptionnot recorded
LlamaIndex JevRerank, communityPyPI llama-index-postprocessor-jev 0.1.1; repository 421afda. Reranker. Counted under FrameworksRelevance score per node; reorders and keeps top_nfallback-retriever orderDefault; the package calls it "fail open". Cut to top_n; 2.5 sfallback-retriever orderno-thresholdOptional flag only; never dropsfail-closedValueError at constructionnot recorded
LiteLLM TypeSafe compaction guardrail 1.104.0Tag v1.104.0 = 7964577. Prompt filter for tool results, not a safety block. Newly counted under Frameworks (8 Oct)Per completed tool result: "is this result still needed for the current task?"fallback-uncompacted requestDefault unreachable_fallback fail_open forwards the request uncompacted; a fail_closed option raises HTTP 502; 30 sfallback-uncompacted requestSame handlerheldResults with noul < 0.2 are removed from the prompt (configurable); a missing answer keeps the resultfail-closedAt start-up: "TypeSafe guardrail requires an API key"recordeddefault_on: False: runs only where configured; small results and protected messages are never sent
jev-reranker 0.1.2 (PyPI)Tag v0.1.2 = d58594b. Reranker. Newly counted under Frameworks (8 Oct)Noul probability per passage; listwise mode batches several passages per request. Sorts, cuts below the threshold, then keeps top_kfail-closed429, 500, 502, 503, 504 and 529 retried up to max_retries=8, then raises; other statuses raise at once; 180.0 s per HTTP requestfail-closedResponseValidationError: invalid JSON, mismatched answer IDs, invalid noul, missing model or usageheldPassage cut: relevance_rerank (threshold 0.2) drops passages below 0.2; rerank (threshold 0.0) keeps all and reordersfail-closed"Set {api_key_env} or pass api_key to use Jev." at constructionrecordedNo call for an empty list or top_k == 0; one passage in pairwise mode gets 0.5 without a call; passages cut to 4,000 characters
LanceDB TypeSafeReranker v0.40.0Release v0.40.0 (7 Oct) = b1b080b. Reranker. Newly counted under Frameworks (8 Oct)Noul per result ("whether the document is relevant"); one request per result or up to batch_size per requestfail-closed"API and response-validation errors propagate"; no timeout set in this file (the SDK default was not read)fail-closedMismatched answer IDs or an invalid probability: ValueErrorno-thresholdSorts by score only; any cut comes from the query's limit, outside this file (not read)not recordedThe SDK reads TYPESAFE_API_KEY; its behaviour with no key was not readrecordedNull documents get score 0 without a request and stay in the list, sorted last
  • fail-closed or held (the item does not go ahead; a person, the host's prompt or, for two passage filters, the code's own cut decides)
  • fail-open or acts-anyway; a "conditional" cell that can act is counted fail-open
  • fallback-<what> (named in the cell)
  • no-decision: nothing returned; the caller must handle it
  • advisory (Jev never decides), or conditional where this page has no count
  • not recorded in the code read, or no-threshold (no confidence cut; the cell says which)

Colours follow the shared legend on when Jev fails. The "What Jev decides" and "recorded" bypass cells are descriptions, not verdicts, and are left uncoloured.

Counts across the 5 rows (below threshold is counted separately from errors):

  • Jev error or timeout: fail-closed 2/5 (jev-reranker, LanceDB), fallback 2/5 (LlamaIndex, LiteLLM), fail-open 1/5 (Spring AI).
  • Malformed answer: fail-closed 2/5, fallback 2/5, fail-open 1/5 (the same rows).
  • Below threshold: held 2/5 (LiteLLM, jev-reranker), acts-anyway 1/5 (Spring AI), no-threshold 2/5 (LlamaIndex, LanceDB).
  • No key: fail-closed 4/5, not recorded 1/5 (LanceDB).
  • Bypass: recorded 3/5 (LiteLLM, jev-reranker, LanceDB), not recorded 2/5 (Spring AI, LlamaIndex).

None of the 5 is a fail-closed safety filter. The one safety screen (Spring AI) is fail-open. The two rows that fail closed (jev-reranker, LanceDB) score relevance, not safety.

On when Jev fails, Spring AI and LlamaIndex stay in the Frameworks domain, where they were already counted. The three new rows (LiteLLM compaction, jev-reranker, LanceDB) are counted there under Frameworks, the nearest existing domain. No new domain is formed: only 2 of the new rows are rerankers, and LiteLLM compaction prunes tool results rather than retrieved passages. Counting LiteLLM compaction is a judgement: its code removes results from the prompt, but those results are tool outputs, not retrieved passages.

Timeouts and retries in general: errors and rate limits. What each route sends and keeps: data privacy. LiteLLM's complexity router is on model routing; only the compaction guardrail appears here.

TypeSafe's cookbooks: examples, not studies

TypeSafe publishes two cookbooks for this use. They are Documented vendor examples; the numbers in them are Reported by the vendor and are not graded. Their dates are not confirmed. Checked

  • "Re-ranking": a noul asks "Could this candidate passage be from the cited precedent?" for each candidate. The vendor reports "top-1 accuracy from 5% to 18%" and "top-10 accuracy from 38% to 62%".
  • "Classifying RAG passages": a second stage between retrieval and generation asks 4 questions per passage over 81 passages. Passages are excluded when injection is above 0.70 or relevance is below 0.45. No accuracy result is cited here.

What was not verified

  • LanceDB: the TypeSafe SDK's default timeout and its behaviour with no key; whether the query's limit cuts reranked results (outside the file read); the lancedb version used by the benchmark (requirements.txt not opened).
  • jev-reranker: the async path was not read separately (it goes through the same _rerank).
  • Spring AI and LlamaIndex bypass cells: not recorded, as published.
  • MindStudio's 28 Sep reranker post: not re-read and not graded. Elastic B14's result is not restated here.
  • jev-pulse: only metrics.ts and the README were quoted; run.ts was reported by a helper; results are README-only. anessbelbati/jev-rerank-bench: README only.
  • The "Classifying RAG passages" cookbook date (27 Aug 2026) came from a fetch summary and was not seen on the page.
  • "No graded pipeline study without a stake" holds for the benchmarks ledger as of 7 Oct plus this run's reads, within the search boundary above.
  • The pool ran nothing: no package installed or run, no benchmark run, no Jev call.

Sources and check times (6 to 8 Oct 2026, UTC)

Source list: R11-S IDs, pinned links, paths and line ranges
  • R11-S11: search suggestions from Google (client=firefox), Bing (osjson) and DuckDuckGo (ac), 8 Oct 2026, 05:16:53–05:17:21 UTC.
  • R10-S215, R10-S216, R10-S223: iwhalen.com post; code at 82d9a0f (main.py). Reported
  • R10-S217, R10-S218: Hugging Face blog post; code at d58594b (examples/eval.py, docs/eval.md). Reported
  • R10-S219, R10-S220, R11-S206 to R11-S208: LanceDB post; lancedb/research at 4616259: reranking/benchmark/bench.py L30, L38, L61–75, L106, L168–169, L290–302, L305–338 (opened 8 Oct, 05:13:50 UTC); README.md and results/README.md. Reported
  • R10-S213, R10-S214: Elastic B14 post and notebook at 3e3d688.
  • R3-S75: Parallel, "Testing out Jev" (B04).
  • R11-S209, R11-S210: mugunthank7/jev-pulse at a1e796f: README.md (Reported), apps/api/src/rerank/metrics.ts L5–11, apps/api/src/rerank/run.ts, data/esci-search.json; 05:14:10 UTC.
  • R11-S211: anweat/jev-websearch-eval at 38dde16 README.md (excluded); 05:14:10 UTC.
  • R11-S212, R11-S213: web search "Jev TypeSafe reranker evaluation nDCG independent benchmark RAG", 05:14:23 UTC, boundary since 15 Sep 2026; anessbelbati/jev-rerank-bench at fecba75 README.md, 05:14:40 UTC. Reported
  • R8-S84: Spring AI TypeSafe 0.4.0, tag v0.4.0 at 51993bb: typesafe-spring-ai/src/main/java/org/springaicommunity/typesafe/rag/JevDocumentFilter.java L118–132, L170–206.
  • R8-S86: LlamaIndex JevRerank at 421afda and PyPI 0.1.1: packages/llama-index-postprocessor-jev/llama_index/postprocessor/jev/base.py L69–83, L103, L145–151, L266–272; utils.py L92–96.
  • R9-S120, R11-S48: LiteLLM 1.104.0, tag v1.104.0 at 7964577: litellm/proxy/guardrails/guardrail_hooks/typesafe/typesafe.py L1–7, L48, L52–55, L162, L167–171, L181–183, L193–203, L278–299, L302–316, L347–371 (read 6 Oct 2026).
  • R11-S200 to R11-S203, R11-S49: jev-reranker 0.1.2 (PyPI), tag v0.1.2 at d58594b: src/jev_reranker/reranker.py L126–128, L233–235, L349–376, L507–517, L546–552, L716–737, L853–854; src/jev_reranker/_client.py L44–67, L100–114, L132–136; src/jev_reranker/_ranking.py; 05:13:20 UTC, runner spot-check 05:19:45 UTC.
  • R11-S204, R11-S205, R11-S49: LanceDB release v0.40.0 (tag object 2fd3239) at b1b080b: python/python/lancedb/rerankers/typesafe.py L43–46, L67–69, L125–127, L158–177, L186, L199–203; 05:13:40 UTC, runner spot-check 05:19:45 UTC.
  • R11-S214 to R11-S216: TypeSafe cookbooks re-ranking, classifying RAG passages and the cookbook index, 05:14:50–05:14:55 UTC.
  • R11-S56: TypeSafe case study (not a retrieval study), 05:14:53 UTC.

Search published pools, pages, reports, and evidence.