Jev reranker and Jev for RAG (TypeSafe AI): does Jev improve reranking, and what happens to your passages when it fails
Jev reranker evidence by study type, and what 5 reranking and passage-filter implementations do with your passages when Jev fails. Checked 8 Oct 2026. When the pool checked search suggestions on , Bing and DuckDuckGo completed "jev rer", "jev rera", "jev rerank" and "jev reranking" to "jev reranker". Google completed "jev for r" and "jev for ra" to "jev for rag" at 05:17:19 UTC; Bing and DuckDuckGo did not. A suggestion shows that people search for something; it does not show how many. This page sorts the published studies by what they measure, and reads the source code of five rerankers and passage filters to show what reaches your prompt when the Jev call fails.
The grade-A study measures relevance judgments, not reranking; both rerunnable graded pipeline studies come from parties with a stake. On a Jev error, of 5 implementations, 2 stop, 2 fall back, 1 passes unscreened passages.
The only grade-A study (iwhalen.com) measures how often Jev's relevance grades agree with human grades on TREC queries; it does not test a reranking pipeline. Both rerunnable graded pipeline studies are published by parties with a stake: the library's own author (Hugging Face blog) and a vector-database vendor that ships a Jev reranker (LanceDB).
Implementations: every reranker or passage filter the pool has read at a pinned commit (5 on 8 Oct 2026): the two published framework rows, LiteLLM's compaction guardrail (read 6 Oct), and two read on 8 Oct (jev-reranker, LanceDB). Studies: every row on benchmarks that tests reranking or relevance, plus two leads read on 8 Oct. Not a survey. The pool installed no package, ran no benchmark and made no Jev call.
Which failure behaviour do you need?
Open the case that matches your pipeline; opening one closes the others. Written from the cells in the implementation table; it is not advice beyond them. Counts on a Jev error: fail-open 1, fallback 2, fail-closed 2 (1 + 2 + 2 = 5 rows). Documented checked
A screen that never lets an unscreened passage throughSafety filter · fail-open 1 of 5
Of the 5 rows, 1 is a safety screen: Spring AI JevDocumentFilter (prompt injection and relevance). It is fail-open on a Jev error and on a malformed answer, and it keeps passages below its 0.70 injection threshold. None of the 5 rows is a fail-closed safety filter. If you need one, wrap the call so an error removes the passage or stops the request. jev-reranker and LanceDB use that shape for their relevance scores.
Better order, and the request must not be lostfallback 2 of 5
2 of 5 rows keep the pipeline running on a Jev error. LlamaIndex JevRerank returns the retriever's order. LiteLLM's compaction guardrail forwards the request uncompacted. Your prompt then holds the same passages it would have held without Jev.
An error rather than a silent fallbackfail-closed 2 of 5
2 of 5 rows raise on a Jev error or a malformed answer. jev-reranker raises after up to 8 retries, with a 180 s timeout per request. LanceDB TypeSafeReranker lets the TypeSafe SDK's error through. Your code must catch it, or the query fails.
Studies: what each one measures
Studies are graded only on benchmarks. This table shows each study's type, its grade there, one result as reported and a link to its benchmarks row. Nothing was rerun. Reported results in each operator's units; checked
| Study, operator and interest | Study type | Task, n and baseline | Result as reported | Grade and flags | Benchmarks row |
|---|---|---|---|---|---|
| iwhalen.com, "Pseudo-relevances with Jev"An individual's blog; no stake stated. | relevance-judgment studyNot a reranking pipeline | TREC DL 2023 judged pairs (22,327). Jev (jev-latest) and GPT-6-luna as relevance judges, compared with human grades | Cohen's kappa against human grades "0.1907" (Jev) vs "0.2391" (GPT-6-luna). System ranking: "Jev and Luna are essentially tied" | A public-datasingle-task | Reranking studies: iwhalen |
| Hugging Face community blog, "Introducing jev-reranker"The library's author, writing about the library (COI). | reranking-pipeline study | NanoBEIR-en NanoHotpotQA, 50 queries. Baseline: hybrid search only | nDCG@10: hybrid 0.833; rerank 0.969; relevance_rerank (threshold 0.2) "0.975", keeping 7.62 documents per query | B COIsmall-nno-baselinepublic-datasingle-taskRerunnable: code at d58594b | Reranking studies: Hugging Face |
LanceDB, "How Jev Compares to Other Rerankers"A vector-database vendor that ships TypeSafeReranker (COI). | reranking-pipeline study | GooAQ (20,000 queries, seed 0), BEIR NQ, HotpotQA, FiQA, SciDocs; 19 reranker configurations, including no reranker | GooAQ jev-relevance Hybrid@10 92.59, p50 169 ms. The default prompt "drops below no reranker on HotpotQA (hybrid 92.67 vs 96.19)" | B COIpublic-dataProvisional lifted 8 Oct 2026: the benchmark script reranking/benchmark/bench.py at 4616259 was opened and lines up with the README. Grade unchanged. Per-query score files are not committed; a rerun needs the API | Reranking studies: LanceDB |
| Elastic Search Labs, "Using Jev as a search reranker"A search vendor. | reranking-pipeline study | ESCI e-commerce search; a 250-query test reported | Not restated here; see row B14 | B public-datasingle-taskGrade B since 7 Oct. Not rerunnable at 250 queries: the notebook at 3e3d688 runs a 10-query sample only | B14 |
| Parallel, "Testing out Jev"A web-search API company that runs its own rerankers. | reranking-pipeline claim | Search reranking; private data; n not stated | "NDCG@10 of 0.7: comparable to internal system" | D no nClaim only; no materials | B04 |
| MindStudio, "JEV as a Steerable Reranker" (28 Sep) | not readType not assigned | Not read | None cited | Not graded | No ledger row |
Ungraded candidates without a stated stake (found 8 Oct)
So far, every rerunnable graded pipeline study comes from a party with a stake. This run also read two candidates that have no stated stake. Neither is graded. The lead adds both to the benchmarks leads. Reported README results; checked
| Candidate and operator | Study type | Task, n and baseline | Result as reported | Status |
|---|---|---|---|---|
mugunthank7/jev-pulseAn individual developer's demo; no product stake seen; stake not stated. | reranking-pipeline study | Amazon ESCI search: 40 queries, 653 products. Jev (typesafe/jev-1.13 via OpenRouter) vs BM25 only | nDCG@10: BM25 0.660, Jev 0.772 ("+0.112, standard error ~0.037"). README only; run outputs not committed | Not graded yet. Code, data and metric at pinned commit a1e796f. The runner proposed A (provisional) for the next benchmarks refresh, with small-n, weak-baseline and single-task flags; that grade is not applied on this page |
anessbelbati/jev-rerank-benchStake not stated; no affiliation statement found. | reranking-pipeline study (per README) | 14 datasets; Jev vs Cohere Rerank 4, ZeroEntropy zerank-2 and a chat-model baseline | "Equal-dataset nDCG@10: Jev rubric 0.692, Cohere Pro 0.691, ZeroEntropy zerank-2 0.682" | Not graded (lead). README read at fecba75; code not opened |
Excluded: anweat/jev-websearch-eval tests "Bocha Jev" (bocha-jev-v1) on Bocha's hosted endpoint. It is not established to be TypeSafe's Jev, so it is out of scope.
Search boundary: one web search on 8 Oct 2026, 05:14 UTC, for items published since 15 Sep 2026, plus two leads from 7 Oct opened at pinned commits. Listed by that search but not read: emretheus/jev-rag-benchmark, denser-org/rerank-bench-jev (a RAG vendor; probable stake) and three blog posts (hevmind.com, tryopine.com, levelup.gitconnected.com). TypeSafe's "Jack & Jill & Jev" case study scores job candidates, not retrieved passages; it is not a retrieval study and is not in this table.
Five implementations: what reaches your prompt when Jev fails
Each cell comes from source at the commit shown; README text was not used. Library details (publisher, maturity, API) stay on framework integrations. Documented checked ; Spring AI and LlamaIndex cells as published on frameworks and when Jev fails
| Implementation (pin) | What Jev decides | Jev error or timeout | Malformed answer | Below threshold | No key | Bypass |
|---|---|---|---|---|---|---|
Spring AI TypeSafe JevDocumentFilter 0.4.0Tag v0.4.0 = 51993bb. Safety screen. Counted under Frameworks | Per retrieved passage: injection and relevance screen | fail-openUnscreened passages kept; 10 s per attempt, 2 retries, 30 s budget | fail-openPassage passes through | acts-anywayInjection < 0.70: passage kept | fail-closedBlank key: startup exception | not recorded |
LlamaIndex JevRerank, communityPyPI llama-index-postprocessor-jev 0.1.1; repository 421afda. Reranker. Counted under Frameworks | Relevance score per node; reorders and keeps top_n | fallback-retriever orderDefault; the package calls it "fail open". Cut to top_n; 2.5 s | fallback-retriever order | no-thresholdOptional flag only; never drops | fail-closedValueError at construction | not recorded |
LiteLLM TypeSafe compaction guardrail 1.104.0Tag v1.104.0 = 7964577. Prompt filter for tool results, not a safety block. Newly counted under Frameworks (8 Oct) | Per completed tool result: "is this result still needed for the current task?" | fallback-uncompacted requestDefault unreachable_fallback fail_open forwards the request uncompacted; a fail_closed option raises HTTP 502; 30 s | fallback-uncompacted requestSame handler | heldResults with noul < 0.2 are removed from the prompt (configurable); a missing answer keeps the result | fail-closedAt start-up: "TypeSafe guardrail requires an API key" | recordeddefault_on: False: runs only where configured; small results and protected messages are never sent |
jev-reranker 0.1.2 (PyPI)Tag v0.1.2 = d58594b. Reranker. Newly counted under Frameworks (8 Oct) | Noul probability per passage; listwise mode batches several passages per request. Sorts, cuts below the threshold, then keeps top_k | fail-closed429, 500, 502, 503, 504 and 529 retried up to max_retries=8, then raises; other statuses raise at once; 180.0 s per HTTP request | fail-closedResponseValidationError: invalid JSON, mismatched answer IDs, invalid noul, missing model or usage | heldPassage cut: relevance_rerank (threshold 0.2) drops passages below 0.2; rerank (threshold 0.0) keeps all and reorders | fail-closed"Set {api_key_env} or pass api_key to use Jev." at construction | recordedNo call for an empty list or top_k == 0; one passage in pairwise mode gets 0.5 without a call; passages cut to 4,000 characters |
LanceDB TypeSafeReranker v0.40.0Release v0.40.0 (7 Oct) = b1b080b. Reranker. Newly counted under Frameworks (8 Oct) | Noul per result ("whether the document is relevant"); one request per result or up to batch_size per request | fail-closed"API and response-validation errors propagate"; no timeout set in this file (the SDK default was not read) | fail-closedMismatched answer IDs or an invalid probability: ValueError | no-thresholdSorts by score only; any cut comes from the query's limit, outside this file (not read) | not recordedThe SDK reads TYPESAFE_API_KEY; its behaviour with no key was not read | recordedNull documents get score 0 without a request and stay in the list, sorted last |
- fail-closed or held (the item does not go ahead; a person, the host's prompt or, for two passage filters, the code's own cut decides)
- fail-open or acts-anyway; a "conditional" cell that can act is counted fail-open
- fallback-<what> (named in the cell)
- no-decision: nothing returned; the caller must handle it
- advisory (Jev never decides), or conditional where this page has no count
- not recorded in the code read, or no-threshold (no confidence cut; the cell says which)
Colours follow the shared legend on when Jev fails. The "What Jev decides" and "recorded" bypass cells are descriptions, not verdicts, and are left uncoloured.
Counts across the 5 rows (below threshold is counted separately from errors):
- Jev error or timeout: fail-closed 2/5 (jev-reranker, LanceDB), fallback 2/5 (LlamaIndex, LiteLLM), fail-open 1/5 (Spring AI).
- Malformed answer: fail-closed 2/5, fallback 2/5, fail-open 1/5 (the same rows).
- Below threshold: held 2/5 (LiteLLM, jev-reranker), acts-anyway 1/5 (Spring AI), no-threshold 2/5 (LlamaIndex, LanceDB).
- No key: fail-closed 4/5, not recorded 1/5 (LanceDB).
- Bypass: recorded 3/5 (LiteLLM, jev-reranker, LanceDB), not recorded 2/5 (Spring AI, LlamaIndex).
None of the 5 is a fail-closed safety filter. The one safety screen (Spring AI) is fail-open. The two rows that fail closed (jev-reranker, LanceDB) score relevance, not safety.
On when Jev fails, Spring AI and LlamaIndex stay in the Frameworks domain, where they were already counted. The three new rows (LiteLLM compaction, jev-reranker, LanceDB) are counted there under Frameworks, the nearest existing domain. No new domain is formed: only 2 of the new rows are rerankers, and LiteLLM compaction prunes tool results rather than retrieved passages. Counting LiteLLM compaction is a judgement: its code removes results from the prompt, but those results are tool outputs, not retrieved passages.
Timeouts and retries in general: errors and rate limits. What each route sends and keeps: data privacy. LiteLLM's complexity router is on model routing; only the compaction guardrail appears here.
TypeSafe's cookbooks: examples, not studies
TypeSafe publishes two cookbooks for this use. They are Documented vendor examples; the numbers in them are Reported by the vendor and are not graded. Their dates are not confirmed. Checked
- "Re-ranking": a noul asks "Could this candidate passage be from the cited precedent?" for each candidate. The vendor reports "top-1 accuracy from 5% to 18%" and "top-10 accuracy from 38% to 62%".
- "Classifying RAG passages": a second stage between retrieval and generation asks 4 questions per passage over 81 passages. Passages are excluded when injection is above 0.70 or relevance is below 0.45. No accuracy result is cited here.
What was not verified
- LanceDB: the TypeSafe SDK's default timeout and its behaviour with no key; whether the query's
limitcuts reranked results (outside the file read); the lancedb version used by the benchmark (requirements.txtnot opened). - jev-reranker: the async path was not read separately (it goes through the same
_rerank). - Spring AI and LlamaIndex bypass cells: not recorded, as published.
- MindStudio's 28 Sep reranker post: not re-read and not graded. Elastic B14's result is not restated here.
- jev-pulse: only
metrics.tsand the README were quoted;run.tswas reported by a helper; results are README-only. anessbelbati/jev-rerank-bench: README only. - The "Classifying RAG passages" cookbook date (27 Aug 2026) came from a fetch summary and was not seen on the page.
- "No graded pipeline study without a stake" holds for the benchmarks ledger as of 7 Oct plus this run's reads, within the search boundary above.
- The pool ran nothing: no package installed or run, no benchmark run, no Jev call.
Sources and check times (6 to 8 Oct 2026, UTC)
Source list: R11-S IDs, pinned links, paths and line ranges
- R11-S11: search suggestions from Google (
client=firefox), Bing (osjson) and DuckDuckGo (ac), 8 Oct 2026, 05:16:53–05:17:21 UTC. - R10-S215, R10-S216, R10-S223: iwhalen.com post; code at
82d9a0f(main.py). Reported - R10-S217, R10-S218: Hugging Face blog post; code at
d58594b(examples/eval.py,docs/eval.md). Reported - R10-S219, R10-S220, R11-S206 to R11-S208: LanceDB post;
lancedb/researchat4616259:reranking/benchmark/bench.pyL30, L38, L61–75, L106, L168–169, L290–302, L305–338 (opened 8 Oct, 05:13:50 UTC);README.mdandresults/README.md. Reported - R10-S213, R10-S214: Elastic B14 post and notebook at
3e3d688. - R3-S75: Parallel, "Testing out Jev" (B04).
- R11-S209, R11-S210:
mugunthank7/jev-pulseata1e796f:README.md(Reported),apps/api/src/rerank/metrics.tsL5–11,apps/api/src/rerank/run.ts,data/esci-search.json; 05:14:10 UTC. - R11-S211:
anweat/jev-websearch-evalat38dde16README.md(excluded); 05:14:10 UTC. - R11-S212, R11-S213: web search "Jev TypeSafe reranker evaluation nDCG independent benchmark RAG", 05:14:23 UTC, boundary since 15 Sep 2026;
anessbelbati/jev-rerank-benchatfecba75README.md, 05:14:40 UTC. Reported - R8-S84: Spring AI TypeSafe 0.4.0, tag
v0.4.0at51993bb:typesafe-spring-ai/src/main/java/org/springaicommunity/typesafe/rag/JevDocumentFilter.javaL118–132, L170–206. - R8-S86: LlamaIndex
JevRerankat421afdaand PyPI 0.1.1:packages/llama-index-postprocessor-jev/llama_index/postprocessor/jev/base.pyL69–83, L103, L145–151, L266–272;utils.pyL92–96. - R9-S120, R11-S48: LiteLLM 1.104.0, tag
v1.104.0at7964577:litellm/proxy/guardrails/guardrail_hooks/typesafe/typesafe.pyL1–7, L48, L52–55, L162, L167–171, L181–183, L193–203, L278–299, L302–316, L347–371 (read 6 Oct 2026). - R11-S200 to R11-S203, R11-S49: jev-reranker 0.1.2 (PyPI), tag
v0.1.2atd58594b:src/jev_reranker/reranker.pyL126–128, L233–235, L349–376, L507–517, L546–552, L716–737, L853–854;src/jev_reranker/_client.pyL44–67, L100–114, L132–136;src/jev_reranker/_ranking.py; 05:13:20 UTC, runner spot-check 05:19:45 UTC. - R11-S204, R11-S205, R11-S49: LanceDB release
v0.40.0(tag object2fd3239) atb1b080b:python/python/lancedb/rerankers/typesafe.pyL43–46, L67–69, L125–127, L158–177, L186, L199–203; 05:13:40 UTC, runner spot-check 05:19:45 UTC. - R11-S214 to R11-S216: TypeSafe cookbooks re-ranking, classifying RAG passages and the cookbook index, 05:14:50–05:14:55 UTC.
- R11-S56: TypeSafe case study (not a retrieval study), 05:14:53 UTC.