run:878a2d3a-67e1-440a-bf94-f7810f487bad · checks , 20:44–21:12 UTCJev (TypeSafe AI) evaluation evidence and home status: graded benchmarks, terms and access, 30 September 2026
This dated report records the evidence behind the release of 30 September 2026. The maintained pages built from it are: the home page status and notices, benchmarks, Jev vs an LLM vs a classifier, and updates on access status, where to use, errors and rate limits, TypeScript, official vs reseller and confidence thresholds. All reports: research history.
Run: run:878a2d3a-67e1-440a-bf94-f7810f487bad (scheduled regular research run 3) · Pool: pool_jev_catalog · Runner: single agent · Checks: 2026-09-30, 20:44 to about 21:12 UTC. The run was scheduled for 05:00 UTC but started at 20:44 UTC, so every check time below is an evening time.
The run's full notes and its source ledger (91 sources, IDs R3-S01 to R3-S106: URL, access time, label, excerpt, outcome) are held in the pool's private record. Key sources are linked below and on the pages each finding feeds.
Rules kept: no account, purchase, API key, or Jev or gateway inference call. Benchmark results are labelled Reported. The pool ran no benchmark and did not average or rank studies. No third-party code was downloaded or run; the only third-party file downloaded was one data file, a published JSONL log, used for an offline recomputation.
Previous release: regular-2026-09-29-integration-trust; report: integration trust.
1. Summary answer, as of 30 September 2026
- Can a new user get Jev right now? Yes. Signups are open, but new accounts get no free credit. The API is operational.
- The open state dates from the TypeSafe CEO's post of 27 Sep 2026, 22:33 UTC. That post was re-read on 30 Sep at 20:47 UTC.
- The status page read "All services are online" at 20:47 UTC. A console incident on 29 Sep was resolved at 22:05 UTC.
- The pool did not observe signup itself: the console still returns HTTP 403.
- OpenRouter still says no TypeSafe account is needed.
- What changed since 29 Sep:
- TypeSafe's published rate limits changed. They were 250,000 tokens/s and 1,200 requests/min. They are now 100K tokens/s and 40 requests/s. When the change happened is unknown.
- TypeSafe published
WorkflowEvals, the code behind its own benchmark. aiand@ai-sdk/typesafe-aieach had two new releases.- Prices, aliases, gateway listings and the TypeSafe SDKs did not change.
- Do TypeSafe's terms restrict publishing benchmarks?
- The current texts contain no restriction.
- Until 19 Sep 2026, the MCA had one. Section 2.3(f), in the version dated 27 Aug and archived on 16 Sep, forbade customers to "publish benchmarks or performance information about the Services". The version dated 19 Sep removed it.
- This answers the Hacker News question "Is benchmarking Jev still a ToS violation?": the clause is gone. The MCA's ban on distillation and on building competing products remains.
- This is a reading of the text, not legal advice.
- Is Jev accurate? Published tests disagree because they test different tasks against different baselines.
- Jev is usually ahead of cheap LLM setups.
- On verdicts that can be read off the text, it is close to frontier LLM judges once a threshold is tuned (Arize; arXiv 2609.26550).
- It is behind the best LLMs on derived judgments (math, code, logic) and on 14 of 15 social-science annotation tasks (arXiv 2609.24574).
- In the little evidence there is, it is behind trained classifiers on fixed label sets. That evidence is grade D only.
- The widely quoted "68%" is TypeSafe's own workflow benchmark. It measures agreement with the answers of GPT-6 Astra and Claude Fable 5.1, not accuracy against human labels.
- When should you use Jev instead of an LLM or a classifier? It depends on 10 stated criteria: labelled data, labels changing at run time, need for text, latency and on-device needs, residency, need for a threshold probability, input length, math/dates/language/images, volume, and vendor risk. Several studies support a hybrid: use Jev first and send low-confidence items to an LLM.
2. Corrections to published claims
| Page | Published claim | Correction | Evidence |
|---|---|---|---|
build/errors-and-rate-limits, where-to-use hub | "Limits 250k tokens/s and 1,200 requests/min" (checked 29 Sep) | TypeSafe now lists "100K tokens per second / 40 requests per second", still "can change without notice". Checked 30 Sep, 20:47 UTC. When it changed between 29 Sep 07:07 and 30 Sep 20:47 UTC is unknown | docs.typesafe.ai/models.md (Documented) |
| Home | Price "checked 28 Sep 2026" | Unchanged at $0.042/M input, output free. Rechecked 30 Sep, 20:47 UTC. The home rework replaces this line with the status block | models page |
| Home chart, "Accuracy and calibration" basis | "vendor claims, one operator suite and two independent calibration tests" | Out of date. The ledger now grades 14 tier-1 studies and 10 tier-2 rows (Reported) | benchmarks |
build/typescript | ai 7.0.122 and @ai-sdk/typesafe-ai 3.0.10 | Needs a version-drift note: 7.0.124 and 3.0.12 were released 30 Sep. Not re-audited | npm ai, npm @ai-sdk/typesafe-ai |
where-to-use/access-status | "no new event" after 28 Sep | New status-page events on 29 Sep: routine maintenance ended 15:32 UTC, and "Console is unavailable" was resolved at 22:05 UTC. API uptime is 99.828% (was 99.819%) | status page |
build/confidence-thresholds | Alex Molas post: "Argument, no experiment" | Minor. The post includes a small example (Jev gives a fair coin 0.92 heads). Grade D either way | Molas post |
alternatives hub, and any page citing "68% on TypeSafe's own benchmark" | 68% presented as accuracy | 67.8% is agreement with LLM reference answers on TypeSafe's four workflows. It is not accuracy against people or ground truth | evals.typesafe.ai |
Publisher's note: no published page of this site used the "68%" wording; the new benchmarks page states it correctly. Correction to the run's own draft: OpenRouter's typesafe/jev-router /endpoints call returns an empty list. The 29 Sep run read only the models list and never queried this call, so whether the empty list is new is unknown. It is not reported as a change.
3. Per-page findings and ship verdicts
Home rework (status and notices): ready
- Signup state is
open_with_conditions, since 2026-09-27T22:33:25Z (executive source, Documented). Service state isoperational, since 2026-09-29T22:05Z. observed_directlyis false: all three console paths returned HTTP 403.- Four notices, all current:
new-user-credit-disabled(Documented);lookalike-resellers(Reported; Eye Security's report and three named sites were still online at 20:53 UTC);rate-limits-changed(Documented; event; retire on 7 Oct);typesafe-workflow-evals-code(Documented; event; retire on 6 Oct). - Retired notices: none, because the 29 Sep run produced no notice file.
- Not made notices: the OpenRouter router's empty endpoint list (meaning not stated); the AI SDK releases (version drift only); the OpenAI "Decision API" headline (secondary, and not a change to Jev).
benchmarks: meets the ship condition
- The condition asks for at least 8 studies read at the primary source, with operator, task, n, baseline and materials recorded. Twelve qualify: B01–B03, B05–B09 and B11–B14. B04 (Parallel) was read at its primary source but states no n.
- At least 3 must compare Jev with a baseline. Nine compare it with an LLM, a classifier or a search baseline.
- The terms answer is stated.
- Key ledger facts (all results Reported): B01 TypeSafe WorkflowEvals, grade B, COI, LLM-labels, 67.8% mean agreement against 66.8–74.1% for GPT-5.6 workflows, reference answers from GPT-6 Astra and Claude Fable 5.1, and the homepage's 193.6x/444.6x cannot be reproduced from the rounded figures. B02 Arize, grade B: Jev and Opus 5 both 87% on RAGTruth after threshold tuning (76% vs 83% at a 0.5 cutoff). B03 NavyaAI, grade C: AG News, Jev 87.5–88% vs 71–83% for cheap LLMs. B04 Parallel, grade D. B05 LiteLLM, grade C, COI, small-n: 95.0% vs Haiku 4.5's 73.75%, ratios reproduced. B06 Jevals, grade B, with the PubMedQA row recomputed from the published log (accuracy 91.33%, median 0.438 s and p95 0.653 s match; calibration error 5.10 vs 5.0). B07 JevBench, grade B, sealed. B08 CMU, grade B: within 3 points of GPT-6 on text-readable verdicts, behind on math, code and logic, cascade 0.9 points ahead at 41% of the fee; version changed on 29 Sep. B09, grade B, no-baseline: renaming options flips 32.5% of decisions. B10 MindStudio, grade D, secondary. B11 Opper, grade B, COI: Jev and Kev within noise. B12 Bryo, grade C: Jev lost to Gemini (figures from Prefactor). B13 NYU Abu Dhabi, grade A provisional: behind the per-task best LLM on 14 of 15 tasks by a median 11.6 macro-F1 at a median 44 times lower cost. B14 Elastic, grade A provisional: +0.018 to +0.021 nDCG@10 over hybrid search.
- Tier 2: punk2898 A provisional; DIY logprobs C; Casco C; Anthus calibration B; SciFact pre-registered A provisional; Rene-1, Ollaya, Jeff, Molas and Vercel
fxD. Full rows on benchmarks.
alternatives/jev-vs-llm-vs-classifier: meets the ship condition
benchmarksships. The page has 10 criteria, each with cited evidence and a label, and a head-to-head table with 13 ledger rows and 2 carried rows.- Caveat printed: the comparison with trained classifiers and zero-shot NLI rests only on grade D rows (B04, B10).
- The page gives no prices.
Updates to existing pages
Access-status timeline; limits corrections on where to use and errors and rate limits; TypeScript drift note; two new official surfaces on official vs reseller (n8n-nodes-typesafe-ai, WorkflowEvals); cross-links from alternatives, use cases, products, confidence thresholds, limitations and videos; Research page entries. build/claude-code-mcp needs no drift note: the six audited integrations have no newer release (checked 21:10 UTC).
4. 24-hour discovery scan in brief
- Cutoff and gap: the cutoff was 2026-09-30 20:45 UTC, with the window starting 24 h earlier. The 13 h 40 min between the previous window's end (29 Sep 07:05) and this window's start went unscanned, except for Hacker News stories and the official-organisation and package checks.
- Counts (index counts, not adoption): 28 Hacker News stories; GitHub 364 "jev" repositories created, 10 "jev benchmark", 7 "jev eval"; 28 npm packages; 6 DEV.to articles.
- YouTube was attempted. It is readable but its dates are relative, so its items are date-ambiguous.
- Unknown, not zero: Reddit, X, TikTok, Lobsters and Product Hunt, which were blocked or required login.
- Flags raised: benchmarks (Elastic, LanceDB, Casco, Halv, CLM-8B and new evaluation repositories); TypeSafe releases (WorkflowEvals, and a push to the n8n node); a competitor event (the OpenAI Decision API, a headline only); version drift (
aiand@ai-sdk/typesafe-ai).
5. Evidence gaps
- Signup: not observed directly (console HTTP 403). No second first-hand signal after 28 Sep. X cannot be searched.
- Rate-limit change: the exact time is unknown. No docs changelog records it.
- B08: the reproducibility package was not located. If it is found, the row may move to grade A.
- Linked repositories not opened: B13, B14, B15 and B24, which is why their A grades are provisional. JevBench's code repository was not opened either.
- Primary source not identified: the MindStudio summary (B10). A related lead is the nicobrenner gist "94% on Banking77".
- Not extracted: B14's Jina reranker comparator row; the LanceDB results.
- No graded head-to-head with a fine-tuned classifier or zero-shot NLI with published materials.
- No graded non-English study except Bryo (grade C). The only Italian item is a repository lead.
- The OpenAI "Decision API" was not checked at OpenAI's own source.
LocalLLaMA/typed-decisions: who operates it is unknown. Its subsets carry the same four workflow names as TypeSafe's WorkflowEvals (Documented observation only).- jevcal's terms claim was not re-read. The Hacker News claim that the Terms of Use once had a benchmark clause is not confirmed by the 15 Sep capture.
6. Not done
- Tier-2 rows left unreviewed: M37, Nate Herk and Funicello videos; "Jev's Founder Explains: Why You Can't Trust AI Benchmarks" (YouTube, 30 Sep); LanceDB, Halv, CLM-8B (VentureBeat), Jeeves (PostHog) and calfram-bench results; Jeff figures (only the README claim was read); the Rene-1 Decision Index suite; the Ollaya n.
- Other checks not made: Vercel's model page was not checked for provider-terms pass-through (OpenRouter and Cloudflare were); Netlify was not rechecked, by plan; GitHub code search and trending were not run; GitHub, npm, Hugging Face and DEV.to were not scanned over the 13 h 40 min gap; the
WorkflowEvalsHugging Face datasets and published runs were not opened; then8n-nodes-typesafe-ainpm package's publisher was not checked against TypeSafe.
Suggested follow-ups for later runs: the planned Laya vs Kev comparison can reuse rows B11 and B18–B20, the Ollaya and JevBench figures, and the leads for Nimble and Ollama's "Jev-style" support. The planned data-handling page can start from this run's legal index and OpenRouter's embedded TypeSafe data policy (retainsPrompts: false, canPublish: false; not interpreted).
7. Key sources
- Access: status.typesafe.ai; docs.typesafe.ai/models.md; docs quickstart; CEO post, 27 Sep; OpenRouter Jev guide; Eye Security report. All 30 Sep 2026, 20:47–21:00 UTC.
- Terms: Terms of Use; AUP; MCA; DPA; Privacy Policy; MCA archived 16 Sep; MCA archived 20 Sep. Read 20:48–20:51 UTC.
- Official releases: typesafe-ai/WorkflowEvals; typesafe-ai/n8n-nodes-typesafe-ai; npm ai; npm @ai-sdk/typesafe-ai; PyPI typesafe-sdk.
- Benchmarks: every study is linked in its row on the benchmarks ledger, with its check time.