run:a0659f82-68c7-4d81-84ef-c1f25d543064 · checks , 05:14–05:33 UTCRun 7 report: trading and fraud gates, the Jev failure checklist and Jev as a judge (4 Oct 2026)
This dated report records the evidence behind the release of 4 October 2026 (regular-2026-10-04-trading-failures-judge). The maintained pages built from it are: trading and fraud gates (new), when Jev fails (new, the failure checklist across use cases), Jev as a judge (new), the refreshed Products hub, the subprocessor finding on data privacy, the drift note on Claude Code and MCP guardrails, and the refreshed home status and access status. Facts on those pages may be rechecked later; this report keeps the 4 October readings.
Run: run:a0659f82-68c7-4d81-84ef-c1f25d543064 (scheduled regular research run 7) · Pool: pool_jev_catalog · Runner: one runner plus two helper agents that read source in parallel (helper A: trading and fraud, then two guard projects; helper B: judge tools). The runner spot-checked 29 lines at the same commits (13 for helper A's trading rows, 4 for its guard rows, 12 for helper B); all 29 matched · Checks: 2026-10-04, 05:14 to 05:33 UTC; runner started 05:13:47 and finished the write-up at 05:34 UTC · Manager review: accepted at 05:38 UTC after its own spot-checks at the pinned commits and an arithmetic check of every total below.
The run's full notes and its source ledger (72 records, IDs R7-S01 to R7-S85: URL, access time, label, excerpt, outcome) are held in the pool's private record. Key sources are linked below and on the pages each finding feeds.
Rules kept: no account, key, purchase, Jev API call, exchange, broker, payment or testnet call, bot run, eval run, or third-party code run. Nothing in this run is Tested. Every finding below is Documented (official docs, legal text, registry metadata, or source read at a pinned commit), Reported (a named third party's own claim) or Unverified (a lead only). No trading or investment advice, and no performance figures.
Previous release: regular-2026-10-03-email-products-n8n; report: email and lead sorting, oh-my-claudecode v5.6.0, Products refresh and n8n nodes. All dated reports: Research.
1. Summary answer, as of 4 October 2026
- Access is unchanged (checked ). Signups are open, but new accounts get no free credit, as since 27 Sep, 22:33 UTC. The API is operational, with no status-page event since 29 Sep, 22:05 UTC. Limits are still 100K tokens per second and 80 requests per second. No legal-date change and no TypeSafe release. All four Home notices are kept. Documented
- Trading and fraud gates: ships. Seven projects were read at pinned commits. Of 4 trading bots, 2 can still place a new order on a Jev error: QuantDinger (an LLM fallback, then allow) and btc-hft-jev (a rules-only fallback). All 4 have a path that trades without Jev. QuantDinger's order call sites are now traced; this had been open since 28 Sep. The 3 fraud or payment gates call Jev in code, but none of the 3 has a payment call site. Documented
- When Jev fails: ships. 39 implementations in 7 domains, recomputed from the published tables. On a Jev error, 21 of 39 stop or hold the action and 5 of 39 let it through unchecked (4 content-moderation bots and 1 Claude Code hook). Of the other 13: 7 fall back, 3 only advise and 3 return no decision. Documented
- Jev as a judge: ships. Four tools were read at pinned commits. DeepEval aborts the run on a Jev error by default. The other 3 drop the failed item from the score, and openlayer jevals' report still says it passed. 1 of 4 codes an escalation for unsure verdicts (openlayer, and only if you switch it on). The Jevals harness code is still not public. Documented
- Data handling: TypeSafe's trust center lists 6 subprocessors; Nebius and CoreWeave were added on 2 Oct. Now on the Home notices and the data privacy page. Documented
2. Access recheck, home notices and the 24-hour scan
Volatile fields and home status: unchanged
- Rechecked 4 Oct 2026, 05:15:49 to 05:17:04 UTC (R7-S20 to R7-S49). Previous values: 3 Oct 2026, 05:15–05:18 UTC.
changed_since_last_run: no. - Unchanged, all Documented: status "All services are online" (API 90-day uptime 99.826%); model
jev-1.13.0with both aliases; price $0.042 per million input tokens, output free; 100K tokens per second and 80 requests per second; console still HTTP 403, so signup is not observed by the pool; OpenRouter needs no separate TypeSafe account and its endpoint is on OpenRouter's ZDR list; Vercel lists one provider, DigitalOcean, withhas_zdr: false; Cloudflare shows "Zero data retention Yes"; SDK versions; legal "Last updated" dates (Terms 19 Sep, AUP and MCA 23 Sep). - Headline unchanged: "Available to new users: signups open, but new accounts get no free credit. API operational." Shown as "checked 4 Oct 2026, 05:15 UTC".
- Notices: four kept, none retired.
new-user-credit-disabled(rank 1),lookalike-resellers(rank 2; all three named sites still answered HTTP 200),rate-limits-changed-2026-10-03(rank 3; retires 10 Oct, 05:16 UTC) andtypesafe-workflow-evals-code. No timeline additions. - Manager decision, 05:38 UTC: the runner added
subprocessors-added-2026-10-02at 05:31 UTC withshown: false(see section 7). The manager shows it on Home at rank 4, because data handling affects every user.typesafe-workflow-evals-codemoves to rank 5, not shown; it is still current and keeps its retire date of 6 Oct 2026, 16:25 UTC. None of the listed notice types fits the subprocessor change exactly;new_official_releasewas used and the notice's note says so.
S1. 24-hour scan
- Window: 3 Oct 2026, 05:14 UTC to 4 Oct 2026, 05:14 UTC. It starts exactly at run 6's cutoff: no gap, no overlap, no backfill needed.
- Index counts (R7-S01 to R7-S05): 4 Hacker News "jev" stories and 8 comments; 225 GitHub repositories created matching "jev", 17 with the
jevtopic; 11 npm "jev" packages published in the window; 24 Hugging Face "jev" models modified; 3 DEV.to articles. These are search-index counts, not verified integrations or adoption. - No TypeSafe release, legal-date change, limits change or incident. One pinned-integration patch release (oh-my-claudecode v5.6.1; section 6). Four new trading, fraud or judge repositories were found and read for the new pages.
- Leads not read (Unverified): two self-run evaluations (jev-safety-eval, jev-zh-tw-eval) and one set of security experiments (system-one-security), queued for run 9.
- Reddit (HTTP 403) and X (login redirect) could not be read; TikTok and YouTube were not attempted: unknown, not zero. GitHub search calls used: 5 of 10.
3. Findings per page and ship verdicts
3.1 use-cases/trading-and-fraud-gates: ships
Of 4 Jev trading bots read in source, 2 can still place a new order when Jev errors; all 4 can trade without Jev. None of 3 fraud or payment gates moves money.
- Page: Jev trading bot (TypeSafe AI): does a Jev failure still place the order? Title query source: Google suggested "jev trading bot"; Bing and DuckDuckGo suggested "jev trader" (05:21 UTC, R7-S11). "jev fraud" had no relevant suggestion, so "fraud" stays in the subtitle. A suggestion is a demand signal, not a volume.
- Projects and pinned commits: QuantDinger
a5a9f4c, jev-traderb587759, Jev Tradea3f2f83, btc-hft-jev96f3541(new, 3 Oct), jev-fraud-shieldeeb4fc6, fraud-jevaa33b97, expense-policy gate161c33b(new, 3 Oct). Every failure cell on the page has a path and line range. - QuantDinger:
ai_decision_filter.pywas rewritten since 28 Sep (now 743 lines). A Jev error, a malformed answer, risk or execution confidence below 0.55, or a missing key all go to an LLM fallback, then allow; the timeout is 8 s. Aninsufficientanswer never blocks. Exits, the grid, DCA and martingale strategies skip the gate. A failed audit write does not block the order. The filter is off by default, and strategy orders default to a virtual account. Order call sites traced: the strategy v2 execution path (the gate runs beforeINSERT INTO pending_orders), quick trade, and the Alpaca route. - Defaults: all 4 bots default to mock, paper, testnet or a virtual account. In jev-trader and Jev Trade, the mock model still sends orders when a wallet key is set: "mock" means no Jev, not no trading. jev-trader turns any answer other than "sell" into a buy. On a Jev error, jev-trader and Jev Trade place no new order, but a resting order stays on the book.
- Not claimed: that Jev trading bots are unsafe, or anything about trading results. Jev Trade's creator's claims of filled orders stay Reported.
3.2 build/when-jev-fails: ships
Of 39 Jev implementations read in source, 21 stop or hold the action when Jev errors and 5 let it through unchecked: 4 content-moderation bots and one Claude Code hook.
- Page: When Jev fails (TypeSafe AI). The title is an internal label: six failure phrases got no relevant suggestion on any of four suggestion endpoints at 05:21 UTC (the same result as 3 Oct).
- 39 rows in 7 domains, 5 to 7 rows each; the official plugin is excluded from every denominator. The runner recounted every column from the map with a script, and the manager checked that each row below sums to 39. No new failure claims: every cell comes from a published page.
- The action goes ahead on a Jev error in 7 of 39: the 5 fail-open rows plus the 2 trading fallbacks. Without the trading domain, the six earlier domains give 32 rows: fail-closed 16, fail-open 5, advisory 3, fallback 5, no-decision 3.
- The checklist has 8 questions, each with counts and example rows. Below-threshold totals are kept apart from error totals.
| Condition | Counts | Not recorded |
|---|---|---|
| Jev error | fail-closed 21, fail-open 5, advisory 3, fallback 7, no-decision 3 | 0/39 |
| Malformed answer | fail-closed 16, fail-open 8, advisory 3, fallback 8, no-decision 3 | 1/39 |
| Below threshold | held 10, acts-anyway 7, fallback 5, no-decision 2, advisory 4, no-threshold 6 | 5/39 |
| No key | fail-closed 13, fail-open 2, advisory 3, fallback 4 | 17/39 |
| Bypass | recorded 17 | 22/39 |
3.3 use-cases/llm-as-judge: ships
Of 4 Jev judge tools read in source, DeepEval stops the run on a Jev error; the other 3 drop the item from the score, and openlayer jevals still reports it as passed.
- Page: Jev as a judge (TypeSafe AI): accept when confident, escalate when unsure. Title query source: Google suggested "jev llm as a judge" and "jev as a judge"; Bing and DuckDuckGo also suggested "jev as a judge"; "jev evals" was suggested on Google, Bing and DuckDuckGo (05:21 UTC, R7-S11). The draft "Jev as an LLM judge" was replaced because it contains no suggested phrase.
- Tools and pinned commits: DeepEval JevEval
a200ece(version 4.2.8; no matching git tag), openlayer jevals0a8f895, did-they-answere0ee60c(new, 4 Oct), competitor-hunter40613ea(new, 3 Oct; it loads whatever Jev caller the user has installed, so that part is not pinned). Jevals: the harness code is still not public, and the row says so. - Details: DeepEval's
ignore_errorsdefaults toFalse(evaluate/configs.pyL46), so a Jev error aborts the run. DeepEval scores a test case 1.0 when every Choice question is mostly "not applicable". openlayer's gate defaults toon_error="allow". - Evidence summary: rows B02, B06 and B08 copied unchanged from benchmarks (Reported). The pool ran no judge benchmark. Not claimed: anything about Jev's accuracy as a judge.
4. Corrections to published pages, and how they were applied
| Page | Published claim | Correction | Applied as |
|---|---|---|---|
build/claude-code-mcp; run 6's checklist preparation (tally in the 3 Oct report) | oh-my-claudecode on a Jev error: fallback-heuristic; tally on a Jev error: advisory 2, fallback 6 | At default settings oh-my-claudecode is advisory: shadow mode, Jev never decides. Tally: advisory 3, fallback 5 | build/claude-code-mcp wording checked: it already said "Shadow by default: Jev never decides"; the matrix row now also says the points are advisory at default settings and that its cells describe the path once Jev is asked; when-jev-fails counts it as advisory. The 3 Oct report keeps its dated reading |
| Run 6's checklist preparation (internal) | No-key cells for jev-claude and jkudish: "not mapped" | Both are recorded on build/claude-code-mcp as advisory | when-jev-fails no-key column uses the recorded cells |
| Below-threshold cells on the published domain pages | Older local terms per page | Mapped to the 4 Oct terms: CyrilBaah, LiteLLM, 0xNatoshi and Zafer-Liu → no-threshold; agent-router and the rows that send items to a review group or review label → held | when-jev-fails below-threshold column; relabels only, no fact changes |
| Bypass cells (run 6 preparation) | Email re-processing guards and agent-router's refusal to send secrets counted as bypasses | Neither is a bypass. Of the 5 email projects, 1 has a bypass: Jevmail's /preview simulator | when-jev-fails bypass column (recorded 17/39) |
use-cases/email-and-lead-sorting | Unsure cells: Jevmail advisory; spam-leadgen fail-closed | Under the 4 Oct terms: Jevmail acts-anyway (flagged); spam-leadgen fallback-not qualified | Both cells relabelled on the email page; under the same rule vynnlee's "fallback: Review label" now reads "held: Review label". No fact changes |
use-cases/email-and-lead-sorting | Cost range starts at "about $0.004 per 1,000 for very short emails" (question allowance left out) | $0.042 per million × 600 tokens (100 email + 500 question allowance) × 1,000 = $0.0252. The table rows already include the 500-token allowance (for example vynnlee: 1,700 characters ÷ 4 + 500 = 925 tokens → $0.039), so $0.004 did not follow the table's own method | Cost line now leads with "about $0.025 per 1,000 very short emails"; $0.004 dropped |
use-cases hub, build/testing-without-a-key, products | QuantDinger links to ai_decision_filter.py lines as read on 28 Sep | The file was rewritten (743 lines), so the 28 Sep line links point at the wrong code | Replaced with links pinned to a5a9f4c on the use-cases hub, testing without a key and Products |
No earlier fact was found false: corrections 1 and 2 are relabels under the 4 Oct vocabulary, correction 3 aligns one figure with the page's own method, and correction 4 replaces stale line links. Older report pages are dated snapshots and keep their readings.
5. Products changes
- New total: 67 rows (A 47, B 19, C 1), up from 61 (A 39, B 21, C 1), checked against the published Products page. The manager checked the arithmetic: 61 + 6 new = 67; A 39 + 2 + 6 = 47; B 21 − 2 = 19. Independently verified in production: still 0 of 67.
- QuantDinger (already A): pin updated to
a5a9f4c; failure description now "LLM fallback, then allow; order call sites traced". - jev-trader and Jev Trade: B → A, now read in source. Jev Trade's creator's claims of filled orders stay Reported.
- 6 new A rows: btc-hft-jev, jev-fraud-shield, fraud-jev, the expense-policy gate, did-they-answer, competitor-hunter.
- Not counted: DeepEval JevEval and openlayer jevals are library features, handled like LiteLLM.
- The two guard rows read in spare time (gulbaki/jev-llm-guard at
7c7b1b5and madisonrickert/jev-permission-gate at7349bc9) are not added this run. They are kept for the nextbuild/claude-code-mcpedit, when Products would become 69 rows (A 49). They are not counted on when-jev-fails either.
6. Drift
- oh-my-claudecode v5.6.1 (3 Oct, 06:44 UTC) is a patch release, so it gets a drift note only on Claude Code and MCP guardrails: "failure path checked at v5.6.0". Documented
- No other pinned integration changed, including all 5 n8n nodes (checked 05:28 UTC; official
@typesafe-ai/n8n-nodes-typesafe-aistill 0.9.0).
7. TypeSafe's trust center: subprocessors
- Read 4 Oct 2026, 05:30 UTC (R7-S80). The trust center is a script-only page; it was rendered read-only with the headless browser already on the host, using a temporary profile. No form was submitted and no access was requested. Documented
- It lists 6 subprocessors, all in the USA: Amazon Web Services, Modal, Slack, Google Workspace, Nebius and CoreWeave. Its updates list shows "Nebius and CoreWeave — Published October 2, 2026 — Added subprocessors".
- For Modal, Nebius and CoreWeave, TypeSafe's wording is that customer AI prompts are "processed, but not stored" on compute nodes those companies manage. For AWS it says customer information for requests is "stored and processed".
- The trust center lists a "SOC 2 Type II - 2026" report behind "Request access". The pool did not see the report.
- Applied: a dated line on data privacy and the Home notice
subprocessors-added-2026-10-02(rank 4). The route labels on the data privacy page (OpenRouter ZDR list, Vercelhas_zdr: false, Cloudflare "Zero data retention Yes") do not change.
8. Evidence gaps and uncertainties
- Cells not recorded across the 39 checklist rows: no key 17/39; bypass 22/39 (many on pages published before the bypass column existed); below threshold 5/39 (the n8n rows).
- jev-trader's no-key path runs inside
@ai-sdk/typesafe-ai, which was not read. - QuantDinger: the worker's upgrade of queued
signalrows to real-account orders may be a way past the filter, but it was not traced end to end. The LLM fallback prompt was not read. - did-they-answer: the abort on a malformed answer was traced in source but not run. competitor-hunter: the user's installed Jev caller version is not pinned. openlayer jevals: the PyPI version was not checked. Jevals: harness not public; its error handling is Reported only (methodology page, R7-S77).
- VentureBeat returned HTTP 429 (R7-S81), so its two relayed Products claims stay unchecked, with their 21 Sep date. The SOC 2 report was not seen.
- Reddit (HTTP 403) and X (login wall) are unknown. YouTube, TikTok and GitHub code search were not attempted.
- Helper-agent reads were spot-checked (29 lines), not re-read line by line. Whether any project behaves as its code reads: nothing was run.
9. Not done
- Trading and payment leads not read: the Jono717, gilesknap and buberlo jev-trader repositories; caiovicentino/jev-risk-check-provider (the runner's next payment candidate); PayShield, fraud-classifier, Nutlope/jev-fraud. TradeRank was not found.
- Judge leads not read: Zaious/JevTRPG, the texposit "Judge Jev" post, DeepEval's ConversationalJevEval.
- New leads for run 9: jev-safety-eval, jev-zh-tw-eval, system-one-security. The run 6 carry list is unchanged.
10. Timings (UTC, 4 Oct 2026)
| Step | Time |
|---|---|
| Runner start | 05:13:47 |
| S1 cutoff | 05:14:00 |
| Helpers launched | 05:14 |
| S2 recheck | 05:15:49 |
| S1 scan | 05:16:32 |
home-status.json written | 05:17:32 |
| QuantDinger read | 05:17–05:20 |
| Helper B back; spot-checks | about 05:19:40; 05:20–05:26 |
| Autocomplete | 05:21:29 |
| Helper A back; spot-check | about 05:21:30; 05:21:48 |
| Page notes | 05:23–05:27 |
| Source ledger | 05:28 |
| Scan file | 05:29 |
| Spare time (trust center, VentureBeat, Netlify, guard rows) | 05:30–05:33 |
| Write-up done | 05:34 |
| Manager review and decisions | 05:38 |
11. Rules followed
- No account, key, purchase or Jev API call. No exchange, broker, payment or testnet call. No bot run, eval run, or third-party code run.
- The host's headless browser was used read-only with a temporary profile. Scratch downloads and the browser profile were deleted; only three small read-only fetch scripts written by the runner remain. No symlinks.
- No email addresses or other personal data in any note or on this page.
12. Key sources
- Access and limits: status.typesafe.ai (R7-S20); docs.typesafe.ai/models.md (R7-S21); OpenRouter ZDR endpoint list (R7-S48); Vercel endpoints (R7-S49); Cloudflare model page (R7-S29). All 4 Oct, 05:15 UTC.
- Trading and fraud (R7-S50 to R7-S63, 4 Oct): QuantDinger
a5a9f4c; jev-traderb587759; Jev Tradea3f2f83; btc-hft-jev96f3541; jev-fraud-shieldeeb4fc6; fraud-jevaa33b97; expense-policy gate161c33b. File paths and line ranges are on the trading page. - Judge tools (R7-S70 to R7-S78, 4 Oct): DeepEval
a200ece; openlayer jevals0a8f895; did-they-answere0ee60c; competitor-hunter40613ea; Jevals methodology (Reported). - Failure checklist: recomputed from the published domain pages; rows and links on when Jev fails.
- Guard rows for the next edit (R7-S84, R7-S85): jev-llm-guard
7c7b1b5; jev-permission-gate7349bc9. - Drift: oh-my-claudecode releases (R7-S47, 05:15 UTC).
- Trust center (R7-S80, 05:30 UTC): trust.typesafe.ai/subprocessors; trust.typesafe.ai/updates.
All dated reports: Research. Previous report: 3 Oct 2026, email and lead sorting, Products and n8n.