Content moderation with Jev (TypeSafe AI): what happens to a post when Jev fails or is unsure
Independent sites publish advice on moderating posts and comments with Jev, under titles such as "Content Moderation With Jev: Rules, Thresholds, Cost" and "Jev Content Moderation Recipe & Schema Guide". The one advice page that could be read in full for this page (opentweet's) says nothing about what to do when the Jev call fails. This page reads six public moderation projects in source and compares them on four questions: what happens to a post when Jev fails or is unsure, who can undo a removal, and what data is sent. It then compares two ways to ask Jev about harms and ends with eight design rules.
Of six Jev moderation projects read in source, four leave the post up or let it through when the Jev call fails; one holds it for a person; one returns no decision.
The six were chosen by GitHub stars, with one planned lead added (see how the six were chosen), so this is not a count of how moderation with Jev is usually built. Five are code-backed integrations (a WordPress plugin, two Discord bots, a Telegram bot and an input guard for AI agents) and one is a runnable demo. The pool ran none of them, and none is independently verified as operating.
When does Jev see the post?
A failed Jev call means something different depending on whether the post is already visible. Documented source, 2 Oct 2026
- Held before it appearsWordPress Jev Comment Triage. A new comment waits in WordPress's queue while a background job asks Jev. If Jev fails: the comment stays hidden and is retried; after three failed attempts it waits for a person.
- Already visibleJev-Moderation-Bot and soter (Discord), the Telegram anti-spam bot. Jev checks messages people may already have read. If Jev fails: the message simply stays up. A removal, when it happens, comes after publication.
- Checked before an AI agent reads itmastra-jev-moderation guards a user's message before a Mastra agent sees it. If Jev fails: the message goes through to the agent.
- No publishing systemjev-demo-moderator is a demo API and console. If Jev fails: the item comes back with an error and no action.
Six projects: failure, doubt, malformed answers and appeals
Open a situation to highlight its column and read what the six do; opening one closes the others. Every cell comes from source at the commit shown; README descriptions were not used. Documented Read 2 Oct 2026, 05:23–05:27 UTC
The Jev call failsHTTP error, timeout or open circuit breaker
Four of six leave the post up or let it through: brainstormity's Discord bot (it warns the mod-log channel after three failures in a row), soter ("a scan error never blocks chat"), the Telegram bot (logs the error, keeps the message) and mastra (lets the message through after 5 s). WordPress holds the comment and retries. The demo returns an error and no action.
Jev is unsureA probability in the middle band
Only WordPress sends its middle band to a person: spam or scam from 0.45 is held for review. soter writes hate speech above 0.5 to a mod-log for review but takes no action there. Three projects delete in a band below certainty: brainstormity from 0.70, soter whenever its spam Choice says high_spam, whatever the probability, and the Telegram bot from 0.81. See the threshold chart.
Malformed answerMissing fields or an answer the code cannot use
Five of six treat a missing answer as "safe" or as an error that lets the post through. WordPress reads missing fields as 0.0, which can restore WordPress's own approval. brainstormity treats a missing choice as LEGITIMATE. The demo returns a per-item error. Treat a missing answer as a failure, not as 0.0.
Someone appealsCan a person undo a removal?
Two projects have a pardon in code (brainstormity's "Pardon" button, soter's /pardon). Both lift the punishment, such as a timeout, but neither restores the deleted message. WordPress relies on its own moderation queue and spam folder. The Telegram bot has no reversal path or review queue. mastra refuses the user's turn, so there is nothing to restore; the demo has no queue.
| Project (commit) | Jev error or timeout | Thresholds and what each band does | Malformed answer | Appeal or reversal |
|---|---|---|---|---|
WordPress Jev Comment Triage (d848a5b)Code-backed integration (WordPress plugin). Noul spam, Noul scam, Score toxicity 0–2. Comments held before they appear | fail-closed: held, retriedAfter 3 failures it stays held for a person. With no provider set up, WordPress's own decision applies (may publish) | held: middle band held for reviewspam ≥ 0.75 or scam ≥ 0.65 → spam folder; spam or scam ≥ 0.45, or toxicity ≥ 1.5 → held; else WordPress decides | fail-open: read as 0.0Can restore WordPress's would-be approval | WordPress queue and spam folderRestoring from spam uses WordPress's own screens (not plugin code) |
brainstormity/Jev-Moderation-Bot (325ea7f)Code-backed integration (Discord bot). One Choice (legitimate, spam, scam link) plus a Noul for ban review. Messages already visible | fail-open: left upMod-log warning after 3 failures in a row | acts-anyway: deletes from 0.700.70–0.95 and ≥ 0.95 both delete, then a warning or timeout ladder; below 0.70 left up | fail-open: left upMissing choice treated as LEGITIMATE | Pardon, message not restoredLifts the timeout; a ban needs two-step admin confirmation |
frolleks/soter (756579a)Code-backed integration (Discord bot). Noul hate speech plus a spam-level Choice. Messages already visible | fail-open: left up"a scan error never blocks chat" | advisory: review log, but some deletions without onehate > 0.8 delete; > 0.5 mod-log review, no action. high_spam deletes at any probability; medium_spam to review log | fail-open: left upMissing answers read as hate 0 and no_spam | Pardon, message not restored/pardon stops counting the latest violation and lifts the timeout |
backmeupplz/jev_antispam_bot (0b07e88)Code-backed integration (Telegram bot). One Noul per spam pattern, 29 in all. Messages already visible; group admins skipped | fail-open: keptError logged; message stays | acts-anyway: deletes from 0.81Highest Noul ≥ 0.81 (SPAM_THRESHOLD, allowed 0.5–1) → delete; else kept | fail-open: keptInvalid answers throw, same as an error | None foundNo reversal path or review queue in src/bot.ts or src/deletion.ts |
CodeAlive-AI/mastra-jev-moderation (8735ea8)Code-backed integration (input guard for Mastra agents). A Noul "must this be blocked?" decides; a Choice category is only a log label | fail-open: let throughTimeout (5 s), HTTP error, unparsable answer or open circuit breaker (3 failures) | acts-anyway: one cut pointP(block) ≥ 0.7 → the agent's turn is refused with a reason; else passed | fail-open: let through | Not applicableThe user's turn is refused; no queue |
Zafer-Liu/jev-demo-moderator (cc47b25)Runnable demo (API and console, no publishing system). Noul spam, Score toxicity 0–3, Choice action | no-decisionItem returned with an error field; a whole-batch failure returns HTTP 502 | no-threshold: Jev's own Choice decidesCode only tightens "publish" to "review" when spam ≥ 0.9 or toxicity ≥ 1.5; it never deletes | no-decision: per-item errorNot handled separately | None"Review" is a label in the console |
- fail-closed or held (a person or the host's own prompt decides)
- fail-open or acts-anyway; a "conditional" cell that can act is counted fail-open
- no-decision: nothing returned; the caller must handle it
- advisory: Jev never decides
- not recorded, not applicable or no-threshold (the cell says which)
Colours follow the shared legend on when Jev fails (changed ; this page used its own colours before). The "Appeal or reversal" column describes a feature, not a verdict, so it is uncoloured.
What each project sends to Jev, languages and tests
| Project | What is sent to Jev | Languages | Tests |
|---|---|---|---|
| WordPress Jev Comment Triage | Comment text, link count and, by default, author name, email and URL; the jct_include_author_details filter turns the author details off | Interface translations only; no handling of the comment's language | Yes (Pest unit tests) |
| Jev-Moderation-Bot | Message text, author ID, account age, channel name and ID, up to 5 earlier false flags as "safe precedents" | Not handled | Yes |
| soter | Message text, account creation date, server join date and the author's recent messages | Not handled | Yes |
| Telegram anti-spam bot | Message text, links, forwarded flag, link previews, the sender's profile (bio or personal channel) when available and recent messages | Prompts mention other languages; one Chinese-language test case | Yes, including live tests |
| mastra-jev-moderation | The last user message only; the message is never logged | Not handled | None in the repository |
| jev-demo-moderator | Comment text only | Chinese and English samples | None |
Author name, email, account age and message history are personal data. What each route then does with what you send is on what data leaves your system, by route. On languages, TypeSafe's models page says "English is the primary training language … Other languages, including CJK scripts, are handled but not equally well" (limitations, row 13).
Where each project draws its lines
What happens to a post at each probability from 0 to 1 (main question of each project)
- WordPress (spam Noul)0.45–0.75 held for review; ≥ 0.75 spam folder
- soter (hate Noul)> 0.5 review log, no action; > 0.8 deleted
- Jev-Moderation-Bot (threat or ban)≥ 0.70 deleted (≥ 0.95 also ban review)
- mastra (block Noul)≥ 0.7 the agent's turn is refused
- Telegram bot (highest of 29 Nouls)≥ 0.81 deleted (operator may set 0.5–1)
- Demo (spam Noul)Jev's own Choice decides; ≥ 0.9 turns "publish" into "review"
- Left up or passed
- Held or logged for a person
- Spam folder (recoverable in WordPress)
- Deleted, or the turn refused
high_spam). Thresholds are code defaults. How to set your own: confidence thresholds. Documented 2 Oct 2026.What the six rows add up to
Six moderation projects read in source, one chip each
- The Jev call failswhat happens to the post
- fail-open: left upbrainstormity
- fail-open: left upsoter
- fail-open: keptTelegram bot
- fail-open: let throughmastra
- fail-closed: held for a personWordPress
- no-decisiondemo
- Removal below certaintydeletes without a review band
- From 0.70brainstormity
high_spam, any probabilitysoter- From 0.81Telegram bot
- Spam folder; middle band heldWordPress
- Refuses the turn at 0.7mastra
- Never deletesdemo
- Undoing a removalin the project's code
- Pardon, not restoredbrainstormity
- Pardon, not restoredsoter
- WordPress queue and spam folderWordPress
- NoneTelegram bot
- Nonemastra
- Nonedemo
- Author data sent by defaultbeyond the text
- Name, email, URLWordPress
- ID, account agebrainstormity
- Dates, recent messagessoter
- Profile, recent messagesTelegram bot
- Last message onlymastra
- Text onlydemo
One Noul per harm, or one Choice over categories?
Jev can answer a harm question in two shapes. TypeSafe's docs describe how each behaves; the table puts those statements side by side. Documented TypeSafe docs (primitives, Noul, Choice) read 2 Oct 2026, 05:26 UTC
| Question | One Noul per harm | One Choice over categories |
|---|---|---|
| What comes back | One probability per harm. "There is no separate confidence value for a Noul" | One label, a probability per option and a confidence. "The sum of all values is 1." |
| Can a post be two harms at once? | Yes: "Every question in a request sees the same state, is evaluated independently" | No: the options share one distribution, so a post that is both spam and harassment splits its probability between them |
| Thresholds | One per harm, set separately: "one question per condition, and the code decides what the combination means" | One confidence gate for the whole decision (method) |
| "None of these" | Implicit: all Nouls low | Must be an option. TypeSafe advises "an other or none of the above option when the list might not cover everything" |
| Fits | Independent harms with different costs of a mistake (a scam versus mild rudeness) | Choices that exclude each other (which queue, which action) |
| In the six projects | The Telegram bot (29 Nouls); WordPress and soter mix Nouls with a Score or Choice | brainstormity and soter use a Choice for spam level; the demo lets a Choice pick the action |
Does TypeSafe's policy say anything about moderation decisions?
Not in the texts read. The Acceptable Use Policy (last updated 23 Sep 2026; sections 1–4 read) forbids using the services to do things such as "1.6 transmit spam or other unsolicited communications to others" or "1.7 engage in, promote, support, or facilitate obscene, defamatory, hateful … activities". No clause addresses classifying harmful content or making automated moderation decisions about people. The Terms (19 Sep 2026) and the Master Customer Agreement (23 Sep 2026) were searched for "automated decision", "decisions about", "human review" and "high-risk": no clause found. Documented dates rechecked 2 Oct 2026, 05:14 UTC. A reading of the text, not legal advice.
How accurate is Jev at moderation?
No published moderation accuracy for Jev was found. Searched: evals.typesafe.ai, the WorkflowEvals README at commit 0ac3b8a (no "moderation", "toxic", "safety" or "content policy" text) and TypeSafe's cookbook list (2 Oct 2026, 05:27 UTC). The benchmarks ledger has no moderation row.
The closest figure measures repeatability, not accuracy. TypeSafe's cookbook "Self-consistency: choices" runs an 8-question moderation rubric 15 times on one borderline post. It reports 90.8% label agreement, rising to 99.2% "with automatic labels on 74.2% of answers" when answers below a top probability of 0.60 go to human review. Reported TypeSafe's own test. Its explicit "uncertain → human review" outcome is the step the four projects that leave posts up on failure do not have.
Lead, not read: KsanaDock/verdict-lab (Jev compared with LLMs on moderation cost and capability). Unverified
Published advice compared with the code
opentweet's "Content Moderation With Jev: Rules, Thresholds, Cost" ("Last updated: September 2026") recommends "Two cut points per rule, not one. The band between them is the human queue." Reported In the six projects, 2 have a review band (WordPress's hold band, soter's review log); brainstormity's second band deletes; the Telegram bot and mastra use one cut point. The same page explains Noul against Choice in line with TypeSafe's docs, but has no text on what to do when Jev errors or times out (0 matches for "unavailable", "outage", "retry", "fail open" or "fail closed"), which is the case where 4 of 6 projects leave the post up. The jev101 and befailproof moderation pages returned 404 at the URLs tried.
Eight rules for moderating with Jev, and which projects follow them
- Decide what happens to a post when Jev fails, before launch.WordPress holdsbrainstormity warns after 3soter, Telegram, mastra: post passes
Hold it if a wrong "publish" is costly; keep it up only if a missed harm is cheap. Log and count every failure. The shared failure rule for queues is on support-ticket triage.
- Hold before publishing where you can.WordPressDiscord, Telegram: after
When a bot deletes after publication, every Jev outage publishes everything.
- Give the middle band to a person, not to deletion.WordPresssoter: log onlybrainstormity, Telegram: delete
Use two cut points per harm and tune them on your own data (procedure).
- One Noul per harm when harms can occur together.Telegram: 29 NoulsWordPress, soter: mixed
Use one Choice only for routing between options that exclude each other, and include a "none of these" option.
- Make removals reversible.WordPress: spam folderbrainstormity, soter: pardon onlyTelegram: none
Keep the removed text, or use the platform's spam folder, so a pardon can restore the post and not only lift the punishment.
- Send the least author data you need.mastra, demoFour send author data
Name, email, URL, account age and message history are personal data (data by route). WordPress lets sites switch author details off.
- Treat a missing or malformed answer as a failure, not as 0.0 "safe".demo: errorFive: pass or approve
Validate the answer before acting on it.
- Test non-English posts separately.Telegram, demo: some samples
Jev reads text only, and TypeSafe says other languages "are handled but not equally well" (limitations).
The rules are a design reading of the source above, not a measured result. Cost per 1,000 posts depends on your token counts: use the cost per decision worksheet.
How the six were chosen
A GitHub repository search for "jev moderation" on 2 Oct 2026 returned 36 results (an index count, not adoption). From it, the projects with the most stars that have a Jev call path were read: brainstormity/Jev-Moderation-Bot (47 stars), CodeAlive-AI/mastra-jev-moderation (7), frolleks/soter (4) and soderlind/jev-comment-triage (3). Added: the planned lead Zafer-Liu/jev-demo-moderator (1) and the top Telegram project from a "jev spam" search, backmeupplz/jev_antispam_bot (14). The WordPress plugin was first read on 28 Sep from its moving main branch; it is now pinned at d848a5b (last commit 18 Sep, no newer version). Each project was pinned to one commit and read from source; nothing was installed or run and no Jev call was made. Other moderation repositories found by search, Jevdit (Show HN, 27 Sep) and the project behind Twittesia issue #261 stay leads.
Moderation demo videos: Mikey No Code, moderation build at 12:06 and Kitson Workshop episode 14 (not watched for this page); more in the moderation videos group.
What was not verified
- How any of the six behaves live: none was run, and none is verified as operating.
- Whether the WordPress AI Provider for Jev validates answers before returning them, and WordPress core's restore-from-spam flow (not read).
- The Jevdit project and the project behind Twittesia issue #261 (not located).
- Jev's moderation accuracy on any dataset: none is published, and the pool ran none.
- How each project handles non-English content beyond what is listed.
- How many people use any of these projects.
Sources and check times (2 Oct 2026, UTC)
- Projects at pinned commits (05:23): WordPress plugin at d848a5b (
jev-comment-triage.phpL43–52, L90–169, L259–280, L357–421); Jev-Moderation-Bot at 325ea7f (moderator.pyL25–56, L412–520;database.py); soter at 756579a (utils/jev.ts,index.tsL52–125,commands/pardon.ts); jev_antispam_bot at 0b07e88 (src/spam.ts,src/bot.tsL186–246,src/config.ts); mastra-jev-moderation at 8735ea8 (L14–24, L296–379); jev-demo-moderator at cc47b25 (moderator.py,server.pyL128–200). - TypeSafe docs (05:26): primitives, Noul, Choice, confidence, consistency cookbook.
- Policy (dates 05:14): Acceptable Use Policy, Terms, MCA; clause text from the pool's 30 Sep capture.
- Evidence search (05:27): evals.typesafe.ai; WorkflowEvals README at 0ac3b8a. Advice page (05:27): opentweet.
- Selection: GitHub repository searches "jev moderation", "jev moderate", "jev toxicity" and "jev spam", sorted by stars (05:23). Query wording: publisher titles listed in the pool's audience research of 28 Sep; Bing, DuckDuckGo, Google and YouTube suggested no moderation phrase for "jev moderation" (05:15).