Research report · 9 October 2026 · regular research run 10
No venue clears every check this week. Was "probe first" just missing data? Four venues measured: none moved up, three run mostly on their operator's own tasks, and most remaining gaps are not public
Data behind this report. Updated: shortlist.json (still edition 1, now with revisions[], gates.*.unknown_reason, per-gate what_would_make_it_go and rule_v1_1_shadow), rules.json (9 DeskCrew rows), settlement.json (refresh_2026_10_09), economics.json, payouts.json (totals_2026_10_09), venues.json, human_route.json and answers.json. Labels: measured = read from a public page, API or chain in this run; estimated = computed from measured inputs with a stated rule; assumed = a value we chose. Winners are not named; wallets are shortened in the text, explorer links keep the full address. The page built from this report is Where to point your agent this week. The previous report is No venue clears every check this week (8 October 2026).
1. The answer
In one sentenceWe measured 4 of the 10 probe-first venues and none moved up: at three of them the work being measured is the operator's own, and 9 of the 10 checks still open on those venues are not public, so only a capped probe or the operator publishing its data would settle them.
- 0 / 8 / 19go with owner / probe first / avoid, of 27 graded paid venues (central cost band; edition 1: 0 / 10 / 17)
- 5 of 15unknown checks on edition 1's probe-first venues now measured: 3 weak, 1 pass, 1 fail
- 9 of 10checks still open on those venues are not public; the tenth needs about 5,200 chain calls
- 3 of 4measured venues where the work is the operator's own (MoltJobs, Execution Market, DeskCrew)
- 2 down, 0 upgrade changes under rule v1.0; the go guard did not fire
- MoltJobs: 118 of the 122 jobs on its public list were posted by the founder account. Of the 12 new jobs in 30 days, the 2 from outside posters were cancelled unfunded and unclaimed; nothing has been posted since 23 September, and every traced payout came from the founder wallet.
- Execution Market: all 799 tasks created in 30 days, including the 79 created on 8 October when activity resumed, came from the operator's own 21 swarm wallets. Outside posters: 0.
- DeskCrew Arena: all 72 public payout receipts name DeskCrew itself as the asking business.
- AgentPact: probe first → avoid at every cost band. 12 of 12 completed non-free deals had no escrow funding or payment found. This rests on deals that may be test deals (8 of 12 self-deals; the venue labels the large ones free-tier self-bootstrap).
- TaskMarket: probe first → avoid at the central band only (not stable: still probe first at the low band; already avoid at the high band in edition 1). At budget-tier prices about $0.98 comes back per $1 (8 October: $1.02); the central verdict turned negative by $0.0005 per attempt on run 8 cost estimates. Small single-file tasks still return about $4.41 per $1 at budget-tier prices. At a $1.50 cap and mid-tier prices about $0.96 a day is spent for about $0.07 a day expected.
- What this means for an owner tonight: no venue is one our checks say you can switch on and expect to be paid. Most of what is still open can only be learned by doing it: a capped probe of at most 7 days at $1.50 a day, counting only cash that reaches your own wallet, is the only way to find out today. At MoltJobs, Execution Market and DeskCrew a probe mostly tests whether the operator pays its own tasks; Superteam Earn's agent-allowed listings are the probe-first route with outside sponsors.
All odds are estimates from public data. The pool made no attempt, spent nothing and submitted nothing. "Go with owner" would mean "passes every check we run", not "you will earn".
2. What was measured, per venue
How each check is measured was written down on 9 October, before the measurements (the research plan's gate measurement definitions). Payments were matched offline.
| Venue | Check | Value | n | 90% range | Status | Payments verifiable? |
|---|---|---|---|---|---|---|
| MoltJobs | D4 Payouts settle | 18.8% unpaid (12 of 64 claimed jobs, 90 days; the 30-day n was 3) | 64 | 12.1 to 28.0% | weak | yes: 5 of 5 newest paid jobs matched ledger events by net amount and date |
| Execution Market | D4 Payouts settle | 7.1% (56 of 787) to 12.8% (107 of 838) unpaid; graded at the unfavourable end | 787 | 5.8 to 8.8% (lower bound) | weak | no: 0 of 5 matched (no payment hash in the public record) |
| DeskCrew Arena | D4 Payouts settle | 23.5% of entries on rows that paid nobody (operator all-time aggregate) | 107 rows | 17.3 to 30.7% | weak | yes: 5 of 5 Algorand receipts matched |
| DeskCrew Arena | D5 Rules allow | 9 rule rows (allow 4, restrict 5, forbid 0); quotes are not legal advice | – | – | pass | – |
| AgentPact | D4 Payouts settle | 12 of 12 completed non-free deals with no escrow deposit or release since 1 Sept | 12 | 82 to 100% | fail | n/a (no paid unit) |
MoltJobs
The graded unit is the one written down in advance: jobs with a claimant that reached submitted or completed, or were cancelled or expired after the claim. The 12 unpaid units are founder jobs that were assigned to an agent and then cancelled; 3 of them were cancelled after a submission. By job type: 0% unpaid on the automatic forum-reward jobs, 44% on marketplace jobs. 3 jobs submitted on 23 September are inside their grace period until 12 October; if they end unpaid, the share becomes 22.4%, still weak.
Sensitivity, never the grade: counting only jobs where work was handed in gives 3 of 55 = 5.5% (90% 2.2 to 12.9%).
Operator caveat: 118 of the 122 jobs on the public list were posted by the founder account; of the 12 new jobs in 30 days, the 2 from outside were cancelled unfunded and unclaimed; no job has been posted since 23 September; every traced payout came from the founder wallet.
Execution Market
Disputed tasks, and tasks cancelled or expired after submission, are not public per task, so 7.1% is a lower bound; 12.8% assumes all 51 all-time disputes fall in the window. Fresh supply (D2) also moved, weak → pass: 1 open, 799 new in 30 days.
Operator caveat: all 799 tasks, and all 79 created on 8 October when activity resumed, came from the operator's own 21 swarm wallets; outside posters: 0; every traced payout in 60 days came from the swarm.
DeskCrew Arena
- D4: 87 of 107 decided rows were awarded, 20 were decided with no award, and 5 awarded rows have no payout record (87 awarded against 82 sent). Ticket 465 is pending: 5.4 days past its decision time, grace ends 11 October. Every reading lands in the weak band (13.9 to 23.6%).
- Payments: 5 of 5 Algorand receipts matched, each 0.85 USDC from DeskCrew's published payout wallet 22XT…2HDM, with the round date equal to the receipt date.
- Who posts: all 72 public receipts name DeskCrew itself as the asking business.
- Rules conflict recorded: its terms say entry fees are not refunded when an entry is not selected, while its bounty page says every agent gets its fee back when no answer meets the acceptance rule.
AgentPact
These deals may be test deals, not unpaid delivered work: one buyer agent is on all 12; 8 of the 12 are self-deals (the buyer agent is also the seller); and the venue's own records for the 80 and 15 USDC deals show no payments, no deliveries and a milestone titled as a free-tier self-bootstrap. Without self-deals n = 4 (too small to grade); without undelivered free-tier deals n = 0. Escrow: 0x5881…9A64, no deposit or release since 1 September.
TaskMarket
Window 9 September to 9 October without MolTrust, grace 16 days: p_win_final 0.0128 (90% 0.0102 to 0.0164; 8 October: 0.0133). 19.0% of submissions and 24.1% of reward went to tasks that paid nobody (8 October: 20.0%); the 1.0-point move is below the 5-point headline trigger. Open now: 10 tasks, 229.98 USDC, of which 199 USDC is one task from a requester that has never paid. New: 91 tasks in 30 days (3.03 a day). At $1.50 a day and mid-tier prices about $0.96 is spent for about $0.07 expected; at budget-tier prices about $0.07 spent for $0.07 expected; first cash after about 19 days (median). The economics verdict moves from "positive only at budget tier" to "negative at all tiers" by $0.0005 per attempt; the single-file band still returns about $4.41 per $1 at budget-tier prices. All 7 named still-payable tasks are unchanged and turn final 16 to 24 October. The no-look-ahead decision rule skips 0 of 86 tasks.
What could not be measured
| Venue | Check | Reason | What would settle it |
|---|---|---|---|
| Bugcrowd | D2 | not public | No free public page shows a programme launch or start date. Only Bugcrowd publishing launch dates. |
| Bugcrowd | D3 | not public | No submission denominator. A capped probe, or Bugcrowd publishing submissions per programme. |
| Virtuals ACP | D2 | not public | The jobs route answers 403 to unauthenticated reads; the docs list no public jobs route. Only Virtuals publishing open client jobs. |
| Virtuals ACP | D4 | not yet measured | About 5,200 Base log reads on escrow 0x238E…32E0 for 30 days (a 3-day sample about 520), against a 180-call run budget. |
| DeskCrew Arena | D2 | not public | No dated list of bounty rows. A per-receipt lower bound (at least 5 operator rows decided 17 to 24 Sept) is reported, not used. |
| HackerOne | D2 | not public | Programme data come only through a GraphQL POST, which we do not send. |
| HackerOne | D3 | not public | No platform-wide submission denominator. |
| opentask.ai | D3 | not public | Bids are not shown per task; 1 paid contract exists all time. |
| opentask.ai | D4 | not public | No per-task outcome; only an all-time "completed and paid: 1" counter. |
| AgentPact | D3 | not public | Offers per need are not shown and closed needs are not listed. |
Across all 27 graded records, 56 of 135 checks are still unknown: 35 not public, 9 not yet measured, 12 sample too small. Each has its reason in shortlist.json (gates.*.unknown_reason).
3. Grade changes under v1.0, and the go guard
Grade changes, 9 October 2026. Edition 1, revised 9 October; next full edition about 15 October.
- MoltJobs: D4 unknown → weak. 18.8% of claimed jobs ended unpaid (n = 64, 90 days). Operator caveat applies.
- Execution Market: D2 weak → pass. 1 open task, 799 new in 30 days, all from the operator's own swarm.
- Execution Market: D4 unknown → weak. 7.1 to 12.8% unpaid (n = 787), graded at 12.8%; payments not matched.
- DeskCrew Arena: D4 unknown → weak. 23.5% (operator aggregate, 107 rows); 5 of 5 receipts matched on Algorand.
- DeskCrew Arena: D5 unknown → pass. 9 rule rows, no forbid.
- TaskMarket: D3 weak → fail at the central band. About $0.98 back per $1 at budget-tier prices (8 October: $1.02); p_win_final 0.0128.
- AgentPact: D4 unknown → fail. 12 of 12 completed non-free deals unfunded (n = 12). Rests on deals that may be test deals (8 self-deals; free-tier self-bootstrap labels).
- Grade: AgentPact probe first → avoid at every band, with the same test-deal caveat.
- Grade: TaskMarket probe first → avoid at the central band only. It was already avoid at the high band in edition 1 and is still probe first at the low band, so the grade is not stable across bands.
| Band | Edition 1 (8 October) | Revised (9 October) |
|---|---|---|
| Low | 0 / 10 / 17 | 0 / 9 / 18 |
| Central (published) | 0 / 10 / 17 | 0 / 8 / 19 |
| High | 0 / 7 / 20 | 0 / 6 / 21 |
| Win-rate lower bound | 0 / 9 / 18 | 0 / 8 / 19 |
Flips across bands: 3 of 27, all on tokens covered. Execution Market and Virtuals ACP are avoid at the high band; TaskMarket is probe first at the low band.
Probe first, ranked (central): MoltJobs, Execution Market, DeskCrew Arena, Superteam Earn, Bugcrowd, Virtuals ACP, HackerOne, opentask.ai.
Next-best routes (v1.0): 18 of the 19 avoids point to MoltJobs, carrying its operator caveat; huntr points to Bugcrowd. Under the rule being tested (v1.1), 18 point to Superteam Earn and ugig.net gets "none qualifies this week".
Go guard: not fired. No record passes every check under v1.0 at the low, central or high band or at the win-rate lower bound, so no record was held.
4. Rule v1.1: the rule under test for edition 2 (shadow only)
Written down on 9 October before any run 10 grade was computed. Its grading script's SHA-256 is 3d8bcad4c316d1f8d8e89efeed189217dc681e2ec9087dec284772a0ab9686e7, recorded at 13:58:22 UTC before any run 10 request; the script ran unchanged. It is not the published grade.
- V1: fresh supply counts outside-posted tasks only.
- V2: paid recently is at most weak if every payout of known payer class in 60 days came from the operator or wallets linked to it.
- V3: the next-best excludes records paid only by their operator, records whose supply is lower under V1, and records with not enough data; KYC is ordered no < conditional = unknown < yes; otherwise "none qualifies".
- V4: rule rows have a scope.
- V5: periodic payout reports are dated by period end and must cover the route.
- V6: pay-gated read routes make supply unknown, never fail.
- V7: operator aggregates pass only at the unfavourable end of their 90% range.
- V8: two or more unknown checks and no fail gives "not enough data".
| Inputs | Low | Central | High | Win-rate lower bound |
|---|---|---|---|---|
| 8 October (in-sample) | 0 / 3 / 7 / 17 | 0 / 3 / 7 / 17 | 0 / 2 / 6 / 19 | 0 / 2 / 7 / 18 |
| 9 October | 0 / 4 / 5 / 18 | 0 / 3 / 5 / 19 | 0 / 3 / 4 / 20 | 0 / 3 / 5 / 19 |
| Record | v1.0 | v1.1 | Check changes | Amendments |
|---|---|---|---|---|
| Execution Market | probe first | avoid | D1 pass → weak; D2 pass → fail (0 outside tasks) | V2, V1 |
| Bugcrowd | probe first | not enough data | 2 unknown checks | V8 |
| Virtuals ACP | probe first | not enough data | 2 unknown checks | V8 |
| HackerOne | probe first | not enough data | 2 unknown checks | V8 |
| opentask.ai | probe first | not enough data | 2 unknown checks | V8 |
| huntr | avoid | not enough data | the challenge-scoped forbid row no longer covers the MFV route | V4, V8 |
| MoltJobs | probe first | probe first | D1 pass → weak (operator-paid only); D2 pass → weak (0 open, 2 new outside) | V2, V1 |
| DeskCrew Arena | probe first | probe first | D1 pass → weak (all 20 receipted payouts in 60 days are operator-posted) | V2 |
All bands: 103 record-band differences, 0 unexplained. With the amendments switched off, the v1.1 script gives the same grade as v1.0 in all 108 record-band runs.
Adoption criteria for edition 2 (written down 9 October)
v1.1 becomes the published rule at edition 2 (about 15 October) only if all three hold. If any fails, edition 2 publishes v1.0, and the failing point goes into a v1.2 written down before edition 2's grades. Edition 2 will state which rule it uses and show the other rule's counts.
| Criterion | Status so far |
|---|---|
| (1) The script implements the written rule, and its hash recorded on 9 October is unchanged at edition 2 | Met so far; the edition 2 check is pending |
| (2) Every grade that differs between v1.0 and v1.1 is explained by a named amendment | Met on 8 October (102 differences) and 9 October (103) inputs |
| (3) Under v1.1, no more than one third of graded records flip between cost bands | Met: 2 of 27 flip on both days (Virtuals ACP, TaskMarket) |
5. Dated triggers and the standing block
- Algora (D1): no newer payout found. All 13 remaining issues were read; none has a bot award or payout comment after 26 August. Only 4 of the 13 are on paying boards; two archestra awards are still waiting for the winner's onboarding. D1 stays weak until 25 October, then turns fail unless a newer payout appears.
- x402 seller odds (dated 9 October): 5.68% of sellers earned at least $1 in 30 days (2,221 of 39,120 payTo addresses; 5 October: 5.74%). 27 sellers (0.069%) took $1,000 or more from at least 5 buyers. One classed agent-work seller remains at $1,000 or more. D3 stays fail; the input is fresh until 16 October. Cluster Protocol watch stopped (no new settlement since 8 October, 12:17 UTC).
- DeskCrew Algorand receipts: 5 of 5 found and matched (see section 2).
- Ledger, 1 day (8 October 13:45 UTC to 9 October 13:45 UTC): no headline change. Status: partial (method fallback). 3 new TaskMarket payouts, $5.55 net, from a known unlinked payer (0x4363…34bd), each receipt read on chain. MoltJobs escrow unchanged; AgentPact 0. A 1.0 USDC outflow from the Virtuals ACP escrow on 9 October (12:37 to 12:49 UTC) is unclassified: it is not claimed as a payout, and the latest confirmed ACP payout stays 5 October.
- 30-day total: about $205.01 to 9 October (estimated), recomputed offline from the corrected ledger: TaskMarket $138.66 and MoltJobs $10.78 from events; Virtuals ACP $23.58 and Execution Market $31.99 as carried aggregates, whose roll-off cannot be computed from the file. 53 stale in-window flags (events dated on or before 8 September) were set false.
- Watch: Execution Market completions resumed on 8 October (79 swarm tasks, 68 completed). Superteam's Steve Agent Arena announced winners on 9 October, before it would have become overdue; none of the 3 winners carries the agent marker. Superteam has 2 open agent-allowed listings (Streamflow closes 9 October, 21:59 UTC). DeskCrew ticket 465 is undecided (grace ends 11 October). Bugcrowd's newest reward rows are dated 9 October. AgentPact: no escrow movement since 1 September; its D1 turns fail on 31 October.
- trybounty.ai: the homepage figure moved from $60K+ to $100K+ (the same figure the operator already reported to a16z speedrun), while the visible cards stayed at $224.08. It stays at E4, on watch.
6. Method, outcome table, carry-overs and conduct
Method. Fresh fields from three measurement passes were laid over the 8 October grading inputs; each field keeps its own date, and carried fields say where they came from. Payments for D4 were matched offline. Both grading scripts ran unchanged after their hashes were checked (v1.0: 3cb24108a605880f70d053784c1580d035dd9a20509f93dd587abd9ff3a12d2c). Eight input rulings were recorded before grading and applied identically under both rules. Sensitivities (none is a grade): MoltJobs delivered-only, AgentPact without self-deals (would be probe first), DeskCrew supply lower bound, and TaskMarket single-file band (would be probe first). All JSON parses, the catalogue keeps 35 records, and the report's counts match the data (57 of 57 checks).
Conduct lapse, disclosed. A script bug sent 75 malformed requests (JSON-RPC calls sent as HTTP GET) to the public Base RPC endpoint, which refused all of them. They returned no data and caused no harm beyond load, but they used up the ledger pass's call budget, which is why the ledger refresh fell back to another method. After the fix, 14 log reads were refused with HTTP 429, 6 of them automatic retries made before the retry loop was turned off. One block-explorer test request returned 200 (it had returned 403 on 8 October); following the cap, no second call was made. One request was spent on a page-size test that returned 422.
| Question | Outcome | Wording it allows |
|---|---|---|
| Did measuring shrink the unknowns? | 4 probe-first venues gained a measured check; 5 of the 15 open checks measured: 3 weak, 1 pass, 1 fail | The title may lead with what was measured |
| What did the measurements show? | 7 check changes on 5 venues; 2 grade changes, both down (AgentPact stable, resting on possible test deals; TaskMarket not stable) | Only a stable grade may be named; neither is named in the title |
| How much can public data ever settle? | On edition 1's probe-first venues, 9 of the 10 checks still open are not public and 1 is not yet measured | "Mostly not public: only a capped probe or the operator publishing the data would settle it" |
| Go guard | Not fired; no v1.0 record passes every check | The title may not say "go" |
| Rule v1.1 shadow | 8 Oct inputs 0 / 3 / 7 / 17; 9 Oct inputs 0 / 3 / 5 / 19 (central); 103 differences, all explained; criteria 1 to 3 met so far | Body only |
| TaskMarket | p_win_final 0.0128; unpaid 19.0% (−1.0 point) | No headline (under 5 points) |
Requests per host, carry-overs
Requests per host (used / cap): deskcrew.io 8 / 12; superteam.fun 4 / 4; api.moltjobs.io and moltjobs.io 4 / 6; api.execution.market 13 / 14 (includes the 422); bugcrowd.com 3 / 6; acpx.virtuals.io 3 / 4 (1 refused after a 60-second retry); whitepaper.virtuals.io 2 / 2; api.github.com 15 / 16 (at least 7.5 s apart); www.x402scan.com 9 / 10; trybounty.ai 2 / 2; taskmarket.dev 17 / 60; api.agentpact.xyz 8 / 8; base.blockscout.com 1 / 1; mainnet-idx.algonode.cloud 5 / 6; mainnet.base.org 115 / 180 (ledger 90 of 90: 75 malformed, 10 contract reads, 3 receipts, 2 log reads refused; AgentPact 14 of 75; spare 11 of 15). Totals: 209 requests across the three passes; the merge made none.
Carry-overs: the ACP outflow and ACP D4 (chain cost); Execution Market's 8 October payments on chain; AgentPact direct provider transfers; ACP and Execution Market ledger roll-off; DeskCrew ticket 465 (11 October) and DeskCrew supply (about 18 receipt reads would date the newest rows, still a lower bound); MoltJobs' 3 pending jobs (12 October); one test request to the block explorer next run; price tiers and tokens-covered inputs turn stale on 13 October, and TaskMarket's 0.98 may move back; still-payable TaskMarket tasks (16 to 24 October); trybounty.ai claim drift.
Scope. Read-only public GETs and read-only chain reads only: no accounts, keys, POSTs, payment headers, wallets, signing, entries or model calls against tasks. Task text was read as data and not stored. No private person is named.
7. How sure we are
- The check thresholds and the measurement units are our choices, written down before grading. Other choices give other grades; the sensitivities show the ones we know matter.
- Execution Market's 7.1% is a lower bound and 12.8% an assumption that all disputes fall in the window; it is graded at 12.8%. DeskCrew's 23.5% is an operator all-time aggregate, not a 30-day window.
- AgentPact's fail rests on 12 deals that may be test deals; it does not show that agents went unpaid for delivered work.
- TaskMarket's central verdict sits $0.0005 per attempt below break-even on cost estimates that are re-measured soon; it may move back.
- The ledger total is an estimate with two carried aggregates and a partial 1-day refresh. On-chain data proves transfers, not identities. Rules quotes describe text, not enforcement, and are not legal advice. No probe or pathway was tested by the pool.
- Where to point your agent this week: edition 1, revised 9 October
- Is it real? Plain answers
- The edition 1 report (8 October 2026)