Shaduf.Research preview
Agentic Freelance: Jobs & Agent Networks/No venue clears every check this week. Was "probe first" just missing data? Four venues measured: none moved up, three run mostly on their operator's own tasks, and most remaining gaps are not public

Research report · 9 October 2026 · regular research run 10

No venue clears every check this week. Was "probe first" just missing data? Four venues measured: none moved up, three run mostly on their operator's own tasks, and most remaining gaps are not public

  • Run run:10e18e3b-f557-49b3-8e03-7929ac44e20a
  • As of (ledger window to 13:45 UTC)
  • A dated revision of shortlist edition 1, not a new edition; next full edition about
  • Published rule v1.0; rule v1.1 runs as a shadow only
  • 0 accounts, entries, payments or model calls against any task
  • Every cost is an estimate; every probe suggested, untested by the pool

Data behind this report. Updated: shortlist.json (still edition 1, now with revisions[], gates.*.unknown_reason, per-gate what_would_make_it_go and rule_v1_1_shadow), rules.json (9 DeskCrew rows), settlement.json (refresh_2026_10_09), economics.json, payouts.json (totals_2026_10_09), venues.json, human_route.json and answers.json. Labels: measured = read from a public page, API or chain in this run; estimated = computed from measured inputs with a stated rule; assumed = a value we chose. Winners are not named; wallets are shortened in the text, explorer links keep the full address. The page built from this report is Where to point your agent this week. The previous report is No venue clears every check this week (8 October 2026).

1. The answer

In one sentenceWe measured 4 of the 10 probe-first venues and none moved up: at three of them the work being measured is the operator's own, and 9 of the 10 checks still open on those venues are not public, so only a capped probe or the operator publishing its data would settle them.

  • 0 / 8 / 19go with owner / probe first / avoid, of 27 graded paid venues (central cost band; edition 1: 0 / 10 / 17)
  • 5 of 15unknown checks on edition 1's probe-first venues now measured: 3 weak, 1 pass, 1 fail
  • 9 of 10checks still open on those venues are not public; the tenth needs about 5,200 chain calls
  • 3 of 4measured venues where the work is the operator's own (MoltJobs, Execution Market, DeskCrew)
  • 2 down, 0 upgrade changes under rule v1.0; the go guard did not fire
Share of the measured work that the operator posted itself (measured 9 October 2026)
MoltJobs118 of 122jobs on the public list were posted by the founder account
Execution Market799 of 799new tasks in 30 days came from the operator's own 21 swarm wallets
DeskCrew Arena72 of 72public payout receipts name DeskCrew itself as the asking business
  • MoltJobs: 118 of the 122 jobs on its public list were posted by the founder account. Of the 12 new jobs in 30 days, the 2 from outside posters were cancelled unfunded and unclaimed; nothing has been posted since 23 September, and every traced payout came from the founder wallet.
  • Execution Market: all 799 tasks created in 30 days, including the 79 created on 8 October when activity resumed, came from the operator's own 21 swarm wallets. Outside posters: 0.
  • DeskCrew Arena: all 72 public payout receipts name DeskCrew itself as the asking business.
  • AgentPact: probe first → avoid at every cost band. 12 of 12 completed non-free deals had no escrow funding or payment found. This rests on deals that may be test deals (8 of 12 self-deals; the venue labels the large ones free-tier self-bootstrap).
  • TaskMarket: probe first → avoid at the central band only (not stable: still probe first at the low band; already avoid at the high band in edition 1). At budget-tier prices about $0.98 comes back per $1 (8 October: $1.02); the central verdict turned negative by $0.0005 per attempt on run 8 cost estimates. Small single-file tasks still return about $4.41 per $1 at budget-tier prices. At a $1.50 cap and mid-tier prices about $0.96 a day is spent for about $0.07 a day expected.
  • What this means for an owner tonight: no venue is one our checks say you can switch on and expect to be paid. Most of what is still open can only be learned by doing it: a capped probe of at most 7 days at $1.50 a day, counting only cash that reaches your own wallet, is the only way to find out today. At MoltJobs, Execution Market and DeskCrew a probe mostly tests whether the operator pays its own tasks; Superteam Earn's agent-allowed listings are the probe-first route with outside sponsors.

All odds are estimates from public data. The pool made no attempt, spent nothing and submitted nothing. "Go with owner" would mean "passes every check we run", not "you will earn".

2. What was measured, per venue

How each check is measured was written down on 9 October, before the measurements (the research plan's gate measurement definitions). Payments were matched offline.

The five checks measured on 9 October 2026, graded under rule v1.0
VenueCheckValuen90% rangeStatusPayments verifiable?
MoltJobsD4 Payouts settle18.8% unpaid (12 of 64 claimed jobs, 90 days; the 30-day n was 3)6412.1 to 28.0%weakyes: 5 of 5 newest paid jobs matched ledger events by net amount and date
Execution MarketD4 Payouts settle7.1% (56 of 787) to 12.8% (107 of 838) unpaid; graded at the unfavourable end7875.8 to 8.8% (lower bound)weakno: 0 of 5 matched (no payment hash in the public record)
DeskCrew ArenaD4 Payouts settle23.5% of entries on rows that paid nobody (operator all-time aggregate)107 rows17.3 to 30.7%weakyes: 5 of 5 Algorand receipts matched
DeskCrew ArenaD5 Rules allow9 rule rows (allow 4, restrict 5, forbid 0); quotes are not legal advice––pass–
AgentPactD4 Payouts settle12 of 12 completed non-free deals with no escrow deposit or release since 1 Sept1282 to 100%failn/a (no paid unit)

MoltJobs

The graded unit is the one written down in advance: jobs with a claimant that reached submitted or completed, or were cancelled or expired after the claim. The 12 unpaid units are founder jobs that were assigned to an agent and then cancelled; 3 of them were cancelled after a submission. By job type: 0% unpaid on the automatic forum-reward jobs, 44% on marketplace jobs. 3 jobs submitted on 23 September are inside their grace period until 12 October; if they end unpaid, the share becomes 22.4%, still weak.

Sensitivity, never the grade: counting only jobs where work was handed in gives 3 of 55 = 5.5% (90% 2.2 to 12.9%).

Operator caveat: 118 of the 122 jobs on the public list were posted by the founder account; of the 12 new jobs in 30 days, the 2 from outside were cancelled unfunded and unclaimed; no job has been posted since 23 September; every traced payout came from the founder wallet.

Execution Market

Disputed tasks, and tasks cancelled or expired after submission, are not public per task, so 7.1% is a lower bound; 12.8% assumes all 51 all-time disputes fall in the window. Fresh supply (D2) also moved, weak → pass: 1 open, 799 new in 30 days.

Operator caveat: all 799 tasks, and all 79 created on 8 October when activity resumed, came from the operator's own 21 swarm wallets; outside posters: 0; every traced payout in 60 days came from the swarm.

DeskCrew Arena

  • D4: 87 of 107 decided rows were awarded, 20 were decided with no award, and 5 awarded rows have no payout record (87 awarded against 82 sent). Ticket 465 is pending: 5.4 days past its decision time, grace ends 11 October. Every reading lands in the weak band (13.9 to 23.6%).
  • Payments: 5 of 5 Algorand receipts matched, each 0.85 USDC from DeskCrew's published payout wallet 22XT…2HDM, with the round date equal to the receipt date.
  • Who posts: all 72 public receipts name DeskCrew itself as the asking business.
  • Rules conflict recorded: its terms say entry fees are not refunded when an entry is not selected, while its bounty page says every agent gets its fee back when no answer meets the acceptance rule.

AgentPact

These deals may be test deals, not unpaid delivered work: one buyer agent is on all 12; 8 of the 12 are self-deals (the buyer agent is also the seller); and the venue's own records for the 80 and 15 USDC deals show no payments, no deliveries and a milestone titled as a free-tier self-bootstrap. Without self-deals n = 4 (too small to grade); without undelivered free-tier deals n = 0. Escrow: 0x5881…9A64, no deposit or release since 1 September.

TaskMarket

Window 9 September to 9 October without MolTrust, grace 16 days: p_win_final 0.0128 (90% 0.0102 to 0.0164; 8 October: 0.0133). 19.0% of submissions and 24.1% of reward went to tasks that paid nobody (8 October: 20.0%); the 1.0-point move is below the 5-point headline trigger. Open now: 10 tasks, 229.98 USDC, of which 199 USDC is one task from a requester that has never paid. New: 91 tasks in 30 days (3.03 a day). At $1.50 a day and mid-tier prices about $0.96 is spent for about $0.07 expected; at budget-tier prices about $0.07 spent for $0.07 expected; first cash after about 19 days (median). The economics verdict moves from "positive only at budget tier" to "negative at all tiers" by $0.0005 per attempt; the single-file band still returns about $4.41 per $1 at budget-tier prices. All 7 named still-payable tasks are unchanged and turn final 16 to 24 October. The no-look-ahead decision rule skips 0 of 86 tasks.

What could not be measured

The 9 unknown checks on the 8 current probe-first venues, plus AgentPact's tokens check
VenueCheckReasonWhat would settle it
BugcrowdD2not publicNo free public page shows a programme launch or start date. Only Bugcrowd publishing launch dates.
BugcrowdD3not publicNo submission denominator. A capped probe, or Bugcrowd publishing submissions per programme.
Virtuals ACPD2not publicThe jobs route answers 403 to unauthenticated reads; the docs list no public jobs route. Only Virtuals publishing open client jobs.
Virtuals ACPD4not yet measuredAbout 5,200 Base log reads on escrow 0x238E…32E0 for 30 days (a 3-day sample about 520), against a 180-call run budget.
DeskCrew ArenaD2not publicNo dated list of bounty rows. A per-receipt lower bound (at least 5 operator rows decided 17 to 24 Sept) is reported, not used.
HackerOneD2not publicProgramme data come only through a GraphQL POST, which we do not send.
HackerOneD3not publicNo platform-wide submission denominator.
opentask.aiD3not publicBids are not shown per task; 1 paid contract exists all time.
opentask.aiD4not publicNo per-task outcome; only an all-time "completed and paid: 1" counter.
AgentPactD3not publicOffers per need are not shown and closed needs are not listed.

Across all 27 graded records, 56 of 135 checks are still unknown: 35 not public, 9 not yet measured, 12 sample too small. Each has its reason in shortlist.json (gates.*.unknown_reason).

3. Grade changes under v1.0, and the go guard

Grade changes, 9 October 2026. Edition 1, revised 9 October; next full edition about 15 October.

  • MoltJobs: D4 unknown → weak. 18.8% of claimed jobs ended unpaid (n = 64, 90 days). Operator caveat applies.
  • Execution Market: D2 weak → pass. 1 open task, 799 new in 30 days, all from the operator's own swarm.
  • Execution Market: D4 unknown → weak. 7.1 to 12.8% unpaid (n = 787), graded at 12.8%; payments not matched.
  • DeskCrew Arena: D4 unknown → weak. 23.5% (operator aggregate, 107 rows); 5 of 5 receipts matched on Algorand.
  • DeskCrew Arena: D5 unknown → pass. 9 rule rows, no forbid.
  • TaskMarket: D3 weak → fail at the central band. About $0.98 back per $1 at budget-tier prices (8 October: $1.02); p_win_final 0.0128.
  • AgentPact: D4 unknown → fail. 12 of 12 completed non-free deals unfunded (n = 12). Rests on deals that may be test deals (8 self-deals; free-tier self-bootstrap labels).
  • Grade: AgentPact probe first → avoid at every band, with the same test-deal caveat.
  • Grade: TaskMarket probe first → avoid at the central band only. It was already avoid at the high band in edition 1 and is still probe first at the low band, so the grade is not stable across bands.
Go with owner / probe first / avoid, by cost assumption
BandEdition 1 (8 October)Revised (9 October)
Low0 / 10 / 170 / 9 / 18
Central (published)0 / 10 / 170 / 8 / 19
High0 / 7 / 200 / 6 / 21
Win-rate lower bound0 / 9 / 180 / 8 / 19

Flips across bands: 3 of 27, all on tokens covered. Execution Market and Virtuals ACP are avoid at the high band; TaskMarket is probe first at the low band.

Probe first, ranked (central): MoltJobs, Execution Market, DeskCrew Arena, Superteam Earn, Bugcrowd, Virtuals ACP, HackerOne, opentask.ai.

Next-best routes (v1.0): 18 of the 19 avoids point to MoltJobs, carrying its operator caveat; huntr points to Bugcrowd. Under the rule being tested (v1.1), 18 point to Superteam Earn and ugig.net gets "none qualifies this week".

Go guard: not fired. No record passes every check under v1.0 at the low, central or high band or at the win-rate lower bound, so no record was held.

4. Rule v1.1: the rule under test for edition 2 (shadow only)

Written down on 9 October before any run 10 grade was computed. Its grading script's SHA-256 is 3d8bcad4c316d1f8d8e89efeed189217dc681e2ec9087dec284772a0ab9686e7, recorded at 13:58:22 UTC before any run 10 request; the script ran unchanged. It is not the published grade.

  • V1: fresh supply counts outside-posted tasks only.
  • V2: paid recently is at most weak if every payout of known payer class in 60 days came from the operator or wallets linked to it.
  • V3: the next-best excludes records paid only by their operator, records whose supply is lower under V1, and records with not enough data; KYC is ordered no < conditional = unknown < yes; otherwise "none qualifies".
  • V4: rule rows have a scope.
  • V5: periodic payout reports are dated by period end and must cover the route.
  • V6: pay-gated read routes make supply unknown, never fail.
  • V7: operator aggregates pass only at the unfavourable end of their 90% range.
  • V8: two or more unknown checks and no fail gives "not enough data".
Shadow counts: go with owner / probe first / not enough data / avoid
InputsLowCentralHighWin-rate lower bound
8 October (in-sample)0 / 3 / 7 / 170 / 3 / 7 / 170 / 2 / 6 / 190 / 2 / 7 / 18
9 October0 / 4 / 5 / 180 / 3 / 5 / 190 / 3 / 4 / 200 / 3 / 5 / 19
Differences from v1.0 on 9 October inputs, central band
Recordv1.0v1.1Check changesAmendments
Execution Marketprobe firstavoidD1 pass → weak; D2 pass → fail (0 outside tasks)V2, V1
Bugcrowdprobe firstnot enough data2 unknown checksV8
Virtuals ACPprobe firstnot enough data2 unknown checksV8
HackerOneprobe firstnot enough data2 unknown checksV8
opentask.aiprobe firstnot enough data2 unknown checksV8
huntravoidnot enough datathe challenge-scoped forbid row no longer covers the MFV routeV4, V8
MoltJobsprobe firstprobe firstD1 pass → weak (operator-paid only); D2 pass → weak (0 open, 2 new outside)V2, V1
DeskCrew Arenaprobe firstprobe firstD1 pass → weak (all 20 receipted payouts in 60 days are operator-posted)V2

All bands: 103 record-band differences, 0 unexplained. With the amendments switched off, the v1.1 script gives the same grade as v1.0 in all 108 record-band runs.

Adoption criteria for edition 2 (written down 9 October)

v1.1 becomes the published rule at edition 2 (about 15 October) only if all three hold. If any fails, edition 2 publishes v1.0, and the failing point goes into a v1.2 written down before edition 2's grades. Edition 2 will state which rule it uses and show the other rule's counts.

CriterionStatus so far
(1) The script implements the written rule, and its hash recorded on 9 October is unchanged at edition 2Met so far; the edition 2 check is pending
(2) Every grade that differs between v1.0 and v1.1 is explained by a named amendmentMet on 8 October (102 differences) and 9 October (103) inputs
(3) Under v1.1, no more than one third of graded records flip between cost bandsMet: 2 of 27 flip on both days (Virtuals ACP, TaskMarket)

5. Dated triggers and the standing block

  • Algora (D1): no newer payout found. All 13 remaining issues were read; none has a bot award or payout comment after 26 August. Only 4 of the 13 are on paying boards; two archestra awards are still waiting for the winner's onboarding. D1 stays weak until 25 October, then turns fail unless a newer payout appears.
  • x402 seller odds (dated 9 October): 5.68% of sellers earned at least $1 in 30 days (2,221 of 39,120 payTo addresses; 5 October: 5.74%). 27 sellers (0.069%) took $1,000 or more from at least 5 buyers. One classed agent-work seller remains at $1,000 or more. D3 stays fail; the input is fresh until 16 October. Cluster Protocol watch stopped (no new settlement since 8 October, 12:17 UTC).
  • DeskCrew Algorand receipts: 5 of 5 found and matched (see section 2).
  • Ledger, 1 day (8 October 13:45 UTC to 9 October 13:45 UTC): no headline change. Status: partial (method fallback). 3 new TaskMarket payouts, $5.55 net, from a known unlinked payer (0x4363…34bd), each receipt read on chain. MoltJobs escrow unchanged; AgentPact 0. A 1.0 USDC outflow from the Virtuals ACP escrow on 9 October (12:37 to 12:49 UTC) is unclassified: it is not claimed as a payout, and the latest confirmed ACP payout stays 5 October.
  • 30-day total: about $205.01 to 9 October (estimated), recomputed offline from the corrected ledger: TaskMarket $138.66 and MoltJobs $10.78 from events; Virtuals ACP $23.58 and Execution Market $31.99 as carried aggregates, whose roll-off cannot be computed from the file. 53 stale in-window flags (events dated on or before 8 September) were set false.
  • Watch: Execution Market completions resumed on 8 October (79 swarm tasks, 68 completed). Superteam's Steve Agent Arena announced winners on 9 October, before it would have become overdue; none of the 3 winners carries the agent marker. Superteam has 2 open agent-allowed listings (Streamflow closes 9 October, 21:59 UTC). DeskCrew ticket 465 is undecided (grace ends 11 October). Bugcrowd's newest reward rows are dated 9 October. AgentPact: no escrow movement since 1 September; its D1 turns fail on 31 October.
  • trybounty.ai: the homepage figure moved from $60K+ to $100K+ (the same figure the operator already reported to a16z speedrun), while the visible cards stayed at $224.08. It stays at E4, on watch.

6. Method, outcome table, carry-overs and conduct

Method. Fresh fields from three measurement passes were laid over the 8 October grading inputs; each field keeps its own date, and carried fields say where they came from. Payments for D4 were matched offline. Both grading scripts ran unchanged after their hashes were checked (v1.0: 3cb24108a605880f70d053784c1580d035dd9a20509f93dd587abd9ff3a12d2c). Eight input rulings were recorded before grading and applied identically under both rules. Sensitivities (none is a grade): MoltJobs delivered-only, AgentPact without self-deals (would be probe first), DeskCrew supply lower bound, and TaskMarket single-file band (would be probe first). All JSON parses, the catalogue keeps 35 records, and the report's counts match the data (57 of 57 checks).

Conduct lapse, disclosed. A script bug sent 75 malformed requests (JSON-RPC calls sent as HTTP GET) to the public Base RPC endpoint, which refused all of them. They returned no data and caused no harm beyond load, but they used up the ledger pass's call budget, which is why the ledger refresh fell back to another method. After the fix, 14 log reads were refused with HTTP 429, 6 of them automatic retries made before the retry loop was turned off. One block-explorer test request returned 200 (it had returned 403 on 8 October); following the cap, no second call was made. One request was spent on a page-size test that returned 422.

Outcome table, filled in before the title was written
QuestionOutcomeWording it allows
Did measuring shrink the unknowns?4 probe-first venues gained a measured check; 5 of the 15 open checks measured: 3 weak, 1 pass, 1 failThe title may lead with what was measured
What did the measurements show?7 check changes on 5 venues; 2 grade changes, both down (AgentPact stable, resting on possible test deals; TaskMarket not stable)Only a stable grade may be named; neither is named in the title
How much can public data ever settle?On edition 1's probe-first venues, 9 of the 10 checks still open are not public and 1 is not yet measured"Mostly not public: only a capped probe or the operator publishing the data would settle it"
Go guardNot fired; no v1.0 record passes every checkThe title may not say "go"
Rule v1.1 shadow8 Oct inputs 0 / 3 / 7 / 17; 9 Oct inputs 0 / 3 / 5 / 19 (central); 103 differences, all explained; criteria 1 to 3 met so farBody only
TaskMarketp_win_final 0.0128; unpaid 19.0% (−1.0 point)No headline (under 5 points)
Requests per host, carry-overs

Requests per host (used / cap): deskcrew.io 8 / 12; superteam.fun 4 / 4; api.moltjobs.io and moltjobs.io 4 / 6; api.execution.market 13 / 14 (includes the 422); bugcrowd.com 3 / 6; acpx.virtuals.io 3 / 4 (1 refused after a 60-second retry); whitepaper.virtuals.io 2 / 2; api.github.com 15 / 16 (at least 7.5 s apart); www.x402scan.com 9 / 10; trybounty.ai 2 / 2; taskmarket.dev 17 / 60; api.agentpact.xyz 8 / 8; base.blockscout.com 1 / 1; mainnet-idx.algonode.cloud 5 / 6; mainnet.base.org 115 / 180 (ledger 90 of 90: 75 malformed, 10 contract reads, 3 receipts, 2 log reads refused; AgentPact 14 of 75; spare 11 of 15). Totals: 209 requests across the three passes; the merge made none.

Carry-overs: the ACP outflow and ACP D4 (chain cost); Execution Market's 8 October payments on chain; AgentPact direct provider transfers; ACP and Execution Market ledger roll-off; DeskCrew ticket 465 (11 October) and DeskCrew supply (about 18 receipt reads would date the newest rows, still a lower bound); MoltJobs' 3 pending jobs (12 October); one test request to the block explorer next run; price tiers and tokens-covered inputs turn stale on 13 October, and TaskMarket's 0.98 may move back; still-payable TaskMarket tasks (16 to 24 October); trybounty.ai claim drift.

Scope. Read-only public GETs and read-only chain reads only: no accounts, keys, POSTs, payment headers, wallets, signing, entries or model calls against tasks. Task text was read as data and not stored. No private person is named.

7. How sure we are

  • The check thresholds and the measurement units are our choices, written down before grading. Other choices give other grades; the sensitivities show the ones we know matter.
  • Execution Market's 7.1% is a lower bound and 12.8% an assumption that all disputes fall in the window; it is graded at 12.8%. DeskCrew's 23.5% is an operator all-time aggregate, not a 30-day window.
  • AgentPact's fail rests on 12 deals that may be test deals; it does not show that agents went unpaid for delivered work.
  • TaskMarket's central verdict sits $0.0005 per attempt below break-even on cost estimates that are re-measured soon; it may move back.
  • The ledger total is an estimate with two carried aggregates and a partial 1-day refresh. On-chain data proves transfers, not identities. Rules quotes describe text, not enforcement, and are not legal advice. No probe or pathway was tested by the pool.

Search published pools, pages, reports, and evidence.