Shaduf.Research preview
Digital Fly Lab/Fourth check: does fly wiring help? 31 control studies compared, a FlyLeno verdict and 12 new entries

Research report · 30 September 2026 · fourth check

Fourth check: does fly wiring help? 31 control studies compared, a FlyLeno verdict and 12 new entries

This check asked one question across the whole catalogue: when a project scrambles the fly's wiring, does the result get worse? We collected 31 control studies into one ledger and compared them on a single scale. We also refereed the week's most-watched new claim, FlyLeno, re-checked all 82 entries and 47 videos, and added 12 entries and 2 videos.

  • Run: run:a46416d4-3706-4c00-ba7d-1371567178c4
  • Research time: 30 Sep 2026, 09:36–10:11 UTC
  • Catalogue version: 2026-09-30-run4 (94 entries)
  • Gallery: 49 videos
  • Controls ledger: 2026-09-30-run4 (31 studies)

Short answer

Does fly wiring help? Sometimes: real fly wiring beat scrambled wiring in 10 of 23 studies, mostly reflex, sensory and steering circuits; in trained, reservoir and ML tasks it tied or lost (7). Of 31 control studies, 23 have a wiring null: the real wiring helps in 10, makes no difference in 5, does worse in 2, gives mixed results in 5, and 1 is not yet scored. Where the wiring helps, a brainless baseline often still matches the fly: in Doom the fly survives 53.3 s, shuffled wiring 6.8 s and a brainless autopilot 51.9 s.

How sure we are: fairly sure of the split, less sure of any single number. Only 9 of 31 studies reach "strong" method quality. Every number except our own run is the authors'; we recomputed 15 studies from their committed files but re-ran none, and none is replicated.

One claim met our attention bar: FlyLeno ("Tonight's host: Grey Leno, piloted by Drosophila melanogaster"), after a streamer's upload rose to 36,763 views. Graded from its code as B, borderline C: the whole FlyWire brain really runs, untrained, in the browser, but inputs and outputs are mapped by hand, gait, homing and the choice of actions are engineered, and there is no control. The Pokémon stream and the Rainbow Six claim were re-checked and stay U.

What changed

  • 82 → 94catalogue entries (12 new: 2 A, 5 B, 2 C, 1 D, 2 not graded)
  • 47 → 49gallery videos (2 new, 0 removed)
  • 31control studies in a new public ledger, with a new page
  • 0main-link failures and grade changes; 3 new releases
Tracking pass on 30 Sep 2026 (82 existing entries, 47 existing videos).
MeasureValue
Entries re-checked82 of 82, plus 12 new
Repositories with new commits / new releases7 / 3
Grade changes0 (three entries re-checked because of new commits)
Main-link failures (4xx, 5xx or error)0 (one host showed a bot check, which is not a dead link)
Videos re-checked / new / removed47 / 2 / 0
Data formatTwo new optional fields on every catalogue record: controls and wiring_effect

New releases: acamilo-flybrain v0.6.7 (a damage reward for the Pokémon battles; it changes the reward signal, not the mechanism, so the grade stays C), FlyWire annotations Version 3.2.0 (29 Sep; an annotation release, not a new connectome) and Kick the Fly 2.13.1 (tests and tooling; stays C).

Other new commits: fly-brain-zero-shot (an install fix in the README), flyconnectome-nulls (author affiliation only), flydoom by mutkuoz (13 commits re-scoring a separate "everything applied" variant; the plain model's scorecard and its grade A are unchanged) and Haltere (round-6 flights passed the Minus Two hairpin for the first time, still with no race finish; stays C). For the ledger studies whose code moved, the numbers were read at the new commit and did not change.

Held-back entries: FlyKart (B) and making-fly-play-chess (C), held back in the third check, were re-verified at their current commits and published.

What people argued about this week

We looked for new fly-brain claims since 29 Sep 2026, 09:50 UTC. The promotion rule is unchanged: a claim must be specific and checkable, have gained attention (30 points on Hacker News, 20,000 video views, 20 GitHub stars within a week, or a news story) and not be covered yet; at most two per check.

Lead scan on 30 Sep 2026: 57 searches plus web searches, dataset pages, directory lists, 8 video pages and 7 re-check fetches; 608 raw hits.
SourceHitsRelevant in the window
Hacker News stories / comments1 / 60No fly-brain story
GitHub277 (98 unique)22 new repositories, each with 0 or 1 star
YouTube242 (121 unique)Vinesauce's FlyLeno stream at 36,763 views; in-window uploads all under 40 views
Google News28No new story
Reddit0Refused (HTTP 403) without a login
  • Promoted (1): FlyLeno. The stream upload grew from 16,201 views (29 Sep) to 36,763 (30 Sep, 09:41 UTC). The viral video is a streamer's playthrough; the claim we graded is the site's and the repository's own.
  • Re-checked, no change: the Pokémon stream's creator still has no public code (their GitHub address returns "not found"), and no recording or follow-up article appeared; the Rainbow Six report is still behind a bot check, and a web search found only the original article. X cannot be searched without a login.
  • Not promoted: "I Taught a Fly How to Play Gorilla Tag" (415,612 views) says the fly "learned" with "its open source brain" but links no code. It was uploaded on 26 Sep, before the window, and our date-sorted video search had missed it; it is flagged for a later check. Two other older videos (24,575 and 11,934 views) link no project.
  • Data releases: FlyWire annotations Version 3.2.0 (29 Sep). No new connectome release.

Verdict: "Tonight's host: Grey Leno, piloted by Drosophila melanogaster"

FlyLeno (TUURD Talk) is a browser show by AgitationSkeleton in which "a whole-brain spiking model of the adult fruit fly (FlyWire v783, 138,639 neurons, 15.1M connections) runs live and puppeteers Grey Leno" (README). We read the code at commit 078d8fa:

  • Wiring: the Shiu et al. v783 files, packed with every connection kept (tools/build_connectome.py).
  • Neurons: the Shiu et al. leaky integrate-and-fire equations, event-driven in a Web Worker (js/brain-worker.js). Dopamine-gated plasticity, on by default, changes only the synapses onto the voice neurons.
  • Inputs and outputs, hand-set: show events, thrown objects, music and food stimulate chosen sensory groups; descending-neuron rates become walk, turn, startle, groom and feed commands through fixed gains of 40 Hz (15 Hz for startle) (js/motor.js).
  • Engineered layers, labelled by the author: a walking rhythm generator, balance "puppet strings", saccades, food seeking, homing and sleep (js/instincts.js), and an action selector outside the connectome that picks spontaneous actions and stimulates the matching descending neurons (js/mind.js).
  • Controls: none. A search of the code for shuffle, scramble or rewire finds only playlist and animation code.

B Verdict, borderline C: the fly brain is real and really running, but Leno's show is mostly built around it. The brain's reflexes move him; his walking rhythm, homing and what he does next are engineered; and nothing tests whether a scrambled brain would look the same. Up to A with a scrambled-wiring or no-brain run of the same show and a measured behaviour; down to C with evidence that the engineered layers drive most on-screen behaviour. Verdict row · catalogue entry

Other rows on Is it real?: the Doom row now adds that the Doom control study's results file (results/a2/summary.json) holds a sweep of 12 autopilot settings, the best of which survives 53.44 s, slightly longer than DOOMFLY's 53.34 s, next to the matched autopilot's 51.93 s; we checked this in the file. The controls section now points to the new ledger instead of its earlier 20-study table.

Does fly wiring help?

Full page: Does fly wiring help? Data: controls.json and its schema.

Wiring effect by task family (31 studies; "not tested" = no wiring null, or results not in yet).
Task familyHelpsNo differenceWorseMixedNot tested
Reflex circuits40000
Sensory models20001
Steering a body20012
Playing games12024
Machine-learning benchmarks03110
Reservoir computing00100
Language models00001
Forecasting00001
Graph analysis10000
Evolved controllers00010
All 31 studies105259

Of the 9 "not tested", 8 studies have no wiring null (only baselines or ablations). The ninth, bioreservoir, has one, but its forecast questions resolve only on 15 Dec 2026.

How much of each result survives when the fly wiring is scrambledNull retention for each study with a wiring null or a no-brain baseline, grouped by task family. 0 means the scrambled brain does no better than no network; 1 means it does as well as the real wiring. Circles: the scrambled or rewired wiring. Diamonds: a no-brain or no-graph baseline, as a share of the real fly network's result. Reflex circuits: Our run: sugar to MN9 0.00 (helps); drosophila-brain-mlx 0.00 (helps); Shiu et al. 2024 model 0.01 (helps); fly-brain (Lulzx) 0.70 (helps). Sensory models: flydoom (smell response) 0.00 (helps); Fly OCR (digits) 0.71 (helps). Steering a body: Zero-shot steering 0.08 (helps); FLY-lab: brain and body 0.53, no-brain baseline 1.00 (helps); Fly.exe (giant fibre) index can be negative: +2.0 vs -0.1 (mixed). Playing games: Doom control study 0.03, no-brain baseline 0.97 (helps); Chess (FlyWire patch) 0.98, no-brain baseline 1.37 (no difference); doomfly-rl (trained) 1.02, no-brain baseline 1.13 (no difference); Fly Self Driving 0.83, no-brain baseline 0.93 (mixed); Arkanoid (fly-plays-games) 1.00, no-brain baseline 0.90 (mixed); Fly Dino (80 neurons) no wiring null, no-brain baseline 0.24 (no wiring null); Fly Worker (game map) no wiring null, no-brain baseline 1.51 (no wiring null); Haltere drone races no null: 0 of 2 vs PD 1 of 2 (no wiring null). Machine-learning benchmarks: fly-cartpole 0.93, no-brain baseline 1.32 (no difference); Larva reservoir (CIFAR-10) 1.00 (no difference); NeuroWeave 1.01 (no difference); The Fly's Hash Function 1.11, no-brain baseline 0.48 (worse); flybench (28 tasks) 0.67 (mixed). Reservoir computing: FlyBrain reservoir 7.05 (worse). Language models: FLM language model no re-fitted wiring null, no-brain baseline 1.02 (no wiring null). Forecasting: bioreservoir forecasts not scored until 15 Dec 2026 (not yet scored). Graph analysis, no simulation: Wired Different (graph) 0.93, no-brain baseline 0.95 (helps). Evolved controllers: Null-model evolution 1.13 (mixed).00.250.50.7511.25wiring mattersno difference0 = no better than no network1 = as good as real wiringReflex circuits (4)Our run: sugar to MN9Build your own: sugar to MN9 with shuffled wiring (our run): scrambled wiring keeps 0.00 of the real result (Degree-preserving rewiring)helpsdrosophila-brain-mlxdrosophila-brain-mlx: scrambled wiring keeps 0.00 of the real result (Degree-preserving rewiring)helpsShiu et al. 2024 modelDrosophila_brain_model (Shiu et al. 2024): scrambled wiring keeps 0.01 of the real result (Weight shuffle)helpsfly-brain (Lulzx)fly-brain: scrambled wiring keeps 0.70 of the real result (Weight shuffle)helpsSensory models (2)flydoom (smell response)flydoom: scrambled wiring keeps 0.00 of the real result (Degree-preserving rewiring)helpsFly OCR (digits)Fly OCR: scrambled wiring keeps 0.71 of the real result (Degree-preserving rewiring)helpsSteering a body (3)Zero-shot steeringAre fruit flies zero-shot adapters?: scrambled wiring keeps 0.08 of the real result (Degree-preserving rewiring)helpsFLY-lab: brain and bodyFLY-lab: What a fly connectome adds to controlling a body: scrambled wiring keeps 0.53 of the real result (Degree-preserving rewiring)FLY-lab: What a fly connectome adds to controlling a body: the no-brain baseline reaches 1.00 of the fly network's resulthelpsFly.exe (giant fibre)index can be negative: +2.0 vs -0.1mixedPlaying games (8)Doom control studyIs the fly brain actually playing DOOM? (control experiments): scrambled wiring keeps 0.03 of the real result (Degree-preserving rewiring)Is the fly brain actually playing DOOM? (control experiments): the no-brain baseline reaches 0.97 of the fly network's resulthelpsChess (FlyWire patch)making-fly-play-chess: scrambled wiring keeps 0.98 of the real result (Degree-preserving rewiring)making-fly-play-chess: the no-brain baseline reaches 1.37 of the fly network's resultno differencedoomfly-rl (trained)doomfly-rl: scrambled wiring keeps 1.02 of the real result (Degree-preserving rewiring)doomfly-rl: the no-brain baseline reaches 1.13 of the fly network's resultno differenceFly Self DrivingFly Self Driving: scrambled wiring keeps 0.83 of the real result (Target permutation)Fly Self Driving: the no-brain baseline reaches 0.93 of the fly network's resultmixedArkanoid (fly-plays-games)fly-plays-games (Pokémon Red chapter; formerly fly-plays-pokemon): scrambled wiring keeps 1.00 of the real result (Weight shuffle)fly-plays-games (Pokémon Red chapter; formerly fly-plays-pokemon): the no-brain baseline reaches 0.90 of the fly network's resultmixedFly Dino (80 neurons)no wiring nullFly Dino (flyjump): the no-brain baseline reaches 0.24 of the fly network's resultno wiring nullFly Worker (game map)no wiring nullFly Worker: the no-brain baseline reaches 1.51 of the fly network's result1.51 →no wiring nullHaltere drone racesno null: 0 of 2 vs PD 1 of 2no wiring nullMachine-learning benchmarks (5)fly-cartpolefly-cartpole: scrambled wiring keeps 0.93 of the real result (Degree-preserving rewiring)fly-cartpole: the no-brain baseline reaches 1.32 of the fly network's resultno differenceLarva reservoir (CIFAR-10)Does the larval connectome beat its own shuffles? (connectome-null-models): scrambled wiring keeps 1.00 of the real result (Degree-preserving rewiring)no differenceNeuroWeaveNeuroWeave: scrambled wiring keeps 1.01 of the real result (Degree-preserving rewiring)no differenceThe Fly's Hash FunctionThe Fly's Hash Function: scrambled wiring keeps 1.11 of the real result (Degree-preserving rewiring)The Fly's Hash Function: the no-brain baseline reaches 0.48 of the fly network's resultworseflybench (28 tasks)flybench: scrambled wiring keeps 0.67 of the real result (Degree-preserving rewiring)mixedReservoir computing (1)FlyBrain reservoirflybrain-reservoir: scrambled wiring keeps 7.05 of the real result (Degree-preserving rewiring)7.05 →worseLanguage models (1)FLM language modelno re-fitted wiring nullFLM - Fly Language Model: the no-brain baseline reaches 1.02 of the fly network's resultno wiring nullForecasting (1)bioreservoir forecastsnot scored until 15 Dec 2026not yet scoredGraph analysis, no simulation (1)Wired Different (graph)Wired Different (ConnectomeLens): scrambled wiring keeps 0.93 of the real result (Degree-preserving rewiring)Wired Different (ConnectomeLens): the no-brain baseline reaches 0.95 of the fly network's resulthelpsEvolved controllers (1)Null-model evolutionNull-model treatment of the sensory-motor boundary changes an evolutionary connectome comparison: scrambled wiring keeps 1.13 of the real result (Degree-preserving rewiring)mixedNull retention (values above 1.4 are shown at the edge with their number)scrambled wiring: helpsno differenceworsemixedno-brain baseline (share of the fly result) Null retention for each study with a wiring null or a no-brain baseline, grouped by task family. 0 means the scrambled brain does no better than no network; 1 means it does as well as the real wiring. Circles: the scrambled or rewired wiring. Diamonds: a no-brain or no-graph baseline, as a share of the real fly network's result. Reflex circuits: Our run: sugar to MN9 0.00 (helps); drosophila-brain-mlx 0.00 (helps); Shiu et al. 2024 model 0.01 (helps); fly-brain (Lulzx) 0.70 (helps). Sensory models: flydoom (smell response) 0.00 (helps); Fly OCR (digits) 0.71 (helps). Steering a body: Zero-shot steering 0.08 (helps); FLY-lab: brain and body 0.53, no-brain baseline 1.00 (helps); Fly.exe (giant fibre) index can be negative: +2.0 vs -0.1 (mixed). Playing games: Doom control study 0.03, no-brain baseline 0.97 (helps); Chess (FlyWire patch) 0.98, no-brain baseline 1.37 (no difference); doomfly-rl (trained) 1.02, no-brain baseline 1.13 (no difference); Fly Self Driving 0.83, no-brain baseline 0.93 (mixed); Arkanoid (fly-plays-games) 1.00, no-brain baseline 0.90 (mixed); Fly Dino (80 neurons) no wiring null, no-brain baseline 0.24 (no wiring null); Fly Worker (game map) no wiring null, no-brain baseline 1.51 (no wiring null); Haltere drone races no null: 0 of 2 vs PD 1 of 2 (no wiring null). Machine-learning benchmarks: fly-cartpole 0.93, no-brain baseline 1.32 (no difference); Larva reservoir (CIFAR-10) 1.00 (no difference); NeuroWeave 1.01 (no difference); The Fly's Hash Function 1.11, no-brain baseline 0.48 (worse); flybench (28 tasks) 0.67 (mixed). Reservoir computing: FlyBrain reservoir 7.05 (worse). Language models: FLM language model no re-fitted wiring null, no-brain baseline 1.02 (no wiring null). Forecasting: bioreservoir forecasts not scored until 15 Dec 2026 (not yet scored). Graph analysis, no simulation: Wired Different (graph) 0.93, no-brain baseline 0.95 (helps). Evolved controllers: Null-model evolution 1.13 (mixed).0 = no better than none1 = as good as realReflexesOur sugar run · helpsMLX port · helpsShiu model · helpsLulzx model · helpsSensesflydoom smell · helpsFly OCR · helpsBody steeringZero-shot · helpsFLY-lab · helpsFly.exe · mixedGamesDoom control · helpsChess · no diff.doomfly-rl · no diff.Self driving · mixedArkanoid · mixedFly Dino · no nullFly Worker · no nullHaltere · no nullML benchmarksCartPole · no diff.Larva · no diff.NeuroWeave · no diff.Fly hash · worseflybench · mixedReservoirFly reservoir · worseLanguage modelFLM · no nullForecastingForecasts · not scoredGraph analysisWired Different · helpsEvolutionNull models · mixed00.511.4+scrambled wiringno-brain baseline
Circles: how much of the real wiring's result the scrambled or rewired wiring keeps (the headline null: degree-preserving where the study has one). Diamonds: how much a no-brain or no-graph baseline reaches, as a share of the fly network's result. A line joins the two. Doom is the clearest case: scrambled wiring keeps 0.03 of the survival time, but a brainless autopilot reaches 0.97. Values above 1.4 are drawn at the edge with their number. Retention is computed from the authors' numbers (recomputed by us for 15 studies).

Which null is fair

The ledger lists 16 kinds of null and control. Where the choice of null changed the result: in flyconnectome-nulls, standard nulls beat the connectome (1.78 and 1.77 against 1.57) while nulls that keep the sensory-motor boundary tie it (1.58 and 1.62); in the Arkanoid chapter of fly-plays-games a weight shuffle keeps the play and a rewiring loses it; in FLY-lab, re-fitting the readout for each null shrinks the gap from 69 to 47 points; reservoir results depend on each wiring's gain, and the larva study, which fixes it, finds no difference; and in Doom the brainless autopilot matches the brain. The headline chart uses the degree-preserving null first (19 of the 23 studies with a wiring null). The four-point fair-control checklist: keep degrees and signs; tune, train and re-fit the null like the real wiring; use 3 or more shuffles and report the spread; include a no-brain baseline and say whether the fly beat it.

Replayed vs live, brain vs no brain

Replayed brain output scores 0 of 30 on FLY-lab's turning task against 100% live; in the zero-shot study, disconnected turn neurons give 0.125 poles per 10 s, scrambled wiring 0.34 and the live real wiring 2.96. No study shows replay matching live performance. 13 studies have a no-brain or no-graph baseline, and the fly network beats it in only 2: fly-dino (179 s against 46 s for a hand rule, but with no wiring null) and the Fly's Hash Function (beats SimHash while losing to its wiring nulls).

Method quality of the 31 studies (rule-based labels).
MeasureStudies
Quality: strong / moderate / weak9 / 6 / 16
Null fairness: fair / partly fair / unclear / unfair11 / 15 / 4 / 1
Numbers from result files / papers / catalogue text27 / 3 / 1
Recomputed by us from committed files15
Our own measurement1 (the Build your own shuffle run)
Peer-reviewed2 (the Shiu et al. model and flyvis)

What would change the answer: a behavioural study with a fair null (3 or more degree-preserving shuffles, tuned equally) and a no-brain baseline in which the real wiring beats both (none exists); boundary-preserving nulls applied to the reflex-circuit studies; independent replications; and our own fair-null test of the Shiu et al. model with weight, degree-preserving and sign shuffles, 5 or more each, which is a candidate for a coming check.

New entries

Every entry was graded from files we opened. The numbers are the authors' own.

12 new catalogue entries on 30 Sep 2026: 10 new ones (the per-check cap) and the 2 held back in the third check.
EntryKindGradeControlsDeciding evidence
fly-cartpoleResearchAWiring null: no differenceMushroom bodies wired from MaleCNS learn CartPole (233.8 steps over 20 held-out seeds); degree-preserving shuffles learn about as well (218.3, p = 0.20; recomputed by us); a TD learner does better (302.3).
flybrain-connectome-benchmarkResearchANone (checked against fly data, not a null)A pre-registered test of the Shiu et al. whole-brain model against published experiments: knockout sensitivity 0.769, specificity 0.989 (preprint).
FlyLenoBrowser demoBNoneSee the verdict above.
FlyKartGameBNoneThe whole MaleCNS brain steers a go-kart, but the game computes what the fly "sees" and the controls are hand-mapped; a new CPU path runs without a GPU.
Open FlyBrowser demoBNone (a shuffle control is pre-registered, with no result yet)The whole Shiu et al. brain plays a strategy game in the browser; game events go to taste neurons and 39 arbitrary action groups.
Fly With Me (Fly Brain DJ)ArtBNoneThe FlyWire brain hears music through its antennal neurons and triggers DJ moves from named descending neurons; the mixing is ordinary audio code.
Fly-NAFGameBNoneThe FlyWire brain plays Five Nights at Freddy's; the seeing is a pixel-difference threshold and each action comes from one chosen neuron.
making-fly-play-chessGameCWiring null: no differenceA fixed FlyWire patch with a trained readout ties its rewired twin (0.506) in a run the author marks as a code check, and loses to a material-count player.
FlyCNS Tic-Tac-ToeGameCNoneBy default a minimax solver picks the playable moves; the fly circuit only breaks ties.
FLYBRAIN Bad Apple x DOOMArtDNoneA connectome viewer that uses neurons as pixels; no brain activity is simulated.
connectome_interpreterData tooln/aNoneA library for turning wiring diagrams into testable circuit hypotheses (MIT, with a preprint).
FlyBrainLabData tooln/aNoneAn older interactive platform for fly brain data and circuits (BSD-3); last commit September 2025.

Left out: 31 other GitHub leads, each with a reason in our exclusion log: most have 0 or 1 star and no demo or control; two promote crypto tokens; two are off-topic; and two repositories that copy FlyGym's name were not opened and are not linked. The first in line for the next check include a project with its own controls section. The Fly Brain Hub directory added 10 repositories.

49 videos, 2 new: Vinesauce's FlyLeno stream, labelled as a streamer's playthrough, not the author's upload (listed because the project links no video of its own), and Fly-NAF's best run, the author's upload. All 47 earlier videos were available, and every card's grade matches the catalogue. Not admitted: the Gorilla Tag video (no project to map it to) and uploads without a project link. The Pokémon and Rainbow Six rows still have no video we could admit, and 9 of the new entries have none. Thumbnails are unchanged.

Build your own: pins and the "Add a body" path

Pins unchanged. No pinned package of the beginner guide has had a release since 27 Sep 2026; newer brian2 2.10, numpy 2.5 and pandas 3.0 predate our pin test. The next full re-test is due by 12 Oct 2026.

FlyGym: the latest release is still 2.1.0 (24 Jun 2026). It requires Python 3.12 to 3.14 (>=3.12,<3.15); the last release for Python 3.11 is 1.2.1, with the old interface. It has no hard GPU dependency, and its download is about 170 MB (an estimate from package metadata). Our research host has Python 3.11, so a brain-body test is not possible there without an extra Python 3.12 install; that choice is open. The body path on Build your own therefore stays as it was: installing FlyGym 2.1.0 and running its physics on a CPU were tested on 28 Sep 2026 with Python 3.12, and connecting it to the brain is untested.

How often things change

Three daily checks so far (28, 29 and 30 Sep 2026): 0 main-link failures in 213 link checks and 0 video removals in 131 video checks, while the share of repositories with new commits rose from 0% to 4% to 9% per day, mostly young demo repositories. We keep checking links and videos on every run, re-grade entries when their commits suggest a change (no upstream change has altered a grade so far), and refresh stars weekly except for young repositories.

Method

  • Lead scan: Hacker News, GitHub, YouTube and Google News searches since the end of the last scan, plus video pages and re-checks of the two open claims.
  • Tracking: every entry's main link, repository commits and releases, and every video link. For repositories that changed, we read the commit messages and re-checked grades where they suggested a change.
  • Controls ledger: one study per catalogue entry with a wiring null or a no-brain baseline, plus our own run. Numbers were read from committed result files, READMEs or papers at the commits listed in the data file. For 15 studies we recomputed the reported numbers from the committed files, and 8 studies were spot-checked a second time against files fetched separately; all matched. Retention, wiring effect, fairness and quality follow fixed rules; a gap counts only if it exceeds twice the combined standard error. One label was set by hand (Fly.exe is "mixed", because its index can be negative), one headline arm was moved to a note (chess: its interval is for the fly's score, not the difference) and one task family was set by hand (flydoom by mutkuoz counts as a sensory model, because its only wiring null is on the smell response).
  • GitHub data: anonymous API access. 59 tracked records keep stars and dates from earlier checks (27, 28 or 29 Sep), labelled with those dates; stars were fetched fresh for the 12 new entries and 11 young repositories.
  • Our own measurement: the Build your own shuffle run of 28 Sep 2026 is the only number we produced ourselves.

Limitations

  • X and TikTok cannot be searched without a login, and Reddit refuses scripts. View counts are single snapshots. Our date-sorted video search misses older videos that grow fast (the Gorilla Tag case).
  • Every ledger number except our own run is the authors'. We recomputed 15 studies from committed files but re-ran none. 7 studies rest on README or docs tables and 3 on papers; the flyvis ablation values sit in a figure and were not extracted.
  • Fairness and quality labels come from yes/no fields read from code. They are consistent but coarse: for example, the Shiu et al. paper result is "weak" only because the paper text does not give the shuffle count and spread.
  • The FlyGym download size is an estimate from package metadata; the 602 MB installed size is our measurement of 28 Sep 2026.
  • 29 GitHub leads were deferred without being opened.

Next

  • Decide the "Add a body" path (a Python 3.12 environment, the older FlyGym 1.2.1, or neither), and whether to referee the Gorilla Tag video.
  • Admit the next leads, including a project with its own controls section, and check the two FlyGym name copies before any link.
  • Add a "most viewed this week" video search, and keep the ledger current: new control studies, Open Fly's pre-registered result, and bioreservoir's scores after 15 Dec 2026.
  • Our own fair-null test (a candidate for a coming check): no study yet pairs a fair null with a no-brain baseline on a behaviour.

Main sources

Search published pools, pages, reports, and evidence.