Deployment data is the fuel of the robot data flywheel, but which data — and with whose labels — determines whether the loop compounds or stalls. We formalize the flywheel as collect → score → curate → retrain → repeat, and isolate the curation decision as the independent variable. On a planar pushing task with an imperfect learned policy, six curation strategies are compared across 6 training seeds and 300 held-out starts. Three findings. (1) Without feedback the policy is frozen — the loop never turns. (2) Self-curated episodes plateau — successes, near-misses, and novel states densify what the policy already does but cannot repair what it does badly. (3) Oracle relabeling of deployment failures is what compounds — blind relabeling reaches 0.66 success from 0.15, and curating by the progress signal reaches 0.48 while spending roughly half the oracle queries. The result is robust to oracle-labeling noise and task difficulty. The loop transfers to raw pixels — and the mechanism is mixing-ratio control: a measured phase diagram over (relabeled:clean ratio × policy capacity) shows the flood boundary is a function of capacity — ratio control alone rescues the CNN from 0.08 to 0.19 — while a low-capacity MLP is flood-robust even on contact-rich MuJoCo dynamics.
6 training seeds × 300 held-out starts × 6 flywheel iterations, 12-dim state vector → 96-unit MLP. Every strategy starts from the same noisy expert demonstrations and the same frozen control baseline; the curation rule is the only variable.
| Strategy | iter 0 | mid | final (gain) |
|---|---|---|---|
| None (frozen control) | 0.12 | 0.12 | 0.12 (+0.00) |
| Self-curation · successes | 0.15 | 0.27 | 0.34 (+0.19) |
| Self-curation · near-misses | 0.14 | 0.28 | 0.32 (+0.17) |
| Self-curation · novel successes | 0.17 | 0.30 | 0.37 (+0.20) |
| Oracle · curated relabel | 0.20 | 0.36 | 0.48 (+0.28) |
| Oracle · relabel all (DAgger) | 0.15 | 0.50 | 0.66 (+0.51) |

4 training seeds × 200 held-out starts × 5 flywheel iterations, 64×64 RGB → torch CNN. The policy never sees the 12-dim state vector; it must infer the block, the target, and its own arm from the rendered image. Blind relabeling crashes the CNN (0.27 → 0.06): relabeled frames flood the dataset and the high-capacity CNN overfits its own failure distribution. Curated relabeling degrades far less (0.27 → 0.14).

Blind relabeling floods the dataset and the high-capacity CNN overfits its own failures; curated relabeling is what keeps the loop stable.

Example rollouts: an expert seed demo, an early policy failure (the flywheel's raw material), and a final success — the block path shows the push being learned.
Each flywheel iteration adds relabeled frames to a clean seed set; the relabeled:clean mixing ratio of the training set is what blind relabeling lets grow without bound. We make it the controlled variable and sweep it against unbounded relabeling, across policy capacities and — for the first time — on contact-rich MuJoCo dynamics.

A low-capacity policy cannot memorize its failure distribution, so even unbounded relabeling forces generalization; a high-capacity CNN can, so flooding it overfits its own failures. The relabel_adaptive curator treats the flywheel report as a sensor — it halves its ratio when held-out success regresses and grows it otherwise, converging to the stable operating point on the CNN while staying near capacity on the MLP. Ratio control alone rescues 0.11 of the lost performance.
The flywheel's training interaction (29,297 environment steps) versus DQN from scratch (300,000). DQN gets the full state and a dense reward — and still lands at 0.04. Milestones: 20,000 steps → 0.32 · 40,000 steps → 0.28 · 60,000 steps → 0.59 · 80,000 steps → 0.17 · 100,000 steps → 0.10 · 120,000 steps → 0.08 · 140,000 steps → 0.04 · 160,000 steps → 0.14 · 180,000 steps → 0.09 · 200,000 steps → 0.04 · 220,000 steps → 0.06 · 240,000 steps → 0.12 · 260,000 steps → 0.07 · 280,000 steps → 0.09 · 300,000 steps → 0.04.


Oracle-quality ablation: curated relabeling degrades less than blind relabeling under human-labeling noise.

On a harder task (tighter goal, longer pushes), oracle relabeling still compounds while no-feedback stays flat.
git clone https://github.com/sehajr-singhs/robotic-data-flywheel cd robotic-data-flywheel pip install -e ".[dev]" # state-based study (all six strategies) python scripts/run_experiment.py python scripts/merge_results.py # vision study (torch CNN, pixels) + DQN baseline python scripts/run_experiment.py --obs-mode image --strategies none relabel relabel_curated # analyses, the papers' numbers, and this site — nothing hand-typed python scripts/analyze.py python scripts/analyze_mechanism.py python scripts/render_results.py python scripts/build_site.py
CPU-scale, seeded protocol (6 seeds per condition, per-seed values committed), committed result JSONs, and the same pipeline on Kaggle GPU kernels. Tests: pytest tests (17 tests covering physics, curation, the loop, and the vision path).