DataFly

A data flywheel for robot manipulation — the loop only spins if the labels come back. Measured across kinematic, pixel, and contact-rich regimes.
Sehaj Randhir Singh
Independent researcher; partial affiliation with NYU Tandon School of Engineering

Deployment data is the fuel of the robot data flywheel, but which data — and with whose labels — determines whether the loop compounds or stalls. We formalize the flywheel as collect → score → curate → retrain → repeat, and isolate the curation decision as the independent variable. On a planar pushing task with an imperfect learned policy, six curation strategies are compared across 6 training seeds and 300 held-out starts. Three findings. (1) Without feedback the policy is frozen — the loop never turns. (2) Self-curated episodes plateau — successes, near-misses, and novel states densify what the policy already does but cannot repair what it does badly. (3) Oracle relabeling of deployment failures is what compounds — blind relabeling reaches 0.66 success from 0.15, and curating by the progress signal reaches 0.48 while spending roughly half the oracle queries. The result is robust to oracle-labeling noise and task difficulty. The loop transfers to raw pixels — and the mechanism is mixing-ratio control: a measured phase diagram over (relabeled:clean ratio × policy capacity) shows the flood boundary is a function of capacity — ratio control alone rescues the CNN from 0.08 to 0.19 — while a low-capacity MLP is flood-robust even on contact-rich MuJoCo dynamics.

Does the flywheel spin?

6 training seeds × 300 held-out starts × 6 flywheel iterations, 12-dim state vector → 96-unit MLP. Every strategy starts from the same noisy expert demonstrations and the same frozen control baseline; the curation rule is the only variable.

Strategyiter 0midfinal (gain)
None (frozen control)0.120.120.12 (+0.00)
Self-curation · successes0.150.270.34 (+0.19)
Self-curation · near-misses0.140.280.32 (+0.17)
Self-curation · novel successes0.170.300.37 (+0.20)
Oracle · curated relabel0.200.360.48 (+0.28)
Oracle · relabel all (DAgger)0.150.500.66 (+0.51)
Held-out push success (mean over seeds); the full per-iteration table with std is in both papers.
Success rate vs flywheel iteration for all strategies
Held-out success vs flywheel iteration (mean ± std over seeds). Oracle relabeling compounds; self-curation plateaus; the frozen control is flat.

Perception flips the balance

4 training seeds × 200 held-out starts × 5 flywheel iterations, 64×64 RGB → torch CNN. The policy never sees the 12-dim state vector; it must infer the block, the target, and its own arm from the rendered image. Blind relabeling crashes the CNN (0.27 → 0.06): relabeled frames flood the dataset and the high-capacity CNN overfits its own failure distribution. Curated relabeling degrades far less (0.27 → 0.14).

Vision policy success curves
The CNN, on raw pixels

Blind relabeling floods the dataset and the high-capacity CNN overfits its own failures; curated relabeling is what keeps the loop stable.

Example trajectories
The raw material

Example rollouts: an expert seed demo, an early policy failure (the flywheel's raw material), and a final success — the block path shows the push being learned.

The mechanism — curation is mixing-ratio control

Each flywheel iteration adds relabeled frames to a clean seed set; the relabeled:clean mixing ratio of the training set is what blind relabeling lets grow without bound. We make it the controlled variable and sweep it against unbounded relabeling, across policy capacities and — for the first time — on contact-rich MuJoCo dynamics.

The flood boundary: final success vs mixing ratio
The flood boundary. a — Perception (CNN): final success falls monotonically as the ratio grows — 0.19 at ratio 0.25 → 0.15 at 1.0 → 0.08 unbounded. b–c — Kinematic and contact-rich MLP: flood-robust at every ratio (0.55 capped vs 0.57 unbounded on MuJoCo). The inset shows the closed-loop curator converging into the stable region without knowing capacity.

A low-capacity policy cannot memorize its failure distribution, so even unbounded relabeling forces generalization; a high-capacity CNN can, so flooding it overfits its own failures. The relabel_adaptive curator treats the flywheel report as a sensor — it halves its ratio when held-out success regresses and grows it otherwise, converging to the stable operating point on the CNN while staying near capacity on the MLP. Ratio control alone rescues 0.11 of the lost performance.

Label efficiency — what does the loop actually cost?

The flywheel's training interaction (29,297 environment steps) versus DQN from scratch (300,000). DQN gets the full state and a dense reward — and still lands at 0.04. Milestones: 20,000 steps → 0.32 · 40,000 steps → 0.28 · 60,000 steps → 0.59 · 80,000 steps → 0.17 · 100,000 steps → 0.10 · 120,000 steps → 0.08 · 140,000 steps → 0.04 · 160,000 steps → 0.14 · 180,000 steps → 0.09 · 200,000 steps → 0.04 · 220,000 steps → 0.06 · 240,000 steps → 0.12 · 260,000 steps → 0.07 · 280,000 steps → 0.09 · 300,000 steps → 0.04.

Success vs environment interactions
Held-out success vs environment interactions used for training (log axis). The label-efficiency argument: the flywheel compounds with an order-of-magnitude less interaction than tabula-rasa RL — and the papers price the labels, reporting the crossover λ* where RL becomes the cheaper route.
Oracle noise ablation
Noisy labels

Oracle-quality ablation: curated relabeling degrades less than blind relabeling under human-labeling noise.

Difficulty robustness
Harder task

On a harder task (tighter goal, longer pushes), oracle relabeling still compounds while no-feedback stays flat.

Reproduce

Committed results on Kaggle

git clone https://github.com/sehajr-singhs/robotic-data-flywheel
cd robotic-data-flywheel
pip install -e ".[dev]"

# state-based study (all six strategies)
python scripts/run_experiment.py
python scripts/merge_results.py

# vision study (torch CNN, pixels) + DQN baseline
python scripts/run_experiment.py --obs-mode image --strategies none relabel relabel_curated

# analyses, the papers' numbers, and this site — nothing hand-typed
python scripts/analyze.py
python scripts/analyze_mechanism.py
python scripts/render_results.py
python scripts/build_site.py

CPU-scale, seeded protocol (6 seeds per condition, per-seed values committed), committed result JSONs, and the same pipeline on Kaggle GPU kernels. Tests: pytest tests (17 tests covering physics, curation, the loop, and the vision path).