Phase I · Foundation (W1–W3)
WEEK 1
Baseline reproduction & infra
Reproduce a π0-style flow-matching VLA baseline (open weights) on Open X-Embodiment subset + LIBERO-style sim; verify 50 Hz receding-horizon execution. Stand up the evaluation harness with intervention-rate and jerk logging from day one. Inventory hardware; define the canonical action frame and write the FK converter.
WEEK 2
Teleop data + subgoal extraction
First 50 expert episodes across the three tasks, logged with raw video. Build the offline subgoal extractor: VLM keyframe selection → language-step labeling → language-conditioned synthesis of novel subgoal images. Segment action chunks at subgoal boundaries; sanity-check stability across repeated demos.
WEEK 3
Procedural sim + subgoal-conditioned training
Procedural simulators for all three tasks with aggressive domain randomization and ground-truth subgoal boundaries. First subgoal-conditioned training runs with the contrastive alignment loss. Scale sweep: subgoal density K vs. horizon — does dense coverage monotonically help beyond a floor?
Phase II · Steerability & safety (W4–W6)
WEEK 4
SMC layer v1 + safety filter — M1
Implement the SMC gated adapter: steering tokens, gate parameterization, anchoring constraint, Grönwall-bound calculator. Train with synthetic steering supervision (random perturbations relabeled as steering) so the adapter learns “absorb a correction, return to nominal” without human-in-the-loop. Implement the CBF–QP filter; measure solve time in the 50 Hz loop; offline falsification over steering amplitude/timing.
WEEK 5
Sim evaluation suite, round 1
All three tasks × held-out difficulty splits; baselines (RT-2-style, Diffusion Policy, ACT, vanilla π0-style, text-only-subgoal hierarchical) vs. the full system. First ablation pass: −SMC, −filter, −visual subgoals, −dense subgoals, −canonical frame. Leaderboard v1 with intervention rate + jerk as costs.
WEEK 6
Data flywheel on sim deployment — M2
Deploy in sim and run the curation loop: score rollouts, curate near-misses, relabel failures with the teleoperator as oracle, retrain, re-measure. Vary the curation rule (self vs. curated-relabel vs. blind-relabel) to confirm the compounding law transfers from the toy study to the VLA setting.
Phase III · Sim-to-real & hardware (W7–W10)
WEEK 7
Sim-to-real bridge
Domain-randomization tuning against real teleop episodes; real-data mixing sweep (0 / 5 / 10 / 20% real teleop in the batch); subgoal supervision on real video. Sim-to-real gap report.
WEEK 8
Real hardware: untangling + tool use
Deploy on the rig for Task A and Task C: 100-episode eval sets, intervention-rate tracking, safety-filter logging. First real steering eval — human nudges mid-task; measure jerk under steering vs. without.
WEEK 9
Real hardware: textile folding
Task B with world-model replanning active; force-limit steering via the filter (contact forces capped in-task). Full three-task hardware suite green.
WEEK 10
Zero-shot ablations on hardware
Held-out morphologies/materials/tools on the real rig, per the pre-registered protocol — the actual zero-shot claim, measured. Component-to-metric attribution table.
Phase IV · Analysis & release (W11–W13)
WEEK 11
Mechanism analysis
Error attribution (which failure class does each component remove?); steering amplitude/timing sweep vs. safety violations; composability study over unseen combinations of morphology × material × tool.
WEEK 12
Figures, video, code cleanup
Monochrome figure set (architecture, safety envelope, flywheel curves, zero-shot leaderboards); demo video with steering overlays; reproduce-everything pass — all numbers from committed artifacts.
WEEK 13
Papers and release — M3
Finalize the NMI and IEEE manuscripts; compile clean; submission checklists. Public release: repo with README, results, and this site.
Budget of effort & risks
| Phase | Weeks | Focus |
|---|---|---|
| I. Foundation | 1–3 | baseline, data, subgoal pipeline |
| II. Steerability & safety | 4–6 | SMC, CBF filter, sim ablations, flywheel |
| III. Sim-to-real & hardware | 7–10 | real untangling / folding / tool use, zero-shot |
| IV. Analysis & release | 11–13 | mechanism study, papers, release |
- Teleop throughput — hedge: procedural sim data is generated continuously from W3, so training never blocks on human hours.
- Real steering interface latency — hedge: steering validated in sim first; the filter bounds deviation regardless of interface.
- Task B hardware difficulty — hedge: folding is last, so failures cannot block the untangling/tool-use results.
- Compute — hedge: the flow expert is 2–3B active parameters; ablations run at fixed budget with early stopping, scheduling-bound rather than compute-bound.