The full paradigm — subgoal conditioning, flow-based action execution, gated steerability, and a verified safety envelope — implemented at miniature scale as a cable-untangling experiment, run end-to-end on GPU. Zero-shot protocol: policies train on two crossing-topology families and are evaluated on held-out topologies and material stiffnesses they never saw.
A deformable cable in a crossed configuration is the smallest task with the proposal's defining failure modes: the state is contact-rich and non-rigid (a random walk never untangles itself — tangles persist unless the gripper resolves them), the space of configurations is effectively unbounded (impossible to enumerate in hardcoded logic), and the task is long-horizon (every crossing must be cleared in the right order without re-folding the cable).
The simulator is a position-based-dynamics chain with fixed segment lengths, free hinges, and localized constraint propagation: a pull bends the strand the gripper holds while distant units stay put — tangles persist until the policy actively resolves them, and a do-nothing policy provably fails.
Every stage of the full proposal has a faithful counterpart in code (src/steerable/), and the study runs as one command.
The oracle teleoperator (a search-based untangling expert) generates demonstrations. Milestone snapshots become subgoals; the first-crossing point relative to the gripper is the "act here" conditioning signal — the miniature's stand-in for a VLM-extracted visual subgoal. Synthetic subgoal noise densifies the conditioning distribution.
A conditional flow-matching model integrates a learned velocity field over action chunks: continuous gripper deltas by flow, the discrete grab decision by a separate Bernoulli head — decoupled exactly as real VLA systems separate grasp from motion. Inference averages multiple flow samples.
The steerable conditioning layer adds a gated correction field anchored at u=0; the control-barrier-function QP filter projects every commanded action onto the safe workspace with discrete forward invariance, and the environment counts bound violations instead of clipping silently.
Pre-registered and committed (scripts/run_experiment.py):
Variants — bc (flat behavioral cloning) ·
flow_flat (flow matching, no subgoals, no steering) ·
ours_nofilter (subgoals + SMC, no safety filter) ·
ours_full (everything) ·
transformer (5.3M-param Transformer backbone, self-attention over cable nodes) — × 3 seeds.
Train — crossing topologies {2, 3}, stiffness × {0.9, 1.1}.
Zero-shot eval — held-out topology {4}, stiffness × {0.6, 1.5}, fresh starts.
Metrics — success · no-intervention success · crossings reduced · steps ·
max jerk · safety violations · oracle interventions (reported as a cost).
Flywheel — the same policy under three curation strategies: none
(control) · near-miss curation · oracle relabeling (DAgger).
Interventions are the deployment-cost channel: if the policy stalls, an oracle assist performs one maneuver (capped), then the policy continues. No-intervention success counts only episodes the policy solves on its own.
Rendered from results/*.json — committed experiment artifacts, never hand-typed. Regenerate with PYTHONPATH=src python scripts/render_results.py.
Loading committed results …
Publication-quality figures generated from committed GPU results (scripts/figures.py).
Fig 1: CBF filter eliminates all workspace violations.
Fig 2: Each component contributes to the overall system.
Fig 3: Expert policy ceiling — 100% train, 92.2% held-out.
Fig 4: No-intervention success — the learning frontier.
Fig 5: Full system comparison — success rates and safety violations.