1 · Title & abstract
Compositional Generalization in Embodied Foundation Models via Multimodal Subgoal Prompting and Flow-Based Action Execution — prepared for Nature Machine Intelligence format.
Generalist robot policies have scaled to thousands of tasks, yet they still fail the test that defines embodied intelligence: performing a long-horizon task on an unseen object, with an unseen morphology, in a world that does not hold still. Cables tangle in ways no training set contains; textiles shift under their own weight; a hook that was never seen must still be used. Scaling monolithic vision-language-action (VLA) weights alone does not produce compositional generalization: the policy either memorizes the data distribution or loses the spatial-temporal anchors that make execution reliable.
We argue that zero-shot generalization across unseen morphologies and chaotic manipulation tasks is achieved not by scaling a single network, but by structuring the problem into a steerable hierarchy: a high-level VLM decomposes the task into dense visual subgoals, a flow-matching action expert executes the segments between them at control frequency, the Steerable Multimodal Conditioning (SMC) layer admits mid-execution corrections with a provable Lipschitz bound on trajectory deviation, and a runtime formal-verification safety envelope (a control-barrier-function filter) certifies every executed action. The program is instantiated on three chaotic benchmarks — cable untangling, textile folding on shifting fabric, and adaptive tool use in clutter — under a pre-registered zero-shot protocol, so that every number in the resulting manuscripts traces to a committed experiment artifact.
Contributions. (1) a steerable two-level hierarchy in which composition happens at the subgoal level, not inside a single weight matrix; (2) the SMC layer with a formal no-jerk guarantee — steering without jerks, by construction; (3) a CBF–QP safety envelope with forward-invariance guarantees and offline falsification; (4) three chaotic benchmarks and a pre-registered zero-shot protocol with intervention rate and jerk as first-class metrics.
2 · System architecture & mathematical formulation
2.1 · Two levels, one shared token space
Observations are multi-view RGB-D plus proprioception (and optional tactile); instructions are language. The planner Λφ maps (ℓ, ot) to a dense subgoal sequence g1:K = {(κk, ℓk)}, where κk is a keyframe image token (decoded to pixels for conditioning) and ℓk its language step. Dense coverage means K scales with task horizon: an untangling with three crossings gets roughly six to ten subgoals.
Text, keyframes, and runtime constraints are fused in a shared token space: the same cross-attention layers in the VLM decoder project language tokens, vision tokens, and steering tokens onto the action-latent stream (π-series parameterization). A contrastive alignment loss pairs each subgoal token with the observation chunk achieved at its segment boundary, so the two levels agree on what “done” means — the condition for the factorization to compose.
Actions live in a canonical cross-embodiment frame: normalized end-effector deltas (Δpee, Δψ) plus gripper, converted from joint space by forward kinematics. One flow expert, many morphologies, zero-shot.
2.2 · Flow matching over action chunks
Let x1 be a demonstrated chunk and x0 an independent noise sample. With the linear interpolation xs = (1−s)x0 + sx1, conditional flow matching trains the velocity field:
Inference integrates the ODE dxs/ds = vθ(xs, s, c) with 2–10 Euler/Heun steps, executed with a receding horizon at 50 Hz. Flow matching is preferred over score-based diffusion because few-step sampling preserves smoothness at control frequency, the marginal path is stable to train at scale, and the ODE formulation yields a velocity field that can be analyzed — the substrate for steering and safety.
2.3 · The Steerable Multimodal Conditioning layer
Steering signals — human corrections (clicked keypoints, drag vectors), world-model replans (stale subgoals replaced), and runtime constraint setpoints (force ceilings, keep-out zones) — enter through a gated adapter on the velocity field:
with the anchoring constraint vθj(·, ·, ·, uj0) ≡ 0 — removing steering recovers the nominal field exactly — and gates λj = σ(wjTφ(xs, s)) with bounded Lipschitz constant Lλ, so signals engage softly over a window rather than as a step in action space.
The no-jerk theorem. If the steered field is Lv-Lipschitz in x and the steering contribution is bounded by M, then over a steering window of flow-time Δs,
a computable smoothness envelope: an operator bounds a priori how far a correction may move the trajectory. Jerk — max over time of ‖d³pee/dt³‖ — is bounded by the same quantity divided by the gate rise time; steering is continuous in flow time by construction, and the bound is verified numerically during training.
2.4 · The runtime formal-verification safety envelope
Let the safe set be 𝒞 = {x : h(x) ≥ 0} for a control barrier function encoding joint limits, collision margins, and contact-force ceilings. At each control step the filter solves a convex QP:
Standard CBF theory gives forward invariance: if h(x(0)) ≥ 0, then h(x(t)) ≥ 0 for all t — the robot provably never leaves the safety set. The composition is the architectural point: steerability is admitted only through the filter, so corrections reshape the trajectory inside the certified tube and never outside it. The QP solves in microseconds inside the 50 Hz loop; the envelope is falsified offline over steering amplitudes and timings before hardware ever sees it.
2.5 · Why the structure composes
- Subgoal factorization. p(a | o, ℓ, g1:K) = Πk p(a[k] | o, gk, ℓ): each flow segment is conditioned on one subgoal, and training data is segmented at subgoal boundaries — the expert learns local “reach the keyframe” skills that compose across tasks.
- Synthetic subgoal coverage. Subgoals are generated procedurally (VLM keyframe extraction + language-conditioned image synthesis), so the conditioning distribution is denser than the demonstration distribution — the expert trains on subgoal states it has never actually reached.
- Canonical actions. The cross-embodiment frame transfers learned segments across morphologies without reparameterization.
3 · Experimental design & benchmarks
3.1 · The three chaotic tasks
- Task A — Cable untangling. Deformable cable (40–80 cm), 2–3 crossings, target knot-free. Zero-shot axes: bending modulus ×0.4–2.2, length, crossing topology, friction. Metrics: untangle success, crossing reduction, time, max jerk, violations.
- Task B — Textile folding on shifting fabric. Fold to a target polyline while the fabric shifts, so subgoal images go stale and must be re-anchored. Zero-shot axes: material classes, fold geometry, grasp visibility. Metrics: fold IoU ≥ 0.8, grasp success, regrasps, slippage events.
- Task C — Adaptive tool use in clutter. Retrieve/lever/pull with unknown tools. Zero-shot axes: tool morphology classes, clutter density, affordance composition. Metrics: task success, tool-selection accuracy, interventions per 100 episodes, contact-force violations.
3.2 · Data pipeline
(1) Teleoperation: 80–120 expert episodes per task on a bimanual rig, converted to the canonical frame. (2) Procedural simulation: an order of magnitude more episodes with aggressive domain randomization and ground-truth subgoal boundaries. (3) Synthetic subgoals: VLM keyframing of demos plus language-conditioned synthesis of novel subgoal images — the conditioning-coverage engine. (4) Data flywheel: deployment rollouts scored and curated; failures that came close are relabeled by the teleoperator and re-ingested — the loop whose curation strategy prior work showed decides whether the loop compounds.
3.3 · Pre-registered zero-shot protocol
Held-out morphologies, materials, and tool classes never appear in any training source — not sim, not teleop, not synthetic subgoals. Three seeds over fixed start sets; every number regenerates from a committed artifact. Intervention rate and max jerk are reported as costs, so a policy cannot win by being dangerous.
3.4 · Baselines and ablations
| Baseline | Description |
|---|---|
| RT-2-style | VLM → discretized actions, no hierarchy |
| Diffusion Policy | flat diffusion over action chunks |
| ACT | action-chunking transformer, no flow |
| π0-style vanilla | flow-matching VLA, no subgoals, no SMC, no filter |
| Hierarchical LLM | text-only subgoals (no visual keyframes) + flow expert |
| Ours (full) | visual subgoal hierarchy + SMC + CBF envelope |
| Ablation | Claimed effect | Primary metric |
|---|---|---|
| − SMC | corrections break smoothness | intervention rate, jerk |
| − safety filter | violations appear | safety violations |
| − visual subgoals | semantic-only conditioning degrades | zero-shot success |
| − dense subgoals | fails with horizon | success vs. horizon |
| − canonical frame | morphology transfer collapses | cross-morphology success |
The ablation set is the mechanism test: each component is claimed to move a distinct metric.