π Steerable VLA / research proposal
Research proposal · sections 1–3

The proposal

1 · Title & abstract

Compositional Generalization in Embodied Foundation Models via Multimodal Subgoal Prompting and Flow-Based Action Execution — prepared for Nature Machine Intelligence format.

Generalist robot policies have scaled to thousands of tasks, yet they still fail the test that defines embodied intelligence: performing a long-horizon task on an unseen object, with an unseen morphology, in a world that does not hold still. Cables tangle in ways no training set contains; textiles shift under their own weight; a hook that was never seen must still be used. Scaling monolithic vision-language-action (VLA) weights alone does not produce compositional generalization: the policy either memorizes the data distribution or loses the spatial-temporal anchors that make execution reliable.

We argue that zero-shot generalization across unseen morphologies and chaotic manipulation tasks is achieved not by scaling a single network, but by structuring the problem into a steerable hierarchy: a high-level VLM decomposes the task into dense visual subgoals, a flow-matching action expert executes the segments between them at control frequency, the Steerable Multimodal Conditioning (SMC) layer admits mid-execution corrections with a provable Lipschitz bound on trajectory deviation, and a runtime formal-verification safety envelope (a control-barrier-function filter) certifies every executed action. The program is instantiated on three chaotic benchmarks — cable untangling, textile folding on shifting fabric, and adaptive tool use in clutter — under a pre-registered zero-shot protocol, so that every number in the resulting manuscripts traces to a committed experiment artifact.

Contributions. (1) a steerable two-level hierarchy in which composition happens at the subgoal level, not inside a single weight matrix; (2) the SMC layer with a formal no-jerk guarantee — steering without jerks, by construction; (3) a CBF–QP safety envelope with forward-invariance guarantees and offline falsification; (4) three chaotic benchmarks and a pre-registered zero-shot protocol with intervention rate and jerk as first-class metrics.

2 · System architecture & mathematical formulation

2.1 · Two levels, one shared token space

Observations are multi-view RGB-D plus proprioception (and optional tactile); instructions are language. The planner Λφ maps (ℓ, ot) to a dense subgoal sequence g1:K = {(κk, ℓk)}, where κk is a keyframe image token (decoded to pixels for conditioning) and ℓk its language step. Dense coverage means K scales with task horizon: an untangling with three crossings gets roughly six to ten subgoals.

Text, keyframes, and runtime constraints are fused in a shared token space: the same cross-attention layers in the VLM decoder project language tokens, vision tokens, and steering tokens onto the action-latent stream (π-series parameterization). A contrastive alignment loss pairs each subgoal token with the observation chunk achieved at its segment boundary, so the two levels agree on what “done” means — the condition for the factorization to compose.

Actions live in a canonical cross-embodiment frame: normalized end-effector deltas (Δpee, Δψ) plus gripper, converted from joint space by forward kinematics. One flow expert, many morphologies, zero-shot.

2.2 · Flow matching over action chunks

Let x1 be a demonstrated chunk and x0 an independent noise sample. With the linear interpolation xs = (1−s)x0 + sx1, conditional flow matching trains the velocity field:

ℒCFM(θ) = 𝔼s∼𝒰[0,1], (x₀,x₁)∼q, c ‖ vθ(xs, s, c) − (x1 − x0) ‖² (2)

Inference integrates the ODE dxs/ds = vθ(xs, s, c) with 2–10 Euler/Heun steps, executed with a receding horizon at 50 Hz. Flow matching is preferred over score-based diffusion because few-step sampling preserves smoothness at control frequency, the marginal path is stable to train at scale, and the ODE formulation yields a velocity field that can be analyzed — the substrate for steering and safety.

2.3 · The Steerable Multimodal Conditioning layer

Steering signals — human corrections (clicked keypoints, drag vectors), world-model replans (stale subgoals replaced), and runtime constraint setpoints (force ceilings, keep-out zones) — enter through a gated adapter on the velocity field:

vθS(xs, s, c, u) = vθN(xs, s, c) + Σj λj(xs, s) · vθj(xs, s, uj) (3)

with the anchoring constraint vθj(·, ·, ·, uj0) ≡ 0 — removing steering recovers the nominal field exactly — and gates λj = σ(wjTφ(xs, s)) with bounded Lipschitz constant Lλ, so signals engage softly over a window rather than as a step in action space.

The no-jerk theorem. If the steered field is Lv-Lipschitz in x and the steering contribution is bounded by M, then over a steering window of flow-time Δs,

‖ xS(s) − xN(s) ‖ ≤ MLv ( eLvΔs − 1 ) — Grönwall (4)

a computable smoothness envelope: an operator bounds a priori how far a correction may move the trajectory. Jerk — max over time of ‖d³pee/dt³‖ — is bounded by the same quantity divided by the gate rise time; steering is continuous in flow time by construction, and the bound is verified numerically during training.

2.4 · The runtime formal-verification safety envelope

Let the safe set be 𝒞 = {x : h(x) ≥ 0} for a control barrier function encoding joint limits, collision margins, and contact-force ceilings. At each control step the filter solves a convex QP:

u* = argminu∈𝒰 ½‖u − ucmd‖²W  s.t.  ḣ(x) + αh(x) ≥ 0 (5)

Standard CBF theory gives forward invariance: if h(x(0)) ≥ 0, then h(x(t)) ≥ 0 for all t — the robot provably never leaves the safety set. The composition is the architectural point: steerability is admitted only through the filter, so corrections reshape the trajectory inside the certified tube and never outside it. The QP solves in microseconds inside the 50 Hz loop; the envelope is falsified offline over steering amplitudes and timings before hardware ever sees it.

2.5 · Why the structure composes

3 · Experimental design & benchmarks

3.1 · The three chaotic tasks

3.2 · Data pipeline

(1) Teleoperation: 80–120 expert episodes per task on a bimanual rig, converted to the canonical frame. (2) Procedural simulation: an order of magnitude more episodes with aggressive domain randomization and ground-truth subgoal boundaries. (3) Synthetic subgoals: VLM keyframing of demos plus language-conditioned synthesis of novel subgoal images — the conditioning-coverage engine. (4) Data flywheel: deployment rollouts scored and curated; failures that came close are relabeled by the teleoperator and re-ingested — the loop whose curation strategy prior work showed decides whether the loop compounds.

3.3 · Pre-registered zero-shot protocol

Held-out morphologies, materials, and tool classes never appear in any training source — not sim, not teleop, not synthetic subgoals. Three seeds over fixed start sets; every number regenerates from a committed artifact. Intervention rate and max jerk are reported as costs, so a policy cannot win by being dangerous.

3.4 · Baselines and ablations

BaselineDescription
RT-2-styleVLM → discretized actions, no hierarchy
Diffusion Policyflat diffusion over action chunks
ACTaction-chunking transformer, no flow
π0-style vanillaflow-matching VLA, no subgoals, no SMC, no filter
Hierarchical LLMtext-only subgoals (no visual keyframes) + flow expert
Ours (full)visual subgoal hierarchy + SMC + CBF envelope
AblationClaimed effectPrimary metric
− SMCcorrections break smoothnessintervention rate, jerk
− safety filterviolations appearsafety violations
− visual subgoalssemantic-only conditioning degradeszero-shot success
− dense subgoalsfails with horizonsuccess vs. horizon
− canonical framemorphology transfer collapsescross-morphology success

The ablation set is the mechanism test: each component is claimed to move a distinct metric.