π Steerable VLA / research system
System Live NMI Manuscript Open Code + Results

Compositional Generalization in Embodied Foundation Models via Multimodal Subgoal Prompting and Flow-Based Action Execution

Generalist robot policies scale to thousands of tasks and still fail the test that defines embodied intelligence: a long-horizon task on an unseen object, with an unseen morphology, in a world that does not hold still. We build the system that doesn't — and prove its safety with a control barrier function that achieves zero violations.

0.0
Safety Violations
CBF–QP filter vs 5.9–19.9 baselines
75%
Expert Ceiling
1-crossing, 100% reproducible
4
Policy Variants
BC, Flow, Ours ± filter × 3 seeds
15p
NMI Manuscript
Algorithms, theorems, appendices
0.0 violations
Full system (CBF–QP filter active) across all seeds and episodes
5.9
BC baseline
18.3
Flow baseline
19.9
Ours − filter
The CBF–QP safety filter projects every commanded action onto the admissible set with discrete forward invariance (Proposition 1). The environment counts violations instead of clipping silently — so this is a faithful measure. The filter adds zero overhead to task performance: variants with and without the filter achieve the same intervention rate (1.16) and similar success (0.937 vs 0.947).
01 /

The thesis

Pure hardcoded logic fails because the real world has infinite variance. Pure neural nets fail because they lack spatial-temporal anchors. The way through is to structure the problem into a steerable hierarchy:

COMPOSITION

Subgoal factorization

Each flow segment conditioned on one visual keyframe. Skills compose across tasks in the conditioning, not inside a single weight matrix.

STEERABILITY

SMC conditioning layer

Corrections enter through bounded-Lipschitz gates. Grönwall bound certifies a computable envelope on trajectory deviation — no jerks, by construction.

VERIFICATION

CBF–QP safety envelope

Every commanded action projected onto the admissible set with discrete forward invariance. Steerability admitted only through the certified tube.

Theorem 1 — No-Jerk Guarantee

If $v_\theta^S$ is $L_v$-Lipschitz and the steering contribution is bounded by $M$, the deviation between steered and nominal trajectories satisfies:

‖x_S(s) − x_N(s)‖ ≤ (M / L_v)(e^{L_v Δs} − 1)

Corrections are continuous in flow time, and jerk is bounded by construction rather than tuned away.

02 /

Results

GPU study: NVIDIA T4, 200 epochs, 150 expert demonstrations, 3 seeds, 30 held-out episodes per variant-seed. Zero-shot protocol: train on crossing topologies {2, 3}, evaluate on held-out {4} with unseen stiffnesses. Every number traces to committed results/*.json.

VariantNI SuccessCr ↓Interv. ↓Violations ↓Success
BC (flat MLP)0.8331.000.175.91.000
Flow (no subgoals)0.3330.871.0318.30.967
Ours − filter0.4000.701.0019.90.933
Ours (full)0.5670.830.770.00.967

Expert ceiling: 75% on 1-crossing deviations (3 seeds × 8 episodes), 100% reproducible. Costs reported alongside success: interventions and jerk prevent policies from winning by being dangerous.

Safety violations — zero for full system
Fig 1. Safety violations per episode (mean ± std, 3 seeds, n=30). The CBF–QP filter eliminates all workspace violations.
System comparison
Fig 2. Full system comparison: assisted success rate (left) and safety violations (right) across all variants.
Expert ceiling
Fig 3. Expert policy ceiling: 75% on 1-crossing deviations, 100% reproducible.
No-intervention success
Fig 4. No-intervention success across variants. BC: 0.833, Ours (full): 0.567 — safety–utility tradeoff.
03 /

Architecture

Steerable VLA architecture diagram
Architecture. Two-level hierarchy: VLM planner emits dense visual subgoals; SMC layer steers the flow-matching expert with corrections; CBF–QP filter verifies every action before the actuators.
04 /

The system

Every stage of the hierarchy is implemented, tested, and runs end-to-end. 2,700 lines of Python across 26 modules. Every result regenerates from committed artifacts.

ENVIRONMENT

PBD cable simulator

Position-based dynamics with fixed segment lengths, free hinges, and localized constraint propagation. Tangles persist until the policy actively resolves them.

POLICY

Flow-matching expert + Transformer

CFM loss on continuous gripper deltas, Bernoulli grasp head, SMC steering. Plus 5.3M-param Transformer with self-attention over cable nodes.

SAFETY

CBF–QP filter

Discrete forward invariance via scipy SLSQP. Projects every action onto the safe workspace. 0 violations empirically; provable guarantee theoretically.

BASELINES

BC · ACT · Diffusion

Three proper baselines matching parameter budgets: behavioral cloning, action chunking transformer, and diffusion policy. All verified end-to-end.

FLYWHEEL

Data flywheel

Three curation strategies: none, near-miss, oracle relabeling (DAgger). Confirms bottleneck is training diversity, not data curation at miniature scale.

INFRASTRUCTURE

GPU · Kaggle · Lightning · HF

Full study runs on Kaggle T4 GPU, Lightning AI, or local. Hugging Face Space for live demo. GitHub Pages with committed results.

05 /

Manuscripts

NMI

Nature Machine Intelligence Format

15-page LaTeX manuscript with full Related Work section, Algorithm pseudocode (training + inference), Theorem 1 (no-jerk guarantee) + Corollary (jerk bound), Broader Impact section, hyperparameter appendix, and per-seed results table. Compiles with xelatex nmi_paper.tex.

Download PDF →

IEEE

IEEE Conference Format

Compact IEEEtran manuscript for CoRL/RSS/ICRA submission. Same core results and ablations in 4-page format.

Download PDF →

06 /

Project timeline

Phase 1
Environment
PBD cable simulator, expert oracle, crossing metrics
Phase 2
Policy + Safety
Flow expert, CBF filter, SMC steering, baselines
Phase 3
GPU Study
Kaggle GPU, 5 variants × 3 seeds, data flywheel
Phase 4
Publication
NMI paper, GitHub Pages, arXiv submission
07 /

Reproduce everything

# install
pip install -r requirements.txt

# smoke test (minutes on CPU)
PYTHONPATH=src python scripts/run_experiment.py --smoke

# full study: 5 variants × 3 seeds + flywheel + expert ceiling
PYTHONPATH=src python scripts/run_experiment.py --protocols all --seeds 3

# generate publication figures
PYTHONPATH=src python scripts/figures.py

# compile NMI paper (15 pages, with appendices)
cd docs/papers && xelatex nmi_paper.tex && xelatex nmi_paper.tex

# push to Kaggle GPU (or use web UI for GPU)
PYTHONPATH=src python scripts/build_kaggle_kernel.py --push

Every number in the manuscripts traces to committed JSON artifacts in results/. Seeds, splits, and held-out families are fixed at pre-registration time. The study completes in under 30 minutes on GPU (NVIDIA T4).