EB‑H‑JEPA

Latent predictors for world models must dissipate. A structural failure atlas and the conditions for engaged dissipation — with three provable predictions.
Sehaj Randhir Singh
Independent researcher; partial affiliation with NYU Tandon School of Engineering

World models that operate in a learned latent space (JEPA; LeCun 2022) are the leading bet for sample-efficient, long-horizon model-based RL (DreamerV3, Hafner et al. 2023). A recurring hypothesis is that imposing physical structure on the latent predictor — Hamiltonian, metriplectic, or GENERIC — should make learned dynamics faithful and stable. We test this hypothesis twice, with instruments of increasing fidelity. First, on a nonlinearly lifted Lorenz-63 system, we audit five structurally distinct latent predictors under a chaos-aware protocol. The result is a failure atlas: each constraint class fails, but on a property disjoint from the one it enforces. A rigid Hamiltonian predictor freezes and cannot dissipate; a naive metriplectic predictor sits on a dead \\(LL^\\top\\) saddle with dissipation pinned at exactly zero; fixing both bugs (constant divergence-free \\(J\\), nonzero \\(R\\) initialization) makes dissipation engage and recovers the true phase-space contraction \\(-13.667\\) to \\(\\sim1\\%\\) accuracy — yet spectral alignment still cannot pin individual Lyapunov exponents. Second, and constructively, we insert the fixed-metriplectic predictor into a compact DreamerV3-style agent and measure sample efficiency and latent stability on Crafter (Hafner 2021), the benchmark of the world-model group that introduced RSSM, with matched 32-bit free-bits budgets and three seeds per arm. The controlled result is a latent-stability finding with matched returns: the categorical RSSM path collapses to its prior within 2k steps (raw KL 0.31 bits, 1% of its budget; drift exactly at the independent-uniform noise floor), while the metriplectic latent stays informative and bounded (raw KL 34.1 bits, 106% of budget; norm 1.49) with dissipation engaged and rising. Greedy evaluation returns sit at the same floor in both arms — structure neither helps nor hurts sample efficiency at this budget, and the stability is free. All code, runs, and evidence are public; every claim traces to a JSON artifact.

The failure atlas: every constraint class fails, on a property disjoint from the one it enforces

We train five structurally distinct latent predictors on Lorenz-63 lifted into a 64-D latent chart, and audit each under a chaos-aware protocol: rollout boundedness, leading Lyapunov exponent, phase-space contraction, attractor overlap, and dissipation engagement. The headline is more precise than "structure doesn't help": each class fails where it is not enforcing anything.

ArmStructureλ₁ (truth ≈ 0.906)tr(R)/n (dissipation)contraction vs −13.667reading
A · unconstrainednone+5.75n/an/aexplodes off the attractor
B · rigid Hamiltonianconstant J, learned H0.000 (by construction)n/afreezes: quasi-periodic tube, near-zero motion
C · naive metriplecticR = LL⊤, zero init~00.000000.0dead saddle: L stuck at 0, no dissipation
D · fixed metriplecticconst. J + nonzero R₀0.810.49−13.80 ± 0.17R engages; true contraction recovered to ~1%
E · fixed + spectral shapingD + QR Lyapunov penalty0.900.53−13.63closest λ₁; still cannot pin individual exponents
3 seeds per arm; every number traces to results/*.json via benchmarks/run_benchmark.py. Same failure modes confirmed on Kuramoto–Sivashinsky (34 DOF, GPU). The cruelest finding: the frozen Hamiltonian arm (B) posts the best attractor-overlap score while being dynamically inert — overlap rewards dead models.
The atlas, one picture

Five structural constraint classes, one chaotic test bed. Each fails on a property disjoint from the one it enforces.

Contraction, recovered

The fixed-metriplectic arm engages R and recovers Lorenz's exact phase-space contraction −13.667 to ~1% — a mechanistic fingerprint, not a curve fit.

Theory: when structure collapses and when it engages

Three propositions predict the empirical findings. These were formulated before the Crafter runs.

Proposition 1 (Posterior collapse basin). The categorical RSSM posterior converges to its prior whenever the GRU captures less than \(\beta / \log_2 K\) of the per-code predictability. For DreamerV3 defaults (\(\beta = 1\), \(K = 32\)), collapse occurs whenever the GRU captures less than 20% of the information. This is precisely the regime we observe in Crafter at 20k steps.

Proposition 2 (Dead saddle). The metriplectic dissipation matrix \(R = LL^\top\) sits on a dead saddle at \(L = 0\): \(\nabla_L R = 0\), so no gradient can lift \(R\) off zero. The fix — nonzero \(L\) initialization at std 0.05 — breaks this saddle and makes dissipation engage.

Proposition 3 (Drift noise floor). For two independent samples from a uniform \(K\)-class prior, the scale-free drift is exactly \(2(1 - 1/K)\). For \(K = 32\): 1.9375. This is precisely the measured RSSM drift — it is sampling noise, not predictive error.

Two chaotic systems. We validate on Lorenz-63 (3 DOF, \(\mathrm{div}\,F = -13.667\)) and Kuramoto–Sivashinsky (34 real DOF with trace-free nonlinearity, \(\mathrm{div}\,F = -4064.9\)). The same failure modes recur: rigid Hamiltonians freeze, naive metriplectics sit on the dead saddle, fixed metriplectics engage dissipation. The Kuramoto–Sivashinsky atlas (5 arms × 3 seeds, GPU) is running on Kaggle now.

Latent stability in the RL loop: DreamerV3 with a pluggable predictor

A failure atlas on Lorenz tells us where structure fails in isolation; it does not tell us whether structure helps or hurts an actual RL world model, where the predictor co-adapts with an actor trained by latent imagination. We build a compact but faithful reimplementation of the DreamerV3 recipe — CNN encoder/decoder, recurrent memory, categorical posterior, KL free bits, symlog targets, TD(λ) imagination — in which the stochastic prior is pluggable. The baseline is DreamerV3's exact categorical RSSM; the treatment is a continuous latent whose prior mean advances under the fixed-metriplectic map from Part 1. Everything else is byte-identical, including a matched total KL free-bit budget (32 bits per step for both arms).

Beyond return, we instrument the thing the paper's title cares about: latent stability. Every 2k environment steps both arms log prior-vs-posterior drift, raw KL (free-bit slack), latent norm, and — for the treatment arm — dissipation engagement \\(\\mathrm{tr}(R)/n\\) inside the RL loop.

metricDreamerV3 RSSMfixed-metriplectic prior
eval return @ 20k steps−0.90 ± 0.00−0.90 ± 0.00
prior–posterior drift (scale-free)1.94 ± 0.00 †1.71 ± 0.09
raw KL (free-bit slack)0.31 ± 0.01 — 1% of 32-bit budget34.06 ± 0.44 — 106% of budget
latent norm0.031 (exactly uniform one-hot)1.49 ± 0.02
dissipation tr(R)/n (treatment only)n/a0.008, rising 0.0065 → 0.0080
Reading. The RSSM stochastic path collapses: raw KL is 1% of its matched free-bits budget, the latent norm is exactly the uniform-one-hot expectation, and its drift sits precisely on the closed-form noise floor for two independent uniform one-hot samples — 2(1−1/32) = 1.94 — so its "drift" is sampling noise, not predictive error. The metriplectic latent stays live: KL above budget (posterior doing real correction), bounded norm, dissipation engaged and rising inside the RL loop — all at zero return cost (returns matched at the −0.90 floor; neither compact arm has a functioning greedy policy at 20k steps). Every number traces to the committed JSONs in results/crafter/.
Crafter stability curves: matched returns, RSSM drift on the noise floor, metriplectic drift below it

Reproduce

Offline env + code dataset · 3-seed × 2-arm runs (T4)

git clone https://github.com/sehajr-singhs/ebh-jepa
cd ebh-jepa
pip install -r requirements.txt
python -m pytest tests/ -q                     # smoke tests: Lorenz pipeline + audit

# Part 1 — the failure atlas on Lorenz-63 (CPU, ~15 min)
python benchmarks/run_benchmark.py --steps 250

# Part 2 — the Crafter RL experiment (GPU; or the Kaggle T4 kernel)
python experiments/crafter/train.py --predictor rssm        --env crafter --train-steps 20000
python experiments/crafter/train.py --predictor metriplectic --env crafter --train-steps 20000

# Analysis → Table 2 + stability curves
python experiments/crafter/analyze.py --results results/crafter_*.json

Seeded protocol (3 seeds per condition), committed result JSONs, an offline Kaggle dataset so the exact experiment environment ships with the repo. No GPU required for Part 1; Part 2 runs in ~1h on a free T4.