Physics Transformers

The equation is input, not only loss — and the pre-registered theory of when that pays was falsified on ten laws, in full.
Sehaj Randhir Singh
Independent researcher; partial affiliation with NYU Tandon School of Engineering

We present PhysFormer, a transformer adjusted for physics, and Law-Conditioned Attention (LCA), a new way to feed physics into it. The governing equation is tokenized into a fixed symbolic vocabulary, embedded into a law vector, and injected as a cross-attention key/value stream in every layer. Physics enters through two channels with different jobs: the loss channel (a physics-consistency layer) buys consistency, 6–8× residual reduction, without held-out accuracy; the input channel — the invention of this paper — is causally active at inference. Across 36 paired runs, LCA reduces trajectory error by 21% (p = 0.0003); swapping the equation signature steers the prediction (p < 0.0001) while a constant-signature control is exactly insensitive. We pre-registered a regime theory and falsified it on a ten-law suite (ρ = 0.07, p = 0.88), reporting the failure analysis. Ten dedicated per-law DeepONets reach median held-out error 0.037 vs. the single generalist's 0.110, which beats them outright on spring and LC.

Why the channel matters

The standard way to make a network physical is a loss term — and the standard way to stop measuring what it does. This paper separates the two effects a single number usually conflates: physics as a penalty (consistency) and physics as input (accuracy). Only the input channel moves held-out error, and it does so causally: remove the equation signature at inference and the prediction changes in a direction the data alone cannot explain.

Causal, p < 0.0001

Swapping the equation signature at inference steers the prediction; a constant-signature control is exactly insensitive. The equation is used, not decoratively present.

Falsified, ρ = 0.07

The pre-registered monotone regime theory failed (p = 0.88). The overlap measure conflates supersets with genuine indistinguishability — the failure analysis is the sharpest section.

0.037 vs 0.110

Ten dedicated per-law DeepONets beat the single generalist on per-law fidelity, but with no cross-law structure; the generalist wins outright on spring and LC.

The pre-registered regime test, in full

Before any ten-law training, the prediction was filed: LCA benefit should be monotone in the token-vocabulary ambiguity of each law, computed from the vocabulary alone. After 6 full trainings (3 seeds × generalist/control): Spearman ρ = 0.07 (p = 0.88), leave-one-out ρ = −0.58, and the group-mean order is violated. The pre-registration is committed at results/pre_registration.json; the eval files that falsified it are in the same directory.

LawAmbiguityLCA benefit
beam1.000+0.719
cantilever1.000+0.340
projectile0.500+0.141
pendulum1.000+0.539
spring1.000-0.407
rc0.667-0.049
damped0.750-0.013
kepler0.667-0.134
lc0.667-0.149
drag0.500+0.110
Measured benefit = 1 − err_real / err_dummy (median over three seeds). beam/cantilever present literally identical parameter tokens — only the equation distinguishes them — and show the largest, most consistent benefits.

The external baseline

DeepONet, the standard operator-network architecture, trained per law on the identical data splits (64 training samples, 6 held out, 250 epochs).

Lawper-law DeepONet error
beam0.025
cantilever0.009
projectile0.011
pendulum0.037
spring0.122
rc0.007
damped0.141
kepler0.066
lc0.192
drag0.011
A dedicated operator network wins on its own law; a single law-conditioned transformer covers all ten. The pooled single-model DeepONet (law identity as an explicit one-hot input — information the generalist never receives) did not complete under available compute and is reported as an attempt, not a result.

Reproduce

git clone https://github.com/sehajr-singhs/physics-transformers
cd physics-transformers
python -m unittest tests.test_physx          # 49 physics tests, 10-law set
python figs/make_figures.py && python figs/make_figures_ext.py
python src/physx/regime_oos.py --out figs/regime_oos.json   # the falsification
python src/physx/train_multi.py --ext --seeds 3             # real + dummy, 3 seeds

Simulation-only, CPU-scale, deterministic seeds. No GPU required.

Sister papers in the series

Seven manuscripts, one codebase, one guarantee: every number traces to a committed JSON and regenerates from a committed script.

AGE-artificial-general-engineer

the system and its project root

physbench

the 12-domain verifiable benchmark

verification-gated-agents

the gate as the missing control in agent evaluation

physics-loss-channel

when physics in the loss helps — and when it only enforces consistency

fewshot-law-acquisition

transfer across laws: what carries the knowledge

field-consistency

the cost of consistency on 2D fields