Adding a physics residual to the loss is the standard way to make a network physical — and the standard way to stop measuring what it does. This study isolates the loss channel on a controlled matrix: identical architectures, data, and training, with the physics term on or off. The physics term cuts the governing-equation residual 19× (p < 1e-7, Cliff's δ = −0.92) — it genuinely enforces consistency — while pooled held-out accuracy stays flat (p = 0.65) and one domain (projectile) gets worse. The full 75-run matrix (5 domains × 3 architectures × 5 seeds) is committed and re-runnable.
A single 'physics-informed' number usually conflates two things: does the physics term make predictions more consistent with the governing equation, and does it make them more accurate? This study measures both, on identical architectures and data, with the physics term as the only difference.

governing-equation violation of predicted trajectories collapses, p < 1e-7, Cliff's δ = −0.92 — the loss channel genuinely enforces consistency.

pooled held-out error is statistically unchanged, and the direction varies by domain — beam improves modestly, projectile worsens.
the whole matrix regenerates from committed stats files; the pooled statistics are exact permutation tests, not asymptotic.
Three model kinds (physics-in-the-loss transformer, no-physics transformer, MLP head) × five domains × five seeds, pooled paired statistics (Wilcoxon signed-rank with exact permutation, Cliff's delta).
| Domain | phys | nophys | MLP | p | δ |
|---|---|---|---|---|---|
| beam | 0.147 | 0.193 | 0.156 | 0.438 | -0.20 |
| cantilever | 0.148 | 0.204 | 0.152 | 0.188 | -0.60 |
| projectile | 0.108 | 0.076 | 0.065 | 0.125 | +0.60 |
| burgers | 0.009 | 0.010 | 0.018 | 0.438 | -0.20 |
| heat2d | 0.015 | 0.014 | 0.016 | 0.625 | +0.60 |
git clone https://github.com/sehajr-singhs/physics-loss-channel cd physics-loss-channel python -m unittest tests.test_physx # 49 physics tests python figs/make_figures.py # figures from committed significance.json python figs/lca_significance.py # re-run the pooled statistics from results/
Simulation-only, CPU-scale, deterministic seeds. No GPU required.
Seven manuscripts, one codebase, one guarantee: every number traces to a committed JSON and regenerates from a committed script.
the system and its project root
the PhysFormer architecture; the falsified regime theory
the 12-domain verifiable benchmark
the gate as the missing control in agent evaluation
transfer across laws: what carries the knowledge
the cost of consistency on 2D fields