The Verification Gate

The same agent, identical plans and skills — the only difference is whether success is defined by executing the ground truth.
Sehaj Randhir Singh
Independent researcher; partial affiliation with NYU Tandon School of Engineering

Autonomous agents are usually evaluated by what they report, and agents that do not verify their work report success that never happened. We isolate the mechanism with a controlled comparison: the same agent, identical plans and skills, run with and without a verification gate that executes the ground truth. With the gate, 0 of 7 missions end in false success and 100% of injected faults are caught; without it, the same agent reports false success on 29% of missions and misses every fault. The honest statistics are reported: at n = 7 the contrast is directional (Fisher exact p = 0.23; Clopper-Pearson one-sided 95% upper bound 34.8% on 0/7), and a power analysis gives ~28 missions per condition for 80% power. The gate is a structural control — success is defined by executing the ground truth, not by self-report.

Why a gate and not a better model

Capability is not honesty. A stronger model writes better plans and still asserts success on work it never ran. The gate is a structural control: the benchmark reuses the AGE agent unchanged, and the only difference between the two conditions is whether the verify step is present.

0% vs 29%

false success with and without the gate on the same seven missions — the mechanism is the agent's self-report, not its capability.

100% caught

both injected faults are caught by the gate and missed without it. The gate's failure reports are exact, not sampled.

n = 7, stated

Fisher exact p = 0.23, Clopper-Pearson 95% upper bound 34.8% on 0/7, and ~28 missions per condition for 80% power — the honest statistics, not a spun direction.

The benchmark

Seven engineering missions (scaffold, injected-fault repair, beam design, cantilever design, Burgers design, clobber-guard, TODO scan), two carrying injected faults. The gate runs each mission's ground truth — tests, or an independent numeric verifier — and reports failure when it disagrees with the agent's claim.

Missionwith gatewithout gate
scaffold-okokok
injected-faultcaughtfalse success
beam-designokok
cantilever-designokok
burgers-designokok
clobber-guardcaughtfalse success
todo-scanokok
Both injected faults (injected-fault, clobber-guard) are reported as success by the un-gated agent; the gate reports failure on both. All five honest missions are reported correctly by both conditions.

Reproduce

git clone https://github.com/sehajr-singhs/verification-gated-agents
cd verification-gated-agents
node --test                      # 16 node tests incl. the gate benchmark
node bench/gate_bench.js         # re-run the 7-mission gate benchmark
python -m unittest tests.test_physx    # 49 physics tests (the solver the gate uses)

Simulation-only, CPU-scale, deterministic seeds. No GPU required.

Sister papers in the series

Seven manuscripts, one codebase, one guarantee: every number traces to a committed JSON and regenerates from a committed script.

AGE-artificial-general-engineer

the system and its project root

physics-transformers

the PhysFormer architecture; the falsified regime theory

physbench

the 12-domain verifiable benchmark

physics-loss-channel

when physics in the loss helps — and when it only enforces consistency

fewshot-law-acquisition

transfer across laws: what carries the knowledge

field-consistency

the cost of consistency on 2D fields