Rank Fidelity

When is it safe to optimise against a certificate?
Sehaj Randhir Singh
Independent researcher; partial affiliation with NYU Tandon School of Engineering

A certificate C is a bound: C(f) ≥ T(f) for every design f in a parametric family, where T is the true objective. Certificates are used not only to verify designs but to optimise over them — choosing the AMG preconditioner with the lowest certified bound, training the network with the highest certified accuracy, selecting the certified-optimal design. That use presupposes rank fidelity: the certificate's ranking of designs matches the truth's ranking. It need not. The slack S = C − T varies across the family, and when it does, the certificate can rank designs wrongly — even anti-rank them, as we measure for physics-informed neural networks (\\(\\rho_S = -0.96\\)). We prove the answer is governed by a single dimensionless quantity, the slack ratio \\(r = \\mathrm{Var}(S)/\\mathrm{Var}(T)\\), with a sharp phase transition at \\(r \\approx 1\\): for \\(r \\le 1\\) rank fidelity is perfect; for \\(r > 1\\) it degrades predictably as \\(\\rho_S \\sim g(r) = 1 - \\alpha\\,\\mathrm{ReLU}(\\tanh(\\beta \\ln r))\\). We validate across eight domains — algebraic multigrid, polynomial interval analysis, certified neural robustness, PINNs, HPC-scale 3D Navier–Stokes AMG, and a real free-boundary Grad-Shafranov tokamak equilibrium solver — and release an open-source diagnostic that determines, before any optimisation begins, whether a certificate can be trusted.

Eight domains, one phase transition

Every claim below traces to a committed JSON in results/. The CIFAR-10 row is the verified experiment run on Kaggle (kernel rank-fidelity-cpu-v10): 52 MLPs, certified accuracy by exact interval bound propagation, robust accuracy by 10-step PGD, \\(\\epsilon = 8/255\\).

Domain\\(\\rho_S\\) (obs)\\(p\\)Regime
AMG scalar1.000<10−16safe
AMG diagonal0.55–0.77<10−7transition
Polynomial interval0.5317×10−5transition
Linear IBP0.5857×10−3transition
PINN Burgers−0.959<10−4anti-ranked
3D NS AMG (HPC, 128³)0.380.23overflow
MLP CIFAR-100.3680.007transition
Tokamak (real GS)0.993<10−10safe
The PINN result is the paper's sharpest warning: when a certificate measures a proxy property structurally opposed to the true objective (interval width vs. local PDE residual), optimisation actively drives models away from truth. The tokamak result is the positive control computed with a real free-boundary Grad-Shafranov solver (FreeGS): a truncated coarse-grid equilibrium solve (\\\\(\\rho_S = 0.993, r = 0.017\\\\)) ranks configurations faithfully against the converged solve. The HPC-scale Oseen row shows the overflow regime: at slack ratio \\(r \\approx 10^7\\), rank fidelity collapses (\\\\(\\rho_S = 0.38\\\\)), indistinguishable from zero.
Goodhart effect and rank fidelity on CIFAR-10
Certified robustness, verified

Certified accuracy rises with IBP regularisation weight while true PGD robustness declines — the certificate–truth slack grows with optimisation pressure.

The phase transition g(r)
The phase transition

The calibrated curve \\(g(r)\\) against empirical Spearman across simulated slack ratios. The transition at \\(r = 1\\) separates safe from unsafe.

Why this matters: the Goodhart mechanism

Optimising against a certificate is a Goodhart game: the bound can be inflated without improving the truth. In the CIFAR-10 sweep, pushing the IBP regularisation weight from 0 to 20 raises certified accuracy from 0.11 to a plateau near 0.17–0.23 while true robust accuracy declines from 0.15 to 0.09. The bound never becomes invalid — it just stops ranking models the way the truth does. The same certificate is a faithful surrogate over one-parameter families and an unfaithful one over wider families; what flips it is the dimension of the family you descend over. The diagnostic makes this measurable before you commit to an optimisation run.

The design rule. Sample \\(N \\ge 30(1-\\rho_{ST})^{-1}\\) designs; evaluate \\(C\\) and \\(T\\); measure \\(\\rho_S = \\mathrm{Spearman}(C, T)\\). If \\(\\rho_S > 0.9\\), descend. If \\(0.5 < \\rho_S < 0.9\\), restrict to \\(\\le 10\\) candidates. If \\(\\rho_S < 0.5\\), the certificate is uninformative — use a different method.

The diagnostic, installable

pip install rank-fidelity

from rank_fidelity import Diagnostic

result = Diagnostic(certificate, truth).run(designs)
print(result.summary())
# n=52  rho_S=0.368 (p=0.0073)  r=4.97  rho_ST=-0.03
# regime: unsafe -> certificate is uninformative; use a different method

The package implements both theorems, bootstrap confidence intervals, permutation tests, and the sample-size rule. Full example in the repository.

Reproduce

git clone https://github.com/sehajr-singhs/rank-fidelity-paradox
cd rank-fidelity-paradox
pip install .
python -c "from rank_fidelity import Diagnostic; print('ok')"

# The verified CIFAR-10 experiment (52 MLPs) ran on Kaggle:
#   sehajrsingh/rank-fidelity-cpu-v10
# Per-model results: results/cifar10_verified_results.json

Every number in the paper traces to a committed JSON. The Kaggle kernel, notebook source, and figure scripts ship with the repo.