Identity Through Interruption

A motion-signature-first decision architecture for multi-object re-identification — staged, hypothesis-holding identity decisions with gait, trajectory-style, and geometry as primary signals, and appearance demoted to a supporting, auditable gate.

Sehaj Randhir Singh · 2026 · simulation + MOT17 public-detection benchmark

0
ID switches, 20 seeds
statistically zero through every constructed interruption — identical-twin crossings, re-entry, the appearance-free ball — with at most one switch on a single occlusion seed
−21%
MOT17 ID switches
1,265 → 995 from the pose-free group-level pile-up mechanism, no keypoints needed
−47%
switches when the CNN gate is removed
995 → 532, with IDF1 rising — the uncalibrated appearance channel was net-harmful
0.7 pp
remaining IDF1 gap to SORT
under a real low-recall detector (MOT17-04), closed by the velocity-consistency gate (IDSW 33 → 15)

The problem

Appearance fails exactly where identity matters most. Person re-identification is usually done by comparing appearance — clothing, face, build. That works when people look different, and fails precisely in the cases where identity is load-bearing: two pedestrians in identical uniforms, a person occluded behind a pillar, a target exiting the frame and re-entering elsewhere, or an object with no stable appearance at all (a ball). This work takes the complementary position: identity is a decision problem, and motion is a first-class identity signal. A decision layer routes every observation through gate → structural → motion-signature → bind/spawn, and — unlike conventional trackers — declines to guess under ambiguity: it holds competing hypotheses open for a bounded window until motion evidence resolves them. Every live ID emits a box every frame, observed where bound and ghost-predicted while hidden, so identity continuity becomes a measurable property of a path, not a label match. The result is deliberately two-sided: statistically zero identity switches through every interruption we constructed, at a measured cost in dense-scene recall; and on real surveillance data the motion-side mechanisms transfer — a pose-free group-level tracker cuts MOT17 identity switches by 21% — while uncalibrated appearance channels actively harm every tracker that trusts them, ours included, until calibrated.

Watch: identity kept through the interruptions

Simulated streams with persistent-ID overlays. Solid boxes are observed detections; dashed boxes are the belief stream — the ghost-predicted trajectory emitted while the target is hidden, so the path never teleports.

Identical twins crossing — the scenario appearance cannot resolve; the geometric gate re-binds before gait is consulted.
Occlusion behind a pillar — ghost prediction carries the identity through the hidden frames (0 switches in 19/20 seeds, at most 1).
Out-of-frame re-entry — the target exits, is edge-parked, and re-binds at a different position on return — 0 switches in 20/20 seeds.
The appearance-free ball — re-bound purely by projectile-physics prediction (hard IDF1 0.817 vs 0.668 best appearance baseline).
Pile-up — co-located identities are tracked as a group through the merged blob; each member re-binds on split. No re-spawn cascade.
Dense repeated merges — the hypothesis-holding cliff: removing it multiplies dense IDSW ≈5.7× (14.9 → 84.4).
Control (distinct walkers) — the appearance channel is genuinely informative here, and the system uses it as a supporting gate.

The architecture

System architecture: two loops + decision layer + observation and belief streams
Two-loop online architecture. A high-frequency loop (SORT-lineage temp tracks + ghost prediction with growing uncertainty) and a low-frequency loop (rolling pose buffers, cadenced ST-GCN gait signatures) feed a decision layer that binds, spawns, or holds hypotheses. Outputs: an observation stream (MOT-comparable) and a belief stream (a continuous predicted trajectory for every non-parked ID, every frame).

Synthetic hard segments: zero switches where appearance breaks

Hard-segment IDF1 / IDSW, 7 scenarios × 20 seeds (identical detection streams for all trackers). Bold = best IDSW; the paper reports both wins and losses honestly.

scenariomotion_reid (ours)SORTDeepSORTByteTrack
crossing (twins)0.929 / 0.00.893 / 0.70.927 / 0.20.940 / 0.1
occlusion0.924 / 0.10.321 / 3.20.937 / 0.10.330 / 2.8
reentry0.912 / 0.00.904 / 0.10.951 / 0.00.904 / 0.1
ball (no appearance)0.817 / 0.00.487 / 3.90.643 / 0.00.668 / 0.5
dense0.467 / 14.90.393 / 40.20.547 / 19.90.425 / 34.9
pileup0.430 / 5.20.557 / 7.00.605 / 7.80.574 / 6.9
distinct (control)0.678 / 10.30.605 / 12.00.824 / 2.80.627 / 12.1

motion_reid and DeepSORT are the only trackers that stay at (statistically) zero IDSW through every interruption. The wins are reported with their costs: hypothesis-holding trades dense-scene IDF1 (0.467/0.430 vs DeepSORT 0.547/0.605) for dramatically fewer switches — a deliberate, quantifiable trade.

ID switches per scenario per tracker
Synthetic IDSW by scenario. The “zero through interruption” panel — SORT and ByteTrack switch on occlusion in every seed; our hard-segment switches are at most one, on one seed.
Ablation IDSW: hypothesis-holding cliff
Ablations. The hypothesis-holding cliff (dense 14.9 → 84.4 without it) and the group-level cluster contrast (full vs −cluster).
Belief stream vs observation stream on re-entry
Tracking, not teleportation. On re-entry, the observation stream fragments into 34 interruptions (shaded); the belief stream stays continuous — 0 interruptions, 0 teleports, easing into re-emergence.
MOT17 hybrid IDSW per run
MOT17 hybrid. The cluster mechanism cuts IDSW 1,265 → 995 (−21%, no pose); removing the uncalibrated CNN appearance gate cuts 995 → 532 (−47%, IDF1 rising); real-skeleton gait leaves the tracker nearly unchanged.

Real-data transfer (MOT17, all seven train sequences)

Hard-segment totals on the identical FRCNN public-detection stream, pose inherited from YOLOv8-pose, appearance from ImageNet-pretrained MobileNet-v3 (DeepSORT receives the same embeddings).

runIDF1IDSWinterpretation
full0.324995both channels live + cluster tracking
+ real-skeleton gait0.3261,000robust to real skeletons (p = 0.75); pose coverage 0–41%
−gait0.3221,039gait positive but small
−appearance0.332532uncalibrated CNN gate → ≈47% of switches, IDF1 rises
−trajectory0.3261,019small positive
−cluster0.3001,265clusters remove 21% of switches, no pose needed
SORT (same stream)0.493263ignores appearance
ByteTrack (same stream)0.496235ignores appearance
DeepSORT (same stream)0.458309trusts the uncalibrated embedding — degraded from 0.522 box-only

Two measured asymmetries carry the real-data story. The architecture transfers; the uncalibrated channels do not. The pose-free cluster mechanism cuts MOT17 identity switches 1,265 → 995 (−21.3%) and raises hard-IDF1 0.300 → 0.324 (MOT17-04 −47.6%, MOT17-11 −30.0%), and the belief stream reports zero interruptions and zero teleports on all seven real sequences. Meanwhile an ImageNet-generic MobileNet-v3 embedding at ≈34 px crop scale is not an identity embedding: used as an appearance gate it produced 463 of 995 switches (47%), and removing it raised IDF1 — the same finding holds for DeepSORT (0.522 → 0.458). The paper formalizes the channel-calibration audit: a channel is flagged untrustworthy if removing it does not lower IDF1. Both appearance channels fail it; appearance-free trackers are invariant by construction.

Detector recall ceiling on 1080p MOT17
Detector-recall ceiling. A real YOLOv8 on 1080p MOT17 recalls only ~34% of visible pedestrians (median person ≈34 px tall). No tracker escapes the detection floor.
Pose-coverage headroom: gait benefit collapses in the real coverage band
Pose-coverage headroom. The gait channel's benefit collapses inside the real MOT17 pose-coverage band (0–41%): the tracker was never gait-limited — it is pose-coverage-limited.

The honest numbers

Every number on this page and in the papers is regenerable from the committed JSONs in results/. python audit_numbers.py re-verifies all 131 headline checks against the ground-truth files.

Papers

versiontargetformatstatus
NMI manuscriptNature Machine Intelligence (Article)~3,500-word main text, ≤6 display items + Supplementarypre-submission draft
WACV 2027 paperIEEE/CVF Winter Conf. on Applications of Computer Visionofficial WACV template, 8 pages + references, anonymized review modeready for Round 2 (Aug 28, 2026)
IEEE paperIEEE journal / conference trackIEEEtran, 7 pagessubmission-ready
IEEE TMM paperIEEE Trans. on Multimedia (Regular Paper)IEEEtran journal, 7 pages (limit 10), single-blindready for ScholarOne submission
ICIP 2027 paperIEEE Int. Conf. on Image ProcessingIEEEtran conference, 4 pages (limit 5+1), double-blindready for Ex Ordo submission

BibTeX

@article{singh2026identity,
  title   = {Identity Through Interruption: A Motion-Signature-First Decision
             Architecture for Multi-Object Re-Identification},
  author  = {Singh, Sehaj Randhir},
  year    = {2026},
  note    = {Pre-submission draft; all measurements reproducible from
             committed result JSONs},
  url     = {https://github.com/sehajr-singhs/motion-reid}
}

Reproduce

python train_gait.py            # ST-GCN triplet model (results/gait_model.pt)
python benchmark.py             # synthetic benchmark (results/benchmark.json)
python ablate.py                # ablations (results/ablation.json)
python make_demo.py             # belief-stream videos (results/demos/*.mp4)
python real_benchmark.py        # MOT17 box-only, 7 seqs
python real_hybrid_benchmark.py # MOT17 hybrid + --ablate variants
python stress_lowrecall.py      # detector-stress + recall sweep
python real_selfcal.py          # self-calibrated transfer constants
python audit_channels.py        # channel-calibration audit
python audit_numbers.py         # re-verify all 131 headline numbers
python -m unittest discover -s tests   # 27 unit tests (core + MOT17 loader + continuity)

Prepared with assistance from opencode / DeepSeek-V4 tooling. Thanks to Vikram Kapila and Karthik Voruganti for guidance and feedback on the manuscript.
MOT17: public-detection protocol, all seven train sequences, deterministic CPU runs. Every number traces to a committed JSON.