A motion-signature-first decision architecture for multi-object re-identification — staged, hypothesis-holding identity decisions with gait, trajectory-style, and geometry as primary signals, and appearance demoted to a supporting, auditable gate.
[Code] [NMI manuscript] [WACV 2027] [IEEE] [BibTeX]
Appearance fails exactly where identity matters most. Person re-identification is usually done by comparing appearance — clothing, face, build. That works when people look different, and fails precisely in the cases where identity is load-bearing: two pedestrians in identical uniforms, a person occluded behind a pillar, a target exiting the frame and re-entering elsewhere, or an object with no stable appearance at all (a ball). This work takes the complementary position: identity is a decision problem, and motion is a first-class identity signal. A decision layer routes every observation through gate → structural → motion-signature → bind/spawn, and — unlike conventional trackers — declines to guess under ambiguity: it holds competing hypotheses open for a bounded window until motion evidence resolves them. Every live ID emits a box every frame, observed where bound and ghost-predicted while hidden, so identity continuity becomes a measurable property of a path, not a label match. The result is deliberately two-sided: statistically zero identity switches through every interruption we constructed, at a measured cost in dense-scene recall; and on real surveillance data the motion-side mechanisms transfer — a pose-free group-level tracker cuts MOT17 identity switches by 21% — while uncalibrated appearance channels actively harm every tracker that trusts them, ours included, until calibrated.
Simulated streams with persistent-ID overlays. Solid boxes are observed detections; dashed boxes are the belief stream — the ghost-predicted trajectory emitted while the target is hidden, so the path never teleports.
Hard-segment IDF1 / IDSW, 7 scenarios × 20 seeds (identical detection streams for all trackers). Bold = best IDSW; the paper reports both wins and losses honestly.
| scenario | motion_reid (ours) | SORT | DeepSORT | ByteTrack |
|---|---|---|---|---|
| crossing (twins) | 0.929 / 0.0 | 0.893 / 0.7 | 0.927 / 0.2 | 0.940 / 0.1 |
| occlusion | 0.924 / 0.1 | 0.321 / 3.2 | 0.937 / 0.1 | 0.330 / 2.8 |
| reentry | 0.912 / 0.0 | 0.904 / 0.1 | 0.951 / 0.0 | 0.904 / 0.1 |
| ball (no appearance) | 0.817 / 0.0 | 0.487 / 3.9 | 0.643 / 0.0 | 0.668 / 0.5 |
| dense | 0.467 / 14.9 | 0.393 / 40.2 | 0.547 / 19.9 | 0.425 / 34.9 |
| pileup | 0.430 / 5.2 | 0.557 / 7.0 | 0.605 / 7.8 | 0.574 / 6.9 |
| distinct (control) | 0.678 / 10.3 | 0.605 / 12.0 | 0.824 / 2.8 | 0.627 / 12.1 |
motion_reid and DeepSORT are the only trackers that stay at (statistically) zero IDSW through every interruption. The wins are reported with their costs: hypothesis-holding trades dense-scene IDF1 (0.467/0.430 vs DeepSORT 0.547/0.605) for dramatically fewer switches — a deliberate, quantifiable trade.
Hard-segment totals on the identical FRCNN public-detection stream, pose inherited from YOLOv8-pose, appearance from ImageNet-pretrained MobileNet-v3 (DeepSORT receives the same embeddings).
| run | IDF1 | IDSW | interpretation |
|---|---|---|---|
| full | 0.324 | 995 | both channels live + cluster tracking |
| + real-skeleton gait | 0.326 | 1,000 | robust to real skeletons (p = 0.75); pose coverage 0–41% |
| −gait | 0.322 | 1,039 | gait positive but small |
| −appearance | 0.332 | 532 | uncalibrated CNN gate → ≈47% of switches, IDF1 rises |
| −trajectory | 0.326 | 1,019 | small positive |
| −cluster | 0.300 | 1,265 | clusters remove 21% of switches, no pose needed |
| SORT (same stream) | 0.493 | 263 | ignores appearance |
| ByteTrack (same stream) | 0.496 | 235 | ignores appearance |
| DeepSORT (same stream) | 0.458 | 309 | trusts the uncalibrated embedding — degraded from 0.522 box-only |
Two measured asymmetries carry the real-data story. The architecture transfers; the uncalibrated channels do not. The pose-free cluster mechanism cuts MOT17 identity switches 1,265 → 995 (−21.3%) and raises hard-IDF1 0.300 → 0.324 (MOT17-04 −47.6%, MOT17-11 −30.0%), and the belief stream reports zero interruptions and zero teleports on all seven real sequences. Meanwhile an ImageNet-generic MobileNet-v3 embedding at ≈34 px crop scale is not an identity embedding: used as an appearance gate it produced 463 of 995 switches (47%), and removing it raised IDF1 — the same finding holds for DeepSORT (0.522 → 0.458). The paper formalizes the channel-calibration audit: a channel is flagged untrustworthy if removing it does not lower IDF1. Both appearance channels fail it; appearance-free trackers are invariant by construction.
Every number on this page and in the papers is regenerable from the committed JSONs in results/. python audit_numbers.py re-verifies all 131 headline checks against the ground-truth files.
| version | target | format | status |
|---|---|---|---|
| NMI manuscript | Nature Machine Intelligence (Article) | ~3,500-word main text, ≤6 display items + Supplementary | pre-submission draft |
| WACV 2027 paper | IEEE/CVF Winter Conf. on Applications of Computer Vision | official WACV template, 8 pages + references, anonymized review mode | ready for Round 2 (Aug 28, 2026) |
| IEEE paper | IEEE journal / conference track | IEEEtran, 7 pages | submission-ready |
| IEEE TMM paper | IEEE Trans. on Multimedia (Regular Paper) | IEEEtran journal, 7 pages (limit 10), single-blind | ready for ScholarOne submission |
| ICIP 2027 paper | IEEE Int. Conf. on Image Processing | IEEEtran conference, 4 pages (limit 5+1), double-blind | ready for Ex Ordo submission |
@article{singh2026identity,
title = {Identity Through Interruption: A Motion-Signature-First Decision
Architecture for Multi-Object Re-Identification},
author = {Singh, Sehaj Randhir},
year = {2026},
note = {Pre-submission draft; all measurements reproducible from
committed result JSONs},
url = {https://github.com/sehajr-singhs/motion-reid}
}
python train_gait.py # ST-GCN triplet model (results/gait_model.pt) python benchmark.py # synthetic benchmark (results/benchmark.json) python ablate.py # ablations (results/ablation.json) python make_demo.py # belief-stream videos (results/demos/*.mp4) python real_benchmark.py # MOT17 box-only, 7 seqs python real_hybrid_benchmark.py # MOT17 hybrid + --ablate variants python stress_lowrecall.py # detector-stress + recall sweep python real_selfcal.py # self-calibrated transfer constants python audit_channels.py # channel-calibration audit python audit_numbers.py # re-verify all 131 headline numbers python -m unittest discover -s tests # 27 unit tests (core + MOT17 loader + continuity)
Prepared with assistance from opencode / DeepSeek-V4 tooling. Thanks to Vikram Kapila and Karthik Voruganti for guidance and feedback on the manuscript.
MOT17: public-detection protocol, all seven train sequences, deterministic CPU runs. Every number traces to a committed JSON.