The idea, made measurable
A servo joint accelerating a payload dissipates ohmic heat into a winding whose rising temperature derates the torque available to the next command, on a thermal time constant three decades slower than the control loop that issued it. Watch any one channel and you see a line wandering. Watch two against each other and you see a shape, and the shape is where the coupling lives, because a lag between torque and winding temperature is invisible in either trace alone and shows up immediately as an oriented loop in the plane they span.
Each channel pair gets nine atoms: monotone association, functional strength, a nonlinearity gap, functional asymmetry, oriented hysteresis, jump share, timescale ratio, support occupancy, and a scaling exponent. The whole array over all pairs is one unit's index of operations. Nine quantiles of each atom across pairs gives a signature of fixed length whatever the channel count, which is the index of operators, the object that lets a 12-channel motor and a 24-channel turbofan be compared at all.
Eight of the nine are computed on within-unit ranks, so they are exactly invariant to any strictly increasing recalibration of each channel separately. That is the property a fleet needs when sensors get replaced.
levy = +0.47, which is the thermal lag made
visible.What held, and what did not
Operating episodes are identifiable, distinct machines only on one axis
Matched from a different stretch of its own operating history, a motor session is recovered from 39 distractors at 32.5% top-1 against 2.5% chance. The scope matters and an earlier draft of this page got it wrong: the motor benchmark is one 52 kW machine recorded in 40 sessions, so this distinguishes operating episodes, not siblings. On simulated fleets of genuinely distinct machines the answer is now controlled rather than a single number: it depends on which axis carries the identity, and the controlled experiment below resolves the earlier wash.
Sibling machines: the atlas is a fingerprint of coupling, not of load
Machine identity has two axes. The level axis is absolute magnitude — payload, how hot a machine simply runs. The relational axis is how one channel couples to another — bearing friction, thermal resistances, servo gain. Three fleets of 48 physically distinct MuJoCo robots each, identical except which axis carries the identity, give a clean inversion on every one of three platforms (UR5e, Panda, iiwa14). Carried by payload alone, the atlas retrieves held-out siblings at 6.2% against 2.1% chance (not significant), while magnitude features are significant. Carried by the six relational parameters alone, the roles invert: the atlas retrieves at 10.4% on the UR5e and 20.8% on the Panda (p < 0.001), and the per-channel calibration block falls to chance. When both axes vary, the atlas reaches 39.6% on the Panda — ahead of the per-channel marginals at 29.2%.
Under simulated re-instrumentation the atlas is numerically unmoved on all nine platform–fleet combinations, to machine precision, while every magnitude-based feature collapses. The deep baseline makes this quantitative: an autoencoder trained on raw telemetry windows outreads the atlas on clean data — 58.3% against 39.6% when both axes vary on the Panda — and falls to 2.1% (chance) after a single per-channel recalibration, losing ~96% of its accuracy. The atlas is exactly where it was. A magnitude fingerprint, learned or hand-built, wins only while the instrumentation never changes; the moment a sensor is replaced, the level axis it was reading is relabelled and the advantage inverts.
Correlation is not what carries identity
Reducing the atlas to its first atom, a Spearman correlation matrix, drops identification to 7.5%. The atom doing the work is oriented hysteresis at 32.5%, and correlation cannot express it: correlation is symmetric, the Lévy area is antisymmetric, so no correlation-based descriptor encodes which of two coupled channels leads.
Exactly one atom can see the arrow of time
Reverse a record end to end and a lag becomes a lead. Any statistic built from the joint distribution of simultaneous values is unchanged, because reversal only permutes the samples: that covers correlation, mutual information, distance correlation and coherence magnitude. Measured on 40 real motor sessions, eight atoms reproduce themselves to between 10−15 and 10−11, while oriented hysteresis satisfies levy(reversed) = −levy to 4.7×10−14. The vocabulary splits into an even sector and a one-dimensional odd sector, the whole standard toolkit lives in the even sector, and lag with a direction is precisely the odd sector's content.
Invariance to recalibration is exact
Under an independent monotone warp on every channel the atlas is unmoved at 32.5%, while per-channel marginal statistics fall from 25.0% to 5.0%.
Classes separate across unlike systems
99.4% ± 0.8% over seven real industrial systems — a five-year gas turbine log, a hydraulic rig, a motor bench, an air compressor, an electricity transformer, a wind farm and a pump-rig anomaly bench — against a 35.6% majority (309 records), where per-channel marginals reach 96.1%. The wider stress test adds a turbofan and traded assets: 93.6% ± 1.4% over 6 classes against a 28.6% majority.
The atlas encodes the machine, not the workload
On simulated fleets of 80 machines whose true physics is known, per channel statistics decode the episode's ambient temperature at R² = 0.99, 0.95 and 0.72 and its duty cycle at 0.82, 0.89 and 0.84. The atlas decodes ambient temperature at −0.05, −0.07 and −0.11, which is no better than predicting the mean. A fingerprint that quietly encodes the task will pass every identification benchmark and fail in the field, and this is the control that catches it.
Unit variation is low-dimensional
Within a class, 80% of the variance across units sits in 43 to 45 directions out of 12,555 to 17,010. That is the quantitative form of the claim that a class prototype fine-tunes to one machine by moving a small number of knobs.
The in-distribution gain over a simple baseline is small
Per-channel means, standard deviations, skews and kurtoses reach 25.0% against the atlas at 32.5%. On a stable, co-calibrated fleet the extra machinery buys little. The separation appears only when instrumentation changes.
Short records defeat it
Atoms need roughly 103 to 104 samples per unit before they stop being noise. On C-MAPSS, whose units are about 200 cycles, identification reaches 1.1% against 0.14% chance and the atlas scores below a plain correlation matrix at class assignment.
Knowing the machine did not help predict the machine
Pretraining a forward model on a fleet and adapting to an unseen unit, conditioning on the unit's atlas gave 0.168 test error against 0.129 for no conditioning at all. The diagnostic is the oracle arm: handed the machine's true payload, thermal resistances and bearing friction, it reached 0.160, also worse than nothing. So the identity is not recovered badly, it is redundant, because a model that already sees a causal history of the state has been told what machine it is on. Anyone conditioning a predictor on a learned machine embedding should run the oracle arm first.
Invariance costs identifiability
A machine that simply runs ten kelvin hotter is identified by that fact, and an invariant is blind to it by construction. Cells cut on absolute load give the best single atom in the whole study (42.5%) and lose three quarters of it under re-instrumentation; cells cut in rank space give the best full atlas and hold. Neither dominates.
Identification, in numbers
Atlases are built from the first half of each session and matched against atlases from the second half, so the target must be recognised from telemetry it did not produce.
| Feature set | Top-1 | Pct. rank | After recalibration |
|---|---|---|---|
| Atlas, invariant 8 atoms | 32.5% | 90.0% | 32.5% |
| Full atlas, 9 atoms | 27.5% | 89.2% | — |
| Per-channel marginals | 25.0% | 82.6% | 5.0% |
| Correlation matrix only | 7.5% | 71.4% | — |
| Chance | 2.5% | 50.0% | 2.5% |
Including the one scale-dependent atom actively hurts: the invariant eight reach 32.5% where all nine reach 27.5%. The scaling exponent is kept for reading physics off a shape, not for identification.
What this is for
The useful reading is a statement about where relational geometry earns its cost. On a fleet that is stable and co-calibrated, per-channel statistics already carry most of what identifies a machine. The moment instrumentation changes, and it does change, because sensors are replaced and drives are re-tuned and thermocouples are re-zeroed, everything built on magnitudes is invalidated at once while the atlas is untouched.
The second reading concerns what kind of quantity carries machine identity. Symmetric descriptors of a pair, which is what nearly the whole relational feature literature uses, cannot express which of two coupled channels leads. In a system whose defining behaviour is a lag with a direction, that is the information worth having, and it is free once the pair is being looked at anyway.
The third reading is negative and deserves equal weight. A unit code that demonstrably carries unit identity did not improve short-horizon forward prediction, and neither did the true physical parameters. That leaves a representation whose value is in identification, auditing and transfer rather than in accuracy. Knowing which machine a record came from, knowing the answer survives a sensor swap, and knowing the answer is not secretly the duty cycle, are worth having on their own.