Quality-Conditional Policies for Adaptive Manufacturing
from Contact Sensor Streams
Manufacturing contact operations—insertion, polishing, deburring—are controlled today by fixed-parameter programs that ignore process variation. We show that a learned policy reading the same sensor streams a process monitor collects (force, vibration, position) can adaptively steer the operation to maintain a target specification envelope—tighter than fixed-parameter control, robust to tool wear and fixture drift. The key mechanism is a quality-conditional policy that conditions on the target specification, not the task, with a residual action space and a specification-adherence training objective that directly measures what the manufacturer cares about.
Sensor streams → Process Encoder (TCN) → FiLM conditioning on specification → Policy head → Trajectory correction. The residual formulation makes it safe and transferable.
Quality-conditional conditioning: The policy is conditioned on the target specification envelope (force range, vibration range), not the task. The same architecture works for insertion, polishing, and deburring.
Residual action space: The policy outputs corrections to a nominal trajectory, not absolute actions. This makes it safe (nominal is always a fallback), easy to learn (small corrections), and transferable across tasks.
Specification-adherence loss: A hinge loss on process signals vs. the target envelope. Zero within spec, linear outside. Directly measures what the manufacturer cares about.
Mind Robotics: Their manufacturing data flywheel needs a quality signal. Tolerance success is binary and sparse; spec adherence is continuous and dense. QC policies make the flywheel closed-loop on quality.
Gimatic: EOAT sensor data needs a controller that consumes tactile signals and outputs adaptive tool paths. QC policies are that controller.
Renishaw: Renishaw Central collects massive process datasets. QC policies turn measurement data into control actions — the missing bridge.