A team had trained a computer vision model to detect defects on a production line, validated against a held-out test set at 94% accuracy. When the model moved to the actual camera and lighting setup on the floor, accuracy dropped to roughly 61% — good enough to erode trust in the system, not good enough to rely on.
The gap wasn't a flaw in the model architecture. It was a mismatch between training data and deployment conditions. The training set was clean, well-lit studio images. The production environment had variable lighting across shifts, motion blur from the line's speed, and lens distortion from the mounted camera angle — none of which the model had ever seen during training.