We're planning presenting reflective components for surface inspection, with Universal Robots UR5e presenting a bright-finished enclosure faceplate to a fixed camera in a cosmetic components workshop. Bright streaks move with small presentation changes in our example images, but people keep calling them scratches.
We haven't commissioned the cell. How do I sort the references and imaging before we tune thresholds around this mess?
Were those scratches confirmed on the physical samples using your agreed appearance method? Start by linking each image to a sample and its inspection result.
@CallumChan1056 Some labels came from the pictures alone. We still have the samples, but those labels aren't tied to physical inspection results. That's irritatingly circular.
Mark those labels uncertain and have the samples inspected properly. Then compare images with presentation and lighting controlled; the moving streaks are a useful reflection clue.
@CallumChan1056 That clue needs a limit: presentation can also change how a real scratch appears. Movement in the image isn't enough to reclassify the sample as acceptable.
You're right to qualify that. The streak behaviour motivates an imaging investigation; it doesn't replace physical classification under the agreed inspection procedure.
Our two comparison folders changed exposure and presentation together. I can't tell which change made the images look better. That comparison needs a caveat too.
I had a separate inspection setup where the 'better lighting' folder used easier samples too. Its name implied a lighting conclusion that the comparison didn't support.
@SarahBell0676 On my setup we kept sample ID, exposure, lighting arrangement and presentation condition. Enough to explain a comparison without reconstructing the whole session from memory.
Use those details to ask your vision specialist for controlled comparisons on representative surfaces. Keep the sample set fixed when assessing an acquisition change.
It still needs trustworthy labels and representative evaluation. A different model doesn't establish whether your reference scratches are real or your acquisition comparison is fair.
@CallumChan1056 Keep physical sample identity in mind when forming evaluation sets. Different images of a sample used for tuning don't provide an independent sample-level evaluation.
Use the held-out samples to report defect misses and false rejections independently. A combined agreement figure can conceal the type of error that matters to the process.
Show them as undecided with their count, outside confirmed-class error rates. That keeps the coverage limit visible without inventing a ground-truth label.