We're planning presenting reflective components for surface inspection, with Universal Robots UR5e presenting a polished metal cover to a fixed camera in a robot-presented camera station. Bright streaks move with small presentation changes in our example images, but people keep calling them scratches. We haven't commissioned the cell.
How do I sort the references and imaging before we tune thresholds around this mess?
Check whether the suspected scratches were confirmed on physical samples under the agreed appearance procedure. Link each image to its sample identity and inspection result first.
Some labels came from the pictures alone. We still have the samples, but those labels aren't tied to physical inspection results. That's irritatingly circular.
Mark those labels uncertain and have the samples inspected properly. Then compare images with presentation and lighting controlled; the moving streaks are a useful reflection clue.
That clue needs a limit: presentation can also change how a real scratch appears. Movement in the image isn't enough to reclassify the sample as acceptable.
I've checked the image folders and found that exposure and presentation both changed between them. Their improved appearance can't be attributed to either change individually.
@RachelChan1131 On my inspection setup, a folder called 'better lighting' also contained easier samples. The folder name was doing a lot of unearned work.
On my setup we kept sample ID, exposure, lighting arrangement and presentation condition. Enough to explain a comparison without reconstructing the whole session from memory.
@FionaChen1195 Use those details to ask your vision specialist for controlled comparisons on representative surfaces. Keep the sample set fixed when assessing an acquisition change.
It still needs trustworthy labels and representative evaluation. A different model doesn't establish whether your reference scratches are real or your acquisition comparison is fair.
@HanaAllen0274 Keep physical sample identity in mind when forming evaluation sets. Different images of a sample used for tuning don't provide an independent sample-level evaluation.
@LouisChan1094 I'll keep all images of one sample together when we split tuning and evaluation sets. Our file names alone wouldn't have caught that overlap.
Show them as undecided with their count, outside confirmed-class error rates. That keeps the coverage limit visible without inventing a ground-truth label.