We're planning checking cosmetic metal surfaces with a fixed camera, with Universal Robots UR5e presenting a brushed stainless trim plate to a fixed camera in a robot-presented camera station. Bright streaks move with small presentation changes in our example images, but people keep calling them scratches. We haven't commissioned the cell.
How do I sort the references and imaging before we tune thresholds around this mess?
Check whether the suspected scratches were confirmed on physical samples under the agreed appearance procedure. Link each image to its sample identity and inspection result first.
I've found that some labels were assigned from the images themselves. The samples are still available, but those labels have no corresponding physical inspection results, making the reference circular.
@IsabelBell0648 Mark those labels uncertain and have the samples inspected properly. Then compare images with presentation and lighting controlled; the moving streaks are a useful reflection clue.
@BrunoArcher0365 A scratch's visibility can change with presentation too. Don't let 'moving streak' quietly become the rule for relabelling something acceptable.
@NoahChen1151 Keep them explicitly undecided, with the sample link. Don't use them as confirmed examples of either class until the physical review settles them.
I've checked the image folders and found that exposure and presentation both changed between them. Their improved appearance can't be attributed to either change individually.
Which details did you retain with each image? I can record sample identity and settings, but I'd like a practical record people will actually maintain.
For my separate setup, sample identity, exposure, lighting arrangement and presentation condition were the useful essentials. They let us understand comparisons without relying on recollection
@OwenBennett0747 Those records can support the vision specialist's comparison of image conditions. Holding the representative sample set constant will make an acquisition change easier to assess.
It still needs trustworthy labels and representative evaluation. A different model doesn't establish whether your reference scratches are real or your acquisition comparison is fair.
I'll group images by physical sample before dividing the tuning and evaluation sets. The current filenames wouldn't reliably expose the same sample appearing in both.
Show them as undecided with their count, outside confirmed-class error rates. That keeps the coverage limit visible without inventing a ground-truth label.
I can explain the reference uncertainty better, although the example images still don't support a performance judgement. Physical classification and image-condition comparisons remain outstanding.