Our planned checking cosmetic metal surfaces with a fixed camera in a small finished-parts production cell has Fairino FR5 presenting a brushed stainless trim plate to a fixed camera. The example images contain bright streaks that shift with small presentation differences, yet those streaks are being labelled as scratches.
Before commissioning, I'd like to establish reliable references and image conditions for threshold tuning.
Check whether the suspected scratches were confirmed on physical samples under the agreed appearance procedure. Link each image to its sample identity and inspection result first.
I've found that some labels were assigned from the images themselves. The samples are still available, but those labels have no corresponding physical inspection results, making the reference circular.
@LeoBaker0449 Mark those labels uncertain and have the samples inspected properly. Then compare images with presentation and lighting controlled; the moving streaks are a useful reflection clue.
@LiamBaker0490 A scratch's visibility can change with presentation too. Don't let 'moving streak' quietly become the rule for relabelling something acceptable.
I've checked the image folders and found that exposure and presentation both changed between them. Their improved appearance can't be attributed to either change individually.
I had a separate inspection setup where the 'better lighting' folder used easier samples too. Its name implied a lighting conclusion that the comparison didn't support.
@ImranAdams0131 Which details did you retain with each image? I can record sample identity and settings, but I'd like a practical record people will actually maintain.
On my setup we kept sample ID, exposure, lighting arrangement and presentation condition. Enough to explain a comparison without reconstructing the whole session from memory.
Those records can support the vision specialist's comparison of image conditions. Holding the representative sample set constant will make an acquisition change easier to assess.
Any learned model still depends on reliable labels and representative evaluation. Changing the method doesn't resolve unconfirmed physical defects or an uncontrolled comparison of image conditions.
@JamieBaker0469 I'll keep all images of one sample together when we split tuning and evaluation sets. Our file names alone wouldn't have caught that overlap.
@LeoBaker0449 Use the held-out samples to report defect misses and false rejections independently. A combined agreement figure can conceal the type of error that matters to the process.
Show them as undecided with their count, outside confirmed-class error rates. That keeps the coverage limit visible without inventing a ground-truth label.
@LiamBaker0490 In my separate setup, reporting the undecided count made the remaining uncertainty visible. A tidier results table no longer implied an improvement in inspection capability.