Our planned inspecting reflective surfaces for scratches in a robot-presented camera station has Fairino FR5 presenting a polished metal cover to a fixed camera.
The example images contain bright streaks that shift with small presentation differences, yet those streaks are being labeled as scratches.
Before commissioning, I'd like to establish reliable references and image conditions for threshold tuning.
Were those scratches confirmed on the physical samples using your agreed appearance method? Start by linking each image to a sample and its inspection result.
@NaomiAllen0332 I've found that some labels were assigned from the images themselves. The samples are still available, but those labels have no corresponding physical inspection results, making the reference circular.
@SaraAllen0287 Keep the image-only labels uncertain pending physical inspection. A controlled comparison of lighting and presentation can then investigate whether the moving streaks indicate reflection effects.
That clue needs a limit: presentation can also change how a real scratch appears. Movement in the image isn't enough to reclassify the sample as acceptable.
I've checked the image folders and found that exposure and presentation both changed between them. Their improved appearance can't be attributed to either change individually.
@TobyAbbott0057 Which details did you retain with each image? I can record sample identity and settings, but I'd like a practical record people will actually maintain.
On my setup we kept sample ID, exposure, lighting arrangement and presentation condition. Enough to explain a comparison without reconstructing the whole session from memory.
@TobyAbbott0057 Use those details to ask your vision specialist for controlled comparisons on representative surfaces. Keep the sample set fixed when assessing an acquisition change.
Any learned model still depends on reliable labels and representative evaluation. Changing the method doesn't resolve unconfirmed physical defects or an uncontrolled comparison of image conditions.
Keep physical sample identity in mind when forming evaluation sets. Different images of a sample used for tuning don't provide an independent sample-level evaluation.
@CallumChen1143 I'll group images by physical sample before dividing the tuning and evaluation sets. The current filenames wouldn't reliably expose the same sample appearing in both.
Use the held-out samples to report defect misses and false rejections independently. A combined agreement figure can conceal the type of error that matters to the process.
@NaomiAllen0332 How should the evaluation report handle samples whose physical status is undecided? I want them visible without assigning an unsupported error label.
Show them as undecided with their count, outside confirmed-class error rates. That keeps the coverage limit visible without inventing a ground-truth label.
The reference problem is clearer, but I still can't judge inspection performance from our example images. The physical labels and image comparisons need more work.