Our planned inspecting reflective surfaces for scratches in a robot-presented camera station has Fairino FR5 presenting a polished metal cover to a fixed camera. The example images contain bright streaks that shift with small presentation differences, yet those streaks are being labelled as scratches. Before commissioning, I'd like to establish reliable references and image conditions for threshold tuning.
Check whether the suspected scratches were confirmed on physical samples under the agreed appearance procedure. Link each image to its sample identity and inspection result first.
Some labels came from the pictures alone. We still have the samples, but those labels aren't tied to physical inspection results. That's irritatingly circular.
Keep the image-only labels uncertain pending physical inspection. A controlled comparison of lighting and presentation can then investigate whether the moving streaks indicate reflection effects.
You're right to qualify that. The streak behaviour motivates an imaging investigation; it doesn't replace physical classification under the agreed inspection procedure.
Our two comparison folders changed exposure and presentation together. I can't tell which change made the images look better. That comparison needs a caveat too.
On my setup we kept sample ID, exposure, lighting arrangement and presentation condition. Enough to explain a comparison without reconstructing the whole session from memory
@YasminAdams0093 Those records can support the vision specialist's comparison of image conditions. Holding the representative sample set constant will make an acquisition change easier to assess.
Any learned model still depends on reliable labels and representative evaluation. Changing the method doesn't resolve unconfirmed physical defects or an uncontrolled comparison of image conditions.
Use the held-out samples to report defect misses and false rejections independently. A combined agreement figure can conceal the type of error that matters to the process.
Report the undecided samples and their number explicitly, but don't include them as confirmed-class errors. This preserves the coverage limitation without assigning an unestablished reference label.
The reference problem is clearer, but I still can't judge inspection performance from our example images. The physical labels and image comparisons need more work.