D
29

Finally clued in that my training data was way too clean

Been teaching a small image classifier for a side project. Spent weeks scrubbing my dataset. Got rid of every blurry photo, every weird angle. Accuracy hit 98% in testing. Then I threw some real-world photos from my phone at it. Tanked to like 60%. My friend who does this for a living laughed. Said real data is always messy. Gotta train on the garbage too. Anyone else spend way too long making things perfect before they needed to be?
2 comments

Log in to join the discussion

Log In
2 Comments
oliver_mitchell
Honestly, "have to train on the garbage too" is the realest thing I've heard in a while. My buddy works on a facial recognition thing for warehouse cameras and he told me about the time they tested their super clean model on actual footage from a loading dock. He said the system straight up flagged a forklift as a person because the training data had zero photos of actual warehouse lighting or dust on lenses. They had to go back and mix in a bunch of blurry, overexposed, and half-lit images from their own security cams before it stopped making dumb mistakes. Took them like two more weeks of just adding random noise and real world mess. Now he just rolls his eyes whenever I talk about perfectly curated datasets. The whole thing taught me that a clean dataset is kinda like a test that's too easy, it doesn't teach your model how to survive the real world.
5
evan_stone
evan_stone10d ago
His buddy's forklift example is a funny story but it doesn't prove much beyond bad data prep. Every entry level ML class warns you about lighting and lens issues, that's like basic stuff you'd catch in a weekend. Two weeks of fixing it sounds about right, not really some huge revelation about the whole field.
1