Spent 9 days fighting a CUDA out of memory error that turned out to be one stupid line my buddy added
I was fine tuning a small image model on my 3090 at home and kept getting hit with CUDA out of memory every 300 or so steps, no matter what batch size I tried. I dropped it to 1, cleared cache between epochs, even reinstalled the drivers twice, and lost almost 3 days thinking my card was dying. Turns out my friend had left one line in the dataloader that kept a full copy of every tensor in memory instead of releasing it, and deleting that single line made the whole run finish in 40 minutes. I get why people say always check your code first, but when you're 6 hours deep at 2am it really feels like hardware. Has anyone else blamed their GPU for something a stray reference in the training loop was doing?
Nine days? I would have thrown the whole rig out the window by day three. The part that gets me is you reinstalled the drivers twice, that's some serious commitment to blaming the hardware when the whole time it was one line of code sitting there like a landmine. A full copy of every tensor just hanging around is brutal, that's not even a subtle leak, that's your buddy basically hiding your keys and watching you tear the house apart. I've definitely stared at a dying fan curve at 2am convinced the card was done when it was just a dumb loop holding onto stuff it should have dropped. Lesson learned the hard way I guess.