A guy at the Seattle AI meetup told me my benchmark charts were garbage, so I changed how I read every new model release | AI newest information - Discusd
A guy at the Seattle AI meetup told me my benchmark charts were garbage, so I changed how I read every new model release
Last month I showed up at the Seattle AI meetup with slides full of leaderboard scores for three new models, and a guy in the back named Dave cut me off and said "you're just reading ads." He explained that a lot of these scores come from the labs' own tests, so I was basically repeating press releases. I went home and dug into the eval docs for the last two big releases, and he was right, the fine print on the test sets was buried and vague. Now I check who ran the evals and on what data before I get excited about any new model. Anyone else here do this, or am I overthinking it?