A guy at a Denver AI meetup asked if I actually trust what my model outputs, and I haven't stopped thinking about it
Last week I went to a small AI meetup in Denver and a guy I'd never met asked me point blank if I trust my own model's outputs. I said yes, because I check them, and he goes "so you don't trust the model, you trust your checking." That hit different. Now I'm split on this: one side says fine-tuning on clean data with a strong eval set makes trust a solved problem, the other side says any model that can hallucinate a fake citation is a liability no matter what numbers you show. I've been running a 7B model locally on 12 gigs of VRAM, and in one test it made up a court case that looked 100% real. So which is it, is trust a data and eval problem, or is it a human process problem that no upgrade fixes? Curious what the builders here actually do in practice.