Showerthought: I fed my logging tool raw chat logs vs tagged examples and the difference was scary
Last Tuesday I finally sat down and tested something I should have tried months ago. I took 500 lines of messy support chat from our shop's Slack and ran it through a small fine tuned model, then ran the same 500 through a plain prompt with no examples. The difference in accuracy was wild. The no example version got maybe 40 percent of the intents right and kept hallucinating order numbers that were not even in the text. The version with just 12 labeled examples, 12, got 88 percent right and never made up a single number. I was so sure more raw data would win because everyone says data is king. Nope. Clear examples beat raw volume every time in my test. If you are throwing unlabeled logs at a model and hoping it figures out your weird internal slang, stop and spend an afternoon tagging a tiny set instead. What example counts are you all using for small custom tasks?
Its like how a recipe with a few good photos works better than a cookbook with no pictures, your brain just needs something to copy. I usually tag around 10 to 20 examples and thats plenty for most small jobs.