14
The stat on AI training data licensing that made me pause mid-scroll
I was reading a report from last month about how something like 70% of commercial AI models are trained on data that technically violates the original terms of use. I found that number in a research paper someone linked on a discord server, not even a big publication. It hit me because I've been arguing with my team about using scraped data for our internal chatbot, and I assumed everyone just accepted the risk. Has anyone else actually checked their own training sources after seeing stats like that, or is it just a theoretical problem until a lawsuit shows up?
2 comments
Log in to join the discussion
Log In2 Comments
dianas5020d ago
Does your team actually talk about what happens when that lawsuit shows up, or is it one of those things everyone pretends doesn't exist until legal sends a panicked email? I read that same stat a few weeks back and it honestly made me dig through our own training logs, and yeah, we found stuff that was iffy at best. It's scary how easy it is to just trust the data dump until someone else gets burned first.
8
kai91419d ago
Whoa wait, 70%? That number feels way too high to just be floating around a Discord server, ngl. @dianas50, you actually digging through your logs is exactly what more teams should do, but I bet most of them freeze up like a deer in headlights when legal finally speaks up. Tbh, I think we're all one big lawsuit away from this becoming a real problem instead of a scary statistic.
4