Skip to content

Comment on Unlocking a Million Times More Data for AI

Comments

What makes this vast private data uniquely valuable is its quality and real-world grounding.

This is a bold assumption. After Enron (financial transactions), Lehman Brothers (customer/population databases, financial transactions), Theranos (electronic health records), Nikola (proprietary research data), Juicero (I don't even know what this is), WeWork (umm ... everything), FTX (everything and we know they didn't mind lying to themselves) I'm pretty sure we can all say for certain that "real world grounding" isn't a guarantee with regards to anything where money or ego is involved.

Not to mention that at this point we're actively dealing with processes being run (improperly) by AI (see the lawsuits against Cigna and and United Health Care [1]), leading to self-training loops without revealing the "self" aspect of it.

[1]: https://www.afslaw.com/perspectives/health-care-counsel-blog...

(OP Here) This is a fair point. Internal datasets can be deceitful just as public ones can. That said, most propaganda lives in the public domain. :)

I'll be surprised if public data is less accurate than private data on average. I've watched many people lie to themselves or others with data within an organization because they are often incentivized to do so.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.