Skip to content

Another Hit Piece on Open-Source AI [video]

youtube.com
30 pointsvignesh_warar11 comments
On HN

Comments

The report (Identifying and Eliminating CSAM in Generative ML Training Data and Models)[0] that this guy is very slowly sumarizing (and seems to largely agree with despite the title) was discussed 3 days ago (38 points, 30 comments)[1]

[0]: https://purl.stanford.edu/kh752sm9123 [1]: https://news.ycombinator.com/item?id=38711135

Tangential, but why didn't the OpenAssistant team (lead by the author of the video) release the OpenAssistant dataset? As far as I know, the project was shut down, and only some initial highly filtered version of the data got released. This dataset could be very valuable for the community that created it.

The effort started at the ~beginning of February and ended at the ~end of October [1]. The dataset you link is from April and had "unsafe" content filtered.

[1] https://m.youtube.com/watch?v=gqtmUHhaplo&feature=youtu.be

My mistake - I guess I assumed that when the dataset was released back in April, that was the end of it, I didn't know collection was ongoing.

Looks like the "final" version was released yesterday:

https://huggingface.co/datasets/OpenAssistant/oasst2

As far as filtered/unfiltered goes, I have no idea.

It's honestly pretty sad that at no time the authors of this paper bothered contacting laion to remove the links and work together to develop better filters. Also pretty interesting, that one of the authors calls, David Thiel himself the "Ai censorship death star". Yannic is probably right that they aren't particularly interested in bettering open source diffusion models and are more in the walled garden camp.

then why does IBM spend money producing this one?

https://www.youtube.com/watch?v=y9k-U9AuDeM

Open source advocates: "With enough eyes, all bugs are shallow."

These researchers: "I see your project includes a non-zero amount of CSAM."

Open source advocates: "How dare you point out an issue? This is a hit piece!"

Weird strawman. His critique wasn't directed at the methodology or the discovery of CSAM. He just lamented the politics and handling of it by the authors. Rather than improving the dataset, the authors published it without attempting to laion to remove the links. Instead, they chose to turn to the media, advocating for the outright banning of certain foss models, contributing to this moral panic around open source ai

Which is obviously the correct approach, because the commercial models have zero CSAM. Just don't ask for proof of this claim.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.