6 ms·
Another Hit Piece on Open-Source AI [video]
- artninja1988 3y agoIt's honestly pretty sad that at no time the authors of this paper bothered contacting laion to remove the links and work together to develop better filters. Also pretty interesting, that one of the authors calls, David Thiel himself the "Ai censorship death star". Yannic is probably right that they aren't particularly interested in bettering open source diffusion models and are more in the walled garden camp.
- Palmik 3y agoTangential, but why didn't the OpenAssistant team (lead by the author of the video) release the OpenAssistant dataset? As far as I know, the project was shut down, and only some initial highly filtered version of the data got released. This dataset could be very valuable for the community that created it.
- nmfisher 3y agoIt was fully released, no? https://huggingface.co/datasets/OpenAssistant/oasst1 https://huggingface.co/datasets/OpenAssistant/oasst1
- Palmik 3y agoThe effort started at the ~beginning of February and ended at the ~end of October [1]. The dataset you link is from April and had "unsafe" content filtered. [1] https://m.youtube.com/watch?v=gqtmUHhaplo&feature=youtu.be https://m.youtube.com/watch?v=gqtmUHhaplo&feature=youtu.be
- nmfisher 3y agoMy mistake - I guess I assumed that when the dataset was released back in April, that was the end of it, I didn't know collection was ongoing. Looks like the "final" version was released yesterday: https://huggingface.co/datasets/OpenAssistant/oasst2 https://huggingface.co/datasets/OpenAssistant/oasst2 As far as filtered/unfiltered goes, I have no idea.
- artninja1988 3y agoThey released the latest dump today :) https://huggingface.co/datasets/OpenAssistant/oasst2 https://huggingface.co/datasets/OpenAssistant/oasst2
- deleted 3y ago[deleted]
- gnabgib 3y agoThe report (Identifying and Eliminating CSAM in Generative ML Training Data and Models)[0] that this guy is very slowly sumarizing (and seems to largely agree with despite the title) was discussed 3 days ago (38 points, 30 comments)[1] [0]: https://purl.stanford.edu/kh752sm9123 https://purl.stanford.edu/kh752sm9123 [1]: https://news.ycombinator.com/item?id=38711135 https://news.ycombinator.com/item?id=38711135
- mistrial9 3y agothen why does IBM spend money producing this one? https://www.youtube.com/watch?v=y9k-U9AuDeM https://www.youtube.com/watch?v=y9k-U9AuDeM
- terminous 3y agoOpen source advocates: "With enough eyes, all bugs are shallow." These researchers: "I see your project includes a non-zero amount of CSAM." Open source advocates: "How dare you point out an issue? This is a hit piece!"
- artninja1988 3y agoWeird strawman. His critique wasn't directed at the methodology or the discovery of CSAM. He just lamented the politics and handling of it by the authors. Rather than improving the dataset, the authors published it without attempting to laion to remove the links. Instead, they chose to turn to the media, advocating for the outright banning of certain foss models, contributing to this moral panic around open source ai
- tadfisher 3y agoWhich is obviously the correct approach, because the commercial models have zero CSAM. Just don't ask for proof of this claim.