3 ms·
The [lack of] integrity of OpenAI (and any other frontier lab) should already be pretty solidified. Among other horrible things, these companies stole millions
by ryan_n 18d ago
The [lack of] integrity of OpenAI (and any other frontier lab) should already be pretty solidified. Among other horrible things, these companies stole millions of IPs and no one seems to care anymore. Regardless of what you think of the product they are making and the success of ai/its impact on humanity, these companies objectively do not have much integrity.
- simonw 18d agoHow do you feel about the integrity of the machine learning researchers over the past twenty years who trained models on scraped internet data that weren't particularly powerful and didn't attract any attention?
- lukewarm707 18d agoi think that this case, if they did train on buckmaster and alpöge, amounts to an attempt to steal the millenium prize, bypassing all attribution. legally speaking, the default privacy notice gives them an irrevocable license to your content. they may read and use the prompts for research. so it is very possible they simply stole the navier-stokes solution. that is the same principle as any other prompt but this would be a concrete example. there would be some difference between simply giving the model some prompts to read, which they are entitled to do on the default policy, and putting it into aggregate training data.
- ryan_n 18d agoIf they scraped internet data in the same way as current day frontier labs do, then I feel the same exact way about them. Why would I feel any different if that is the case?
- simonw 18d agoMy point is that researchers and academics really have been doing this for decades - it's the reason projects like Common Crawl and LAION exist. I think it's notable that nobody was calling out those researchers for their lack of integrity, because the systems they were building did not seem like a threat to anyone. OpenAI etc get accused of a lack of integrity on this precisely because the systems they are building work, and are profitable. My personal opinion here is that integrity is more about what you build with the data. I think saying "scraping means you lack integrity" is a simplification.
- smcg 18d agoyou massively collapsed what AI companies have been doing by comparing it to old internet-scraping. Facebook flat-out admitted that they scanned copyrighted books for their AI. The image generators most definitely trained on copyrighted images.
- ryan_n 18d agoLAION and Common Crawl both scraped copyrighted images. From what I can tell (I'm not an expert in this domain at all), the main difference between those two and frontier labs is in how they stored and used the data. CC and LAION seem to be actually open (unlike "Open"AI) and are more centered around publicly sharing the data they scrape to support research and innovation. OpenAI et al also stole everything from everyone. But then they raised billions of dollars from that data and sell back their LLM to people (again, among other things). They are also very much NOT open in any way, aside from sharing their benchmarks of new models.
- needfish 17d agoPersonal two cents, I have friends whose music work posted on YouTube were scraped to be in LAION-DISCO-12M, so yeah not very open.
- ryan_n 17d agoWhat I meant more is that the dataset they scrape is openly available for download by anyone, unlike any of the frontier labs. Not that they don’t scrape copyrighted content. Still sketch, but at least they don’t call themselves “OpenLAION”. Also my understanding was they’re not storing the actual music, but the metadata and a link to the YouTube video.
- ccgreg 17d agoCommon Crawl is text-only.
- simonw 18d ago
- fn-mote 17d ago> who trained models on scraped internet data The strongest complaint is that they trained on a huge corpus of pirated copyrighted works. It’s a large step above “scraping” and well into the “everyone acknowledges this is illegal” territory.
- deleted 17d ago[deleted]
- deleted 18d ago[deleted]