6 ms·
I don't think getting training data is that hard still, the biggest platforms that locked down their APIs still use them for their mobile apps and can easily be
by Shekelphile 3y ago
I don't think getting training data is that hard still, the biggest platforms that locked down their APIs still use them for their mobile apps and can easily be reverse engineered to find keys or undocumented endpoints (or in the case of reddit, an entirely different internal API with less limits and a lot more info leaks...)
- bloqs 3y agoCan you explain the reddit one?
- 4death4 3y agoAssuming the Reddit app does not use certificate pinning, you can use your computer to provide internet to your phone and then use an app like Charles Proxy to inspect requests being made from an app. Pretty easy to reverse engineer the API. If the app does use certificate pinning, then you can use an Android phone and a modified app that removes the logic that enforces certificate pinning. This is more involved but also not impossible.
- philistine 3y agoThat does not sound like the proper way to do an openAI 2.0. If Reddit ever hears that's how an AI company scraped them, they'll get sued for fun and profits.
- az226 3y agoYou can legally scrape anything that does not require a login in the US. You can also legally train an AI on it for now.
- ejstronge 3y agoAre you referring to the LinkedIn case? There has not been a decision on the legality of scraping in that matter
- 4death4 3y agoThe point is that the data is easily accessible. If you wanted to get your hands on the data while simultaneously keeping them clean, contract with a Russian contracting company to give you a data dump. You don't need to know how they got it.
- twoodfin 3y agoWell, until discovery, wherein your deliberate not knowing will be a pretty big deal.
- mr_toad 3y agoSubcontracting out your crimes isn’t going to fly in court.
- crotchfire 3y ago[flagged]
- 4death4 3y agoReally? It's done pretty regularly to limit liability.
- Nasrudith 3y agoThey make a point out of not directly asking for the crime when they do that. Just increasing pressure on subcontractors that leads to cutting corners including the law. It is harder to prove to a "should have known" standard compared to say buying stolen speakers from the back of a truck for 20% of the list price.
- 4death4 3y agoThere’s an implicit assumption in your argument that you’re going to directly ask for a crime to be committed. Why are you assuming that? You’ll go to a contractor and say “we want Reddit data.” Anyone with even mild technical competence can figure out how to get it.
- wahnfrieden 3y agoyou're aware openai trained on a boatload of pirated ebooks? they "steal" access to data because the LLM launders it on the other end
- philistine 3y agoThat is frustrating to no end. If I pirate one book I should pay a hefty fine. If a company does it it's unlocking untapped value.
- bko 3y agoWhat do you base this on? Llms know the contents of books because they are analyzed, reviewed and spoken about everywhere. Pick some obscure book that doesn't show up on any social media and ask about it's contents. GPT won't have a clue
- wahnfrieden 3y agohttps://qz.com/openai-books-piracy-microsoft-meta-google-chatgpt-bard-1850757064#:~:text=20)%2C%20the%20Atlantic%20revealed%20how,internet%2C%20to%20train%20its%20models https://qz.com/openai-books-piracy-microsoft-meta-google-cha.... What's your evidence contrary to this? Sounds like your common sense rather than inside knowledge
- bko 3y agoDid you read the article (this one misstates the case but if you look at the one linked about the lawsuit)? This is a lawsuit. Nothing has been proven. Burden of proof is on you
- Shekelphile 3y agoIt's essentially impossible to prove in court that training data was obtained or used improperly unless you go and tell on yourself. And even then it requires you to actually make someone with a lot of money mad, or to not have enough money yourself. Certainly microsoft would have already caught lots of flak for training their models on every github repo, instead they got a minor paddling from the public eye that went away after not much time had passed.
- mongol 3y agoIt is not impossible. You can call witnesses, refer to emails, source code etc.
- patcon 3y agoYeah! <3 https://github.com/mitmproxy/android-unpinner https://github.com/mitmproxy/android-unpinner
- gumballindie 3y agoWhy y’all desperate to steal data to train non intelligent software? Reddit and others should sue for license violations.
- Shekelphile 3y agoThe reddit app uses an undocumented graphql based api seperate from the publicly available rest api used by third party apps.
- thunkshift1 3y agoI think its a lot harder, while you still have lots of lawsuits coming in against AI models
- monocasa 3y agoEasier than that would just be downloading the torrent of all of Reddit through Sept 2023. https://academictorrents.com/details/89d24ff9d5fbc1efcdaf9d7689d72b7548f699fc https://academictorrents.com/details/89d24ff9d5fbc1efcdaf9d7...
- q7xvh97o2pDhNrh 3y agoThat's fascinating that the total size is so tiny — only 2.4 TB‽ I assume this must be only the text portion, and heavily compressed?
- lxgr 3y agoText really doesn't take up that much space, and in addition it compresses pretty well. The entire English language Wikipedia is only around 60GB in a format that can be readily searched and randomly accessed (ZIM), for example: https://kiwix.org/ https://kiwix.org/
- lmm 3y agoDoes Kiwix actually work? I see people hyping it here but I could never get it to actually, y'know, download the file and display the wikipedia on my phone.
- lxgr 3y agoIt works perfectly for me, both on iOS and macOS.
- pc2slow4webpack 3y agoJust downloaded. Doesn't seem to want to download Wikipedia on my phone, it says "detecting if filesystem supports 4gb files"
- vatueil 3y agoKiwix worked for me. IIRC there may be difficulties opening an archive that was downloaded outside of the mobile app, but archives downloaded in-app were fine. For the mobile app I used one of the smaller Wikipedia subsets, since I didn't want to take up too much space on my phone. The full offline Wikipedia download is saved to my laptop.
- mongol 3y agoThat would still pose a legal problem.