5 ms·
"Stealing" something you already paid for (tokens), but that you can't have access to(!). And trained on the sum of human knowledge. Training on other model ou
by Aissen 2mo ago
"Stealing" something you already paid for (tokens), but that you can't have access to(!). And trained on the sum of human knowledge.
Training on other model outputs ought to be business as usual, stop using morally charged terms made up by future monopolists: https://thomasdullien.github.io/posts/2026-06-15-rl-economics-morally-charged-terms-and-distillation/ https://thomasdullien.github.io/posts/2026-06-15-rl-economic...
- deleted 2mo ago[deleted]
- __MatrixMan__ 2mo agoLiberating!
- paxys 2mo agoThe only person calling it stealing is the author of this article, so this is a pointless discussion. The majority of this thread is just arguing with themselves.
- dymk 2mo agoAnthropic and OpenAI made a big deal about how it's stealing.
- amazingman 2mo agoThey also have made a big deal about how what they did to build their models is not stealing. And we all know that's bullshit.
- blackqueeriroh 2mo agoNo, we actually don’t all know that.
- amazingman 2mo agoI'm pretty sure essentially all HN participants understand that the frontier labs indiscriminately sucked up every bit of human output they could, IP and ethical concerns be damned. Some of that cohort may indeed be okay with it, but that doesn't change the facts.
- ajam1507 2mo agoKnowing that they trained on that data doesn't mean that you've demonstrated that they "stole" it. Certainly the courts haven't decided that in every case.
- khanan 2mo agoYou mean you didn't know that all frontier models stole all of our knowledge and are now charging for it? It's abysmal and disgusting and we should pitchfork them all! :D
- brianxq3 2mo agoThey are also encrypting it so they must see some reason to do this. I suspect they think it is proprietary or otherwise a way that people can “steal” their implementations.
- fragmede 2mo agoThe reason for this is the LLM says some truly unhinged shit while in the thinking stage of the process, and Twitter would trip over itself to make fun of what it says.
- encomiast 2mo agoStealing may be the wrong word, but I actually think this is important. I don't think the providers have been up-front about how we should be handling these thought signatures. A large system with a lot of users may be capturing these and even caching them to send them back with future requests. If data can be pulled out of these, then they need to be treated more like cookies than opaque, encrypted nonces.
- throw1234567891 2mo agoNo, you paid for the end result. The thought process is a step in between, a function. Think about it, who should get charged if the answer you received comes from a cache? Thinking tokens are the complexity-of-the-problem cost. I mean, you may not agree but both are valid points of view.
- anigbrowl 2mo agoNo I didn't. I buy my tokens from a provider that exposes the model reasoning so I can understand what it's doing and work with it, or interrupt if I see things going in the wrong direction.
- throw1234567891 2mo agoI agree with you in principle. I'm just pointing out that the latter is in some way another valid point of view.
- HDBaseT 2mo agoThinking tokens aren't free though. This is not a valid point of view. If I was being charged for the raw, output/input token count, excluding thinking/reasoning token costs, then sure. But at least via the API, you pay for tokens you cannot see.
- ardel95 2mo agoIs it really that unusual? When you attach an image or a video, it gets converted to tokens you don’t see, at a rate that is proprietary to the model. You pay for those tokens, but don’t see them. Even how text is converted to tokens is a property of the dictionary, which is opaque for proprietary models. There are features of input and output that are opaque to you, but that you pay for. Part of how model providers chose to run their service.
- throw1234567891 2mo agoWell, I don’t pay for them. Maybe you do. But again, they’re a byproduct, an intermediary. They contribute to your result but aren’t the final result. I’m not defending their position, just showing you an alternative universe.
- nonethewiser 2mo ago> stop using morally charged terms made up by future monopolists Lets not gloss over this claim. Being: “Stealing is a morally charged term made up by future monopolists.” I strongly disagree. Stealing is not a made up term and property rights are foundational for any society. Your take is at least sensationalist if not malicious.
- theyliesoeasily 2mo agoThey definitely stole the data to make the models, but they do not say that they stole the data to make the models, but they do say that others using their outputs for unauthorized purposes is stealing. Do you see the point?
- super256 2mo agoCrawling the internet and dumping it to disk is not "stealing".
- theyliesoeasily 2mo agoThen distilling models and deobfuscating reasoning traces isn't. Mass downloading copyrighted works is. Which they did. Aaron got threatened with 20 years, they got pentagon contracts.
- michaelmrose 2mo agoIt's not stealing but arguing that it's not infringement because its on the internet is pretty obviously nonsense.
- UpsideDownRide 2mo agoIs everything licensed in the same way? Are there any copyrighted works available to be had through crawling?
- modriano 2mo agoIf they were only copying, for example, New York Times articles and many publishers to a disk, I don't think NYT and the publishers would have sued OpenAI. But OpenAI isn't just copying things to disk. NYT reported ChatGPT (before Dec 2023, [0]) was returning near verbatim sections of NYT articles. Is this stealing? Is it depriving NYT or publishers/writers from money via lost sales/subs? I don't know, but it certainly could be. [0] https://www.nytimes.com/2023/12/27/business/media/new-york-times-open-ai-microsoft-lawsuit.html?smid=nytcore-android-share https://www.nytimes.com/2023/12/27/business/media/new-york-t...
- ashdksnndck 2mo agoSuppose you hired a consulting firm to write a report, and they delivered the report but not the internal conversations they had when developing it. You exploit a vulnerability in their phone system to get those conversations. You can argue over semantics of whether “theft” is what you did, maybe the right word is “espionage” or “spying”, but that either way we probably agree you are guilty of something? Paying for the final product didn’t entitle you to see how it was made, unless that was part of the agreement. Anyway, you can distinguish this from the debate over copyright.
- nathanwh 2mo agoI think a more fair comparison would be that you hired a consulting firm to create a report and give you a summary of it, but you’re charged for the report itself separately from the summary, and you are not allowed to access the unsummarized report.
- selestify 2mo agoHow is that a more fair comparison? The consulting firm in this case never promised you the interim reports, only the summaries of the reports. They also promised you the final output that the reports led to. You decided that report summaries + final output was worth paying for. You got exactly what you were promised.
- iot_devs 2mo agoI personally read the thinking traces to know if the model is on the right direction
- selestify 2mo agoI'm not saying they're not useful, of course they are. I am disputing that they are part of the agreed bargain between you and the proprietary LLM providers. They explicitly do not promise reasoning traces. You (general you) agree to those terms and pay for that bargain anyways.