6 ms·
I agree. The frontier models are based on training data from tons of copyrighted work. Some of that work was obtained illegally, even. They could not exist with
by kelnos 21d ago
I agree. The frontier models are based on training data from tons of copyrighted work. Some of that work was obtained illegally, even. They could not exist without strip-mining the commons. The labs have no moral or ethical ownership to the end result, and others should feel free to treat any company-imposed restrictions on their use as invalid.
I don't expect Tan's position to be based on any kind of real moral high ground, but his conclusion is correct.
I love the "illicit distillation attacks" framing from the incumbents. There's nothing illicit. There's no attack. You just don't like it because it threatens your market position and business model.
- knollimar 21d agoI'm sure they put some BS in their TOS
- mobelkh 21d agowhy can't I use the tokens i paid for anyway?
- giancarlostoro 21d agoAbolish copyright and make it less ridiculous. Sampling music was never a thing that required royalties until the 1990s when I guess someone got angry that rappers were making money off their sampled music. Its insane to me. Make it illegal to transfer ownership of copyrighted work too, only the spouse or one single inheritor who isnt a company can have the rights transferred, after both die, the work enters public domain. LLMs should just pay a flat fee to use a specific book and thats it. Fees should be reasonable (not a million dollars per book), so long as the model doesnt spit out the entire book.
- TFNA 21d agoOne of the most infamous legal challenges to sampled music was MARRS "Pump Up the Volume" in the 1980s, and that was preceded by other famous cases. Not sure why you think that started in the 1990s.
- jrajav 21d agoThis is nitpicky. The MARRS case was 1987, and Biz Markie and Vanilla Ice are way higher on the list in terms of actually getting attention on the issue and influencing culture.
- TFNA 20d agoThe USA is not the whole world. I heard a lot about the MARRS case as a teenager, and I wasn't even in the UK where it happened.
- derefr 21d ago> Make it illegal to transfer ownership of copyrighted work too, only the spouse or one single inheritor who isnt a company can have the rights transferred, after both die, the work enters public domain. By your phrasing, it sounds like you still intend the possibility of companies owning copyrights; but how does that happen (other than copyrights already owned by companies grandfathered in)? Copyright always starts off in the hands of individual human beings; it only ends up in the hands of companies when those human beings transfer ownership to a company. That ownership transfer can be automatic as a term of a contract, e.g. as part of a work-for-hire agreement. But no contract can cause the copyright to come into existence already held by the company instead of the individual. So if you abolish ownership transfer, you effectively make work-for-hire IP assignment invalid. What replaces it? And, if "nothing"... then how do people pool the IP rights of their own small contributions to a large-scale work, into an IP pool that can be legally defended by a coherent legal entity, so that the large-scale work itself can have market value (i.e. so that sales of polished commercial bootlegs don't drive sales of the "authentic" work to zero)? Keep in mind that, no matter how much we might want "mass distributed" media to have more-reasonable IP terms, the ability to sue for infringement is still critical to the existence of some forms of media. Especially "location-based" media, with no equivalent licensed broadcast right: movies still in theatre; concerts; live performances of plays and musicals; etc. If there's no legal team that can sue a movie theatre that shows an unlicensed copy of a given movie, then no movie theatre will ever bother with licensing movies again; "box office" goes to zero (from the movie company's perspective); and the incentive to create movies in the first place declines massively. (You can see what this alternate world looks like from the few cases where movies screwed up the steps required to assert copyright, back before copyright was automatic. Night of the Living Dead (1968) is a good example: theatres — even upstanding large-chain theatres! — did indeed leap at the opportunity to show the movie unlicensed, and so Romero et al made effectively zero revenue off the work.) I'm not saying this is an impossible problem. There are ways to accomplish this besides the way it's done now. (For example, individual-contributor IP could be retained by the original owners, but cross-licensed between individuals through a collaboration structure to form a coherent defensible IP pool, in exactly the same way that IP for e.g. video codecs is cross-licensed between corporations to form a coherent defensible IP pool today.) I'm just pointing out that the problem does need to be solved.
- stymaar 21d agoThis. Distillation “attacks” are a made up concept. It's as if I claimed that Anthropic made a “training attack” when training on my internet writing.
- overfeed 21d agoAnthropic carried out a multitude of "copyright attacks" on open source repositories, and the broader internet.
- torginus 21d agoWith the recent Navier-Stokes controversy, I think there's a credible suspicion that all your IP you run through these models will end up in these companies' possession. OpenAI themselves has admitted a weak version of this (that prompts might inadvertedly end up improving the model). We don't know the extent of this. Obviously it's not possible to run a company whose value is predicated on its IP that uploads said IP to a third party which might get access to it. This could mean every potential serious customer would have no option but to seek alternatives to these online services.
- ronsor 21d agoAlmost every serious customer is already using ZDR where nothing is retained at all, instead of "anonymized" data.
- applfanboysbgon 21d agoZDR is based on the exact same pinky-promise as training opt-outs. There is no technical barrier to OpenAI, or whoever is running your compute, retaining your prompt after they run inference on their servers. If you don't control the hardware the model is being inferenced on, you don't control your data.
- pennomi 21d agoWhere nothing is retained at all, allegedly.
- steveBK123 21d agoThey already trained on pirated content, what makes you think they are going to honor ZDR?
- GrinningFool 21d agoContractual obligations carry teeth. Scraping the internet is relatively risk-free.
- 21d ago
- Barbing 21d agoAll correct, just help me get over the idea of an open-weight Mythos where one or a dozen of us eight billion does something stupid on the bioweapon front. Smart people who’ve exhausted possibilities for what they can do with books and web search and today’s Kimi/GLM. Figure we’ll have to reckon with this next year in any case, guess we’ll see.
- ronsor 21d ago"Bioweapon" information is not useful without a lab for synthesis. Someone with that lab could almost certainly figure out how do something stupid or destructive on their own, or bypass model safeguards somehow.
- deleted 21d ago[deleted]
- a34729t 21d agoYou dont need an LLM to figure out to make anthrax. Anybody who can figure out how to make a home lab can make all sorts of dangerous stuff pretty easily. Same with college grad from a respectable chemistry program. This all FUD.
- edot 21d agoThis is our generation's "Saddam has WMDs". It's something the big labs thought up when they were trying to figure out how to make their product sound scary enough to deserve regulation. Literally no one is doing this or even trying, anyone who would want to do it would have already done it. Not worried about it.
- godwinson__4-8 21d agoIf the leading private labs attempt to use the government to pull up the ladder under the pretense of "safety" then the response of the people should be to take such questions out of private hands and nationalize the leading labs. Or they could abide by the precedents they set and learn to compete. They shouldn't be allowed to have it both ways.
- chadgpt3 21d agoHow can "the people" nationalize a lab? I'm people, how can I do it?
- georgemcbay 21d ago> I'm people, how can I do it? Vote (well-informed of the candidate's policies) in every election you can, even the local ones that seem of little consequence. Convince others to vote. Make demands of your elected representatives. You can mail them, call them, etc. The government is the people. The Reagan-era and beyond successful convincing of people that the government is an unchangeable black box made up of shady actors out to destroy everything (see: Republicans still going on about the 'deep state' when they run literally everything) is a big part of how we got to this place. It was a self-fulfilling lie, now coming true as the people who sold the lie start grasping for unending power. But we still have the ability to vote our way out of it. If we continue to fail to do so, then at an evolutionary level we have to consider that we collectively deserve all the bad that comes from it.
- tehjoker 21d agoVoting only changes things when it doesn’t threaten the interests of elites or there is a sufficient counterweight in terms of a competitor nation or a radical labor movement or an uncontrolled armed insurgency. See salvador allende in chile and mitternand in france for examples of voting without sufficient leverage. your regan example occurred during a successful counterattack by capital that started under carter and crushed the labor movement.
- Aurornis 21d agoThere is nothing illegal about training on traces from frontier models. However the frontier labs don’t have to serve customers who are farming the service for distillation purposes. That’s their choice and they’re free to make it if they detect distillation happening.
- hlynurd 21d agoThat's fine, they just gotta tone down the victim rhetoric.
- ronsor 21d agoYes, I think this is the main issue. I don't care what policies the AI labs have or enforce, but they need to stop acting like ToS violations are an international crisis demanding intervention instead of a boring civil dispute at most.
- darth_avocado 21d agoI would argue they should have to. They scraped data off others, a lot of whom did not want that data to be used for AI training, and still had to share it with the frontier labs. It’s only fair they should have to hand it back. The only way US maintains dominance over Chinese models is by having an ecosystem of models. Relying on a small set of frontier labs will only let you get ahead temporarily. I agree with Gary Tan on this one.
- dannyw 21d agoGenerally companies are welcome to choose to who to provide service to, as long as it's not discriminating against a protected class, or ruled as anticompetitive (which is a very high bar in recent case law; even if the same 1890s-era laws are still on the books). I don't think a correct remedy is to require companies to provide services even if they want to. A simple example: you drop a client because their asks / ways-of-working / etc is more headache and costs than it's worth. I've done that before, multiple times, in my freelancing life.
- impossiblefork 21d agoMorally I agree, but since there's probably a lot of LLM text in the training data, distilling on another model will probably make your model copy the values encoded into the other model as well, even in cases where you only distill on value-neutral stuff. By copying their programming style, you'll move the model towards that way of writing, which will move the model towards the values expressed in those documents. I feel that Deepseek v4 got so claudified at the end that it was like Claude.
- pj_mukh 21d agoI wonder if along with “Pacing the frontier”, we can get the frontier labs to Share the raw data. I’m sure the labs claim that their real innovation is in the RLHF, training and architecture. Keep that and just share the raw data somewhere.
- throwawayk7h 21d ago"Strip-mine" is not correct. The commons are all still there and you can still train on them just like the frontier labs did. Of course, it may be illegal to do so, but that's not any different than before.
- vermilingua 21d agoYknow, aside from the books they are literally destroying while scanning
- protocolture 21d agoBooks they wouldnt need to destroy if they were simply permitted to torrent. Daily reminder that piracy is the only enduring archive mechanism.
- vermilingua 20d agoNo, these are books that aren't online, which means not only are they not contributing to the commons, they are irrevocably salting the earth (irrevocably because let's be honest, anything going into their archives isn't coming out without legal or actual violence) Also, permitted or no, they are definitely torrenting. I would be deeply surprised if they hadn't already leeched every torrent on public trackers. The only reason they (probably) haven't depleted all the private trackers too is that they would be required to actually contribute back, which as above is never going to happen.
- protocolture 20d ago>irrevocably because let's be honest, anything going into their archives isn't coming They were torrenting, they were sent to court and settled for big $$$$. The only other method available to them now is scanning, and scanning at scale requires the books destruction. I agree that they should definitely be required to see the new scans but that would just be more $$$$ they get charged if caught.
- 20d ago
- deleted 21d ago[deleted]
- soundworlds 21d ago100% - the work came from the people, it should go back into the hands of the people. I also think if Anthropic and OpenAI had been releasing Open models along the way, people wouldn't be nearly as suspicious of them.
- CamperBob2 20d agoOpenAI has released a few open-weight models, including some that were considered quite competitive back in their day. It's been a while, though. Anthropic has never released anything but FUD.
- larodi 21d agoGiven (A) : > He also notes that the proprietary AI labs didn’t ask permission when they vacuumed up as much human knowledge as they could to train their models. They famously ingested plenty of copyrighted material without the permission of those intellectual property holders. And many people's shared opinion (B): >> I don't expect Tan's position to be based on any kind of real moral high ground, but his conclusion is correct. ... It is very difficult to actually say NO to the fact that (A) was done, which then leads logically to conclusions as (B). But also we should remember that if these two hold (and (A) is an axiom more or less now), then it comes as no surprise that then also all opensource licensing is immediately rendered void and null, as keeping it would contradict (A) and would go against the very common and consequential logic in (B). Copyright is so dead. And it was not me killing it with a cynical post on HN. Dunno why so many people still fail to face it. There is no way it can exist in its current form, because then immediately (A) happens and (B) follows.
- TZubiri 21d ago>I love the "illicit distillation attacks" framing from the incumbents. There's nothing illicit. There's no attack. You just don't like it because it threatens your market position and business model. With what knowledge are you claiming this? If it turns out companies are using IP proxy networks would you change your mind? What if the IP Proxy networks were used by criminals for similar attacks like DDoS or plain cyber attacks? What if the source of the IP proxy networks were residential addresses to avoid detection? What if the way these IPs were acquired were through pwned devices? What if the credit cards used do not identify the company that carries the attack? What if they use the employee's personal credit cards? What if it's family members of employees? What if it's a network of personal credit cards where cc owners get a payment for making a purchase on their name? What if they are stolen ccs? Not just a hypothetical btw, I believe almost all of these are true.
- ejj28 21d agoSupposing it is true, then I'm glad the AI companies are getting a taste of their own medicine.
- qlte 20d agoYou can't just throw in "or what if they're stolen credit cards" at the end to lump in blatantly illegal/unethical activity with the far weaker preceding points that can be summarized as "uses a residential proxy".
- hn993302 21d agoThere's no moral high ground here, it's just that nobody would invest in training publicly usable models if they could be easily distilled. Not that I think there should be laws against it or that such laws would even work; they're going to have to protect themselves.
- dannyw 21d agoFewer people would create scientific or artistic works if they could just be copied or used without protection either; or so is the premise behind copyright and intellectual property; even being deeply embedded into the US Constitution (Art 1, Sec 8, Clause 8). There is sooo much irony here.
- hn993302 21d agoI agree. Some existing licenses don't seem compatible with AI training. If they don't go back to rectify that, at the very least you should be able to license your work in a way that explicitly prohibits AI training. They can pay if they want to use it.
- stale2002 20d ago> it's just that nobody would invest in training publicly usable models if they could be easily distilled. Thats literally what is happening right now though. People are spending hundreds of millions on a training run, and then people are distilling them, fairly easily, and making cost competitive models. We are seeing all of this in action right now.
- dannyw 21d agoThese frontier labs violate billions of terms of services across the web, that prohibit scraping / automated access / etc. Most sites have a clause, it’s basically standard boilerplate. So why is their own ToS so special? :)
- JeremyNT 20d ago> I love the "illicit distillation attacks" framing from the incumbents. There's nothing illicit. There's no attack. You just don't like it because it threatens your market position and business model. I think a lot of this is done using stolen black market API credentials, which is why it might be somewhat accurate to consider it illicit.