4 ms·
As described, this would not be the same thing. If the AI is looking at the source and effectively porting it, that is likely infringement. The idea instead sho
by robmccoll 7mo ago
As described, this would not be the same thing. If the AI is looking at the source and effectively porting it, that is likely infringement. The idea instead should be "implement Minecraft from scratch" but with behavior, graphics, etc. identical. Note that you'll need to have an AI generate assets or something since you can't just reuse textures and models.
- Gigachad 7mo agoAI models have already looked at the source of GPL software and contain it in their dataset. Adding the minecraft source to the mix wouldn't seem much different. Of course art assets and trade marks would have to be replaced. But an AI "clean room" implementation has yet to be legally tested.
- reverius42 7mo agoFor copyright purposes I think there is an important legal distinction between training data (fed in once, ahead of time, and can in theory no longer be recovered as-is) and context window data (stored exactly for the duration of the model call). I'm not sure there should be, but I think there is.
- NewsaHackO 7mo agoThat's why he is saying it's not equivalent. For it to be the same, the LLM would have to train on/transform Minecraft's source code into its weights, then you prompt the LLM to make a game using the specifications of Minecraft solely through prompts. Of course it's copyright infringement if you just give a tool Minecraft's source code and tell it to copy it, just like it would be copyright infringement if you used a copier to copy Minecraft's source code into a new document and say you recreated Minecraft.
- paxys 7mo agoIs there a legal distinction between training, post-training, fine tuning and filling up a context window? In all of these cases an AI model is taking a copyrighted source, reading it, jumbling the bytes and storing it in its memory as vectors. Later a query reads these vectors and outputs them in a form which may or may not be similar to the original.
- SatvikBeri 7mo agoJudges have previously ruled that training counts as sufficiently transformative to qualify for fair use: https://www.whitecase.com/insight-alert/two-california-district-judges-rule-using-books-train-ai-fair-use https://www.whitecase.com/insight-alert/two-california-distr... I don't know of any rulings on the context window, but it's certainly possible judges would rule that would not qualify as transformative.
- deleted 7mo ago[deleted]
- deleted 7mo ago[deleted]
- derangedHorse 7mo agoThe context window is quite literally not a transformation of tokens or a "jumbling of bytes," it's the exact tokens themselves. The context actually needs to get passed in on every request but it's abstracted from most LLM users by the chat interface.
- alpaca128 7mo agoWhat if Copilot was already trained with Minecraft code in the dataset? Should be possible to test by telling the model to continue a snippet from the leaked code, the same way a news website proved their articles were used for training.
- NewsaHackO 7mo agoI feel as though the fact that you are asking a valid question shows how transformative it is; clearly, while the LLM gets a general ability to code from its training corpus, the data gets so transformed that it's difficult to tell what exactly it was trained on except a large body of code.
- RhythmFox 7mo agoThen the training itself is the legal question. This doesn't seem all that complicated to me.
- Gigachad 7mo agoThis would still be true of the case where you ask an LLM to rewrite a program while referencing the source. Unless someone was in the room watching or the logs are inspected, how would they know if the LLM was referencing the original source material, or just using general programing knowledge to build something similar.
- phendrenad2 7mo agoIt's not equivalent, but it's close enough that you can't easily dismiss it.
- sneak 7mo agoYou are confusing training data with context (prompts).
- NiloCK 7mo agoA room "as clean" as the one under dispute (chardet) is very easy to replicate. AI 1: - (reads the source), creates a spec + acceptance criteria AI 2: - implements from spec AI 1 is in the position of the maintainer who facilitated the license swap.
- smsm42 7mo ago"Behavior, graphics, etc." would likely constitute separate IP from the code. I am not sure there's a model that allows you to make AI reproduce Minecraft without telling it what "Minecraft" is - which would likely contaminate it with IP-protected information.
- yunnpp 7mo ago> Note that you'll need to have an AI generate assets or something since you can't just reuse textures and models. As far as I know, you can as long as you own a copy of the original. In other words, you can't redistribute the assets, but you can distribute the code that works with them. This is literally how every free/libre game remake works. The copyright of your new, from-scratch code, is in no way linked to that of the assets.