4 ms·
> I love ML for the sake of ML. I wanted to prove that individual researchers can contribute in a way that matters. Sounds like someone so focused on "if they
by srhtftw 3y ago
> I love ML for the sake of ML. I wanted to prove that individual researchers can contribute in a way that matters.
Sounds like someone so focused on "if they could" but not giving much thought about "if they should".
I don't want you or any responsible ML researcher to go to prison. Responsible researchers make mistakes, but responsible researchers also consider the consequences of their actions. How much consideration was given to those? And given everything you know now, where would you draw the line? At what point do the consequences of this sort of research become indistinguishable from theft? Suppose the research enabled the piracy of a high fidelity copy of a single author's work without their permission? Would that be wrong? What about ~2e5 low fidelity copies of who knows how many man years of labor?
- sillysaurusx 3y agoI’m not sure. Someone pointed out that I should’ve put a non commercial license on books3. But books3 isn’t my data to license. As for actual impact, books3 made its way into BloombergGPT and LLaMA. I’m just happy it made LLaMA a little better. Would I do it again? Well, given that there’s little benefit other than warm feelings about LLaMA and pride that hackers can make an open source ChatGPT competitor with this, that’s a tough question. As for the line, I invite you to download books3 and try reading one of the books. The experience is positively awful. Notepad isn’t a good book reader; it’s missing images, which renders lots of the technical books less useful; and the markdown format isn’t even modern markdown syntax. Code snippets are indented with 4 spaces instead of surrounded with triple backticks. Meanwhile anyone seriously interested in reading a book will go straight to libgen and get the real book themselves. So I’m skeptical of claims of harm. As for whether it’s wrong, let history be my judge.
- srhtftw 3y ago> As for the line, I invite you to download books3 and try reading one of the books. I'm not going to and I suspect anyone who values their time won't either, especially when my public library meets my needs. But the line I'm asking you to look for isn't the one around a single book - it's the aggregate value previously locked in hundreds of thousands of them now released by the power of research like yours. Even if you don't agree, I hope you appreciate that authors whose works are protected by copyright will feel entitled to a share of that value. Most of them didn't write their works to further ML research. > Would I do it again? ... that’s a tough question. I am reminded of when I was young how accessing an unprotected computer wasn't in itself a crime. People like me hacked into systems mainly out of curiosity. Then the media started making movies about teenagers starting nuclear wars with dial-up modems. Then journalists jumped on the hype train to sell their books and convinced people hackers needed to be punished. Then some went to jail and we got things like the CFAA. In a few short years what might have been wrong but wasn't explicitly illegal became so. You are clearly a smart and intellectually honest person but society often treats such people poorly - especially when confronted by scary new technology - and the tone of your answers are giving me serious Aaron Swartz vibes. There's a lot of money at stake here.
- anewhnaccount2 3y agoIt's fairly clear to me that trained models are a transformative rather than derivative works. However, without distributing forms which are not transformed, open/distributed/permissionless work in this space is either severely hampered, or basically impossible. I think a lot of people instinctively feel that removing barriers to their work is particularly important, since if we don't it could remain impossible for anyone outside a few megacorps -- who have the resource to obtain whatever they need in multiple ways without the additional liability of distributing non-transformed copies externally -- to build highly scaled up machine learning models.
- srhtftw 3y agoI largely agree. I would also like to see more research into how to give authors incentives to license their works for transformative purposes such as training and ways to credit them when the results of inference draw disproportionately on a particular author's work.