3 ms·
> This means that OpenAI should be sued every single time someone manages to extract what is considered as copyrighted material from their software. I agree! I
by subroutine 3y ago
> This means that OpenAI should be sued every single time someone manages to extract what is considered as copyrighted material from their software.
I agree! If GPT4 is outputting copyrighted material beyond what is considered fair-use (i.e. substantively more than what is provided by say, google books), I agree that is copyright infringement.
Indeed it is about the output, and making stuff available that people would otherwise have to pay for (or more precisely, enough of the copyrighted work that a person would have reason not to pay for the original work, causing a material loss to the original author) - that is a fineable violation imo.
Something else to think about... I work in biotech and have published articles in scientific journals on cellular and molecular level disease sequelae (such articles are also protected by copyright). Models trained on scientific literature are now being used for novel drug discovery and disease treatment pathways. These models are already outputting suggestions that seem very promising. Shall we also not provide these models access to the full corpus of scientific literature? It would significantly handicap these models to not have access to copyrighted scientific works. On one hand, some proportion of researchers will retain their jobs that would have otherwise been outsourced to LLMs (perhaps even myself). On the other hand, some amount of future patients will suffer or die from a disease that would have otherwise been cured.
- palata 3y agoThat actually brings another point: if you train LLMs on scientific papers, at least in some domains it will make it easier to write a lot of papers. I am not an academic, but it is already my impression that there are a lot of low-quality papers out there. What if now many more get generated by LLMs? Won't that be a problem?
- subroutine 3y agoThe low quality problem with primary research publications is not the writing but poor experimental design, misrepresenting experimental results, shoddy statistical analysis, and putting null results into file cabinets. Summarizing research findings isn't the crux of the problem, so if anything if an LLM can help the author perform a clearer and more concise writeup I'd see it as a net benefit.