3 ms·
> counted word occurrences and presented statistics about them. I don't think you even need fair use for this, because this is something you obviously are allow
by warning26 3y ago
> counted word occurrences and presented statistics about them. I don't think you even need fair use for this, because this is something you obviously are allowed to do, without any permissions
You're pretty much describing exactly what an LLM "learns" about text. I agree that it should obviously fall under fair use, but as the author of this article found out, there are quite a few who (very vocally) disagree.
- tensor 3y agoI think there is a big difference in terms of data recovery though. You can't take a compression algorithm, for example, and claim that its "just some statistical analysis" when it can reproduce the original perfectly. Heck, even if it can reproduce it approximately, that's a lot different than what we see in this particular example, where the data could not be used to reproduce a text at all.
- make3 3y agogeneration of related text vs analysis of human understandable facts is very different in the mind of most people. I think that using an LLM to get insights on the text should be ok, it's the generation part that scares them. probably rightly so.