3 ms·
Yes, that is interesting. It sounds like he was aware of the theoretical possibility of a book being regurgitated verbatim. Do you know if he was aware it had b
by brlewis 2mo ago
Yes, that is interesting. It sounds like he was aware of the theoretical possibility of a book being regurgitated verbatim. Do you know if he was aware it had been done? https://news.ycombinator.com/item?id=49000742 https://news.ycombinator.com/item?id=49000742
If he was not aware, I wonder if he still would have described the process as "exceedingly transformative" had he been aware.
- FeepingCreature 2mo agoNote that they're testing for 100-word passages. This is a level of memorization that avid readers can credibly also reach. Note also that Sonnet 3.7 had to be jailbroken. Note also that they got high memorization for a few books that were widely quoted. The books in question can probably also be "retrieved" by putting phrase prefixes into Google, which is probably why Sonnet 3.7 knows them with the precision of a fanboy. Material being widely repeated in the training set is a well-known cause of memorization.
- globular-toast 2mo agoNo "avid reader" could recall anywhere near that much text. That takes dedicated effort to commit to memory. Copyright was never meant to stop people copying books anyway, it was meant to stop machines (ie. printing presses) copying them. Edit: Apologies, I misread it as "100 pages". My point about copyright still stands, though.
- FeepingCreature 2mo agoI disagree that avid readers cannot complete entire passages from books they've read several times when fed a prefix.
- leni536 2mo agoAnd can these avid readers publish these recited passages without infringing copyright?
- FeepingCreature 2mo agoI mean, the debate would then turn on whether publishing the passages and publishing the model is the same sort of thing. I think there's mainly two views: "we know the passages are in there, so publishing the model is publishing the passages is copyright violation", and "nothing happens until you go through considerable effort to elicit the passages, so the user is committing copyright violation using the model as a tool." Personally I think our legal system is just not set up for a world where we can download mindstates in numeric form. Would a sufficiently detailed recording of my brain violate copyright? If simulated, it could certainly be elicited to commit violations. edit: At any rate, Anthropic are not publishing the Sonnet 3.7 weights.
- JAlexoid 2mo agoI used to use the initial letters of a whole paragraph from Lord of The Rings as my password. Some of us have a good enough memory.