8 ms·
That's why copyright holders for reference works have been using copyright traps for ages. That's where you include a fictional town in a map, a nonsense word i
by humanistbot 3y ago
That's why copyright holders for reference works have been using copyright traps for ages. That's where you include a fictional town in a map, a nonsense word in a dictionary, or a fake person in your phone book. If your competitors reproduce the trap, then that's clear evidence you can use in court.
https://en.wikipedia.org/wiki/Copyright_trap https://en.wikipedia.org/wiki/Copyright_trap
- iudqnolq 3y agoIf you look at the Legal Action section of your link you'll see the line "However, the case was dismissed" quite a few times. That's because data isn't copyrightable. Edit: As sroussey points out s/isn't copyrightable/isn't copyrightable in the USA
- sroussey 3y agoNot in the USA, but it is in the EU and elsewhere.
- whiplash451 3y agoHow do you define the geolocation of data? If my website is hosted in EU but a company scans it from the internet in the US, how could they possibly know it is hosted in EU?
- iudqnolq 3y agoWhich country's laws apply and what remedies you can get if they were violated is far more complicated than geolocation of data. But very broadly speaking you would need to sue in an EU court to enforce EU law. And you could sue a US company in specific EU country's court if the company had more than some minimum level of connection to the that country. The country the data is hosted in isn't key, though it can be evidence of connection to that country.
- z3t4 3y agoWhere the data is stored does not matter much. Laws deal with people and companies, so it matters where you live or where your company operates. So if you live in the US you don't have to worry about EU laws unless you do buisness in EU.
- concordDance 3y agoHence why you should live on that unclaimed but of land in Africa. :D
- AnthonyMouse 3y agoThe other problem with these "copyright traps" is that they do nothing to prove someone copied the legitimate parts of the data. Suppose you recreate the entire dataset from scratch. Then someone notices (e.g. using an automated comparison) that the "trap" is in the other dataset but missing from yours, and submits it to you to add. This is arguably too small an addition to be copyrighted on its own, but regardless of that, it would then be all you have to remove to get back to a clean version. And since it's erroneous data, you would want to remove it anyway.
- plasticchris 3y agoMy favorite of these was a town founded to match the map. Pretty sure I heard an npr story on it.
- jprete 3y agoThe relevant line is “information alone without a minimum of original creativity cannot be protected by copyright”. There is definitely creativity in writing code; it’s not a completely deterministic translation of even a complete specification.
- iudqnolq 3y agoOh absolutely. I was speaking only about the comment I replied to.
- cxr 3y agoIt's occasionally explained—but still not widely understood, I'd wager—that this is the reason why so much GNU code is hard to follow. In the US legal system the merger doctrine is a concept whereby a given expression cannot be granted protection if it's not sufficiently creative—and there only so many ways to express something when stripped down to its fundamentals. In response to this, RMS and Moglen encouraged contributors from very early on to try to express the inner workings of GNU utilities in creative and non-obvious ways out of caution against the possibility that the copyleft obligations of the GPL wrt a given package could be nullified by a finding in court that it did not pass the threshold for creativity.
- noirscape 3y agoGNU code is partially hard to follow because of RMS paranoia, but that mostly manifests itself in the code being weirdly structured. The far bigger reason is that GNU code tends to run with really strange optimizations and project decisions since they want their tools to be able to run on ancient mainframes that practically nobody uses anymore, so everything is overoptimized for that.
- netfortius 3y agoI think I mentioned this before, in another context: the solution is known as "honeytoken", and it is equally applicable in computer security.
- jschrf 3y agoAlso, re: maps, fake streets and cul-de-sacs that don't exist. I've set a "trap" myself years ago in code in a novel solution at the time for uploading photos from iOS non-interactively after the fact. It was to support disconnected field workers taking photos from iPhones/iPads, with the payloads uploaded at a later date. Chunked form data constructed in userland JS was the solution. Chunk separator was 17 dashes in a row (completely arbitrary), company name in 1337 speak, plus 17 more dashes. Found a competitor that had copied the code, changing only the 1337 speak part. 17 dashes remained on each side. Helped me realize that they had unminified and indeed ripped off our R&D work. Wonder if Copilot could be gamed the same way.
- peteradio 3y agoHow did you manage to find that your competitor copied your code?
- aetch 3y agoJavascript
- jschrf 3y agoYeah. The feature set offerred by the competitor was similar to ours, and we went through the wringer building that solution, so i unminified their code and sure enough...it easn't exactly theirs. Oh yeah and they ripped off our website too. That was the first clue haha.
- tedivm 3y agoWe don't need the copyright traps here though as Github openly admits to using the public code for training. They just don't care that they're essentially license laundering code since they can make money doing it. That said we used copyright traps at Malwarebytes, which is how we found out that iobit was stealing our database.
- peytoncasper 3y agoWhat happens if GitHub didn't use GPL licensed code, but still generated code that was identical to GPL licensed code?
- layer8 3y agoThey’d have to prove to the court that the former is true despite the latter happening, which I imagine would be difficult to do in practice.
- tedivm 3y agoWe know that isn't the case because we can see code being reproduced even with comments, and Github has been open about the fact that they used everything they had in training. That said, lets say there's a new model that explicitly excluded closed source and copyleft licenses. Well, the MIT, MPL, Apache, BSD- they all say you can't strip their licensing off. Okay, so to get to the spirit of your question, lets say Github managed to program a model that worked using only their own code or code that was explicitly put in the public domain. If Github managed to reproduce code that wasn't in the training set, then it can't be accused of copying it. At that point the argument could be made that it independently created it. At the same time algorithms can't be copyrighted, but implementations of an algorithm can be, so if Github was basically just spitting out an algorithm that just happened to be implemented similarly to how some other code it wasn't trained on implemented it, then I would say there was no copyright violation.
- bryanrasmussen 3y ago>We know that isn't the case because we can see code being reproduced even with comments If the comment is something like //check fromIndex is greater than toIndex then that is not any more individualistic or different than the actual function. Sadly, many people comment like this, on the other hand if it reproduced a comment with typos or something more complicated like /* this hack is because Firefox's implementation of SVG z-indexing does not match how Chrome or Safari does it - please read this article ...url...*/ then yeah, then you would have something
- ljm 3y agoI first saw this in action on StackOverflow when, during an interview, a candidate copy-pasted a solution verbatim including the attribution. Didn't even give it a second thought, like they didn't even read the code or what it was doing. It wasn't the right solution to the problem in question, for what it's worth. Just manually did what GPT does now.