17 ms·
GitHub repository for Sedgewick's Algorithms is taken down
- seaman1921 5y agoIf this was a youtube channel takedown people would be mad at Google, but good to see at least here HN is blaming the bad actors like they should - ie. the copyright laws and its abusers.
- maweki 5y agoPeople are mad at google for _automatic_ takedowns, as they enable bad actors and are arguably not in the spirit of the law. In this case here, github does not employ an overzealous automated system. This is a manual DMCA claim and the issuer is the single actor here, who could be to blame.
- judge2020 5y agoPeople got mad at Google for automatically honoring DMCA takedowns, which is what GitHub is doing, as they both have to if they want to maintain not being legally liable. Content ID is also a system people are mad with but has nothing to do with DMCA besides Google saying 'pretty please only enroll content you own into Content ID'.
- maweki 5y agoFor github this is not an automated process, nor is the policy based on a presumption of guilt: https://docs.github.com/en/github/site-policy/dmca-takedown-policy#a-how-does-this-actually-work https://docs.github.com/en/github/site-policy/dmca-takedown-...
- judge2020 5y agoAssuming the claimant doesn’t mess up the complaint it might as well be automated - GH has only voluntarily not honored a request once with YouTube-dl (that I’m aware of) and that still was after YT-dl removed the README reference of using it to download music videos.
- matsemann 5y agoDifference is that Github follows DMCA practices, while YouTube has their homemade system that bans you without a fair way of appealing.
- dan-robertson 5y agoGoogle have their own system for copyright claims which people get mad about. GitHub follows the DMCA laws which they have to follow. I put it to you that your implied claim that these situations are basically the same is false.
- leeuw01 5y agoContext?
- marktucker 5y agoKevin Wayne is one of the authors of Algorithms 4th ed. by Robert Sedgewick and Kevin Wayne and he has a github repository called "algs4". Pearson sent github a DMCA notice for this repository and another. I'm not sure exactly what content he hosted in "algs4".
- hvdijk 5y agohttp://web.archive.org/web/20201016112929if_/https://github.com/kevin-wayne/algs4 http://web.archive.org/web/20201016112929if_/https://github....: > This public repository contains the Java source code for the algorithms and clients in the textbook Algorithms, 4th Edition by Robert Sedgewick and Kevin Wayne. This is the official version—it is actively maintained and updated by the authors. [...]
- mrighele 5y agoThe DMCA notice linked on the page [1] gives a few more details. We have received information that the domain listed above, which appears to be on servers under your control, is offering unlicensed copies of, or is engaged in other unauthorized activities relating to, copyrighted works published by Pearson Education, Inc. Copyright work(s): Deep Learning with R 9781617295546 Algorithms 9780321573513 Copyright owner(s) or exclusive licensee: Pearson Education, Inc. Copyright infringing material or activity found at the following location(s): https://github.com/jjallaire/deep-learning-with-r-notebooks https://github.com/jjallaire/deep-learning-with-r-notebooks https://github.com/kevin-wayne/algs4 https://github.com/kevin-wayne/algs4 [1] https://github.com/github/dmca/blob/master/2021/04/2021-04-20-pearson.md https://github.com/github/dmca/blob/master/2021/04/2021-04-2...
- troelsSteegin 5y agoThat seems odd, given that Kevin Wayne is co-author of the 4th edition of "Algorithms".
- dkjaudyeqooe 5y agoWhat seems odd is that the code is freely licenced to the public, no matter who wrote it. These DMCA requests are created and acted upon without the slightest checking, with no penalty. There should be some sort of modest cost somewhere to stop this nonsense.
- viraptor 5y agoAs long as they're properly filled out, they have to be acted on if the provider wants to keep their safe harbour status.
- qayxc 5y ago> There should be some sort of modest cost somewhere to stop this nonsense. Playing the Devil's Advocate here: But wouldn't that in turn incentivise actual copyright infringement on account of it potentially being cheaper to mass publish copyrighted material as opposed to request it being taken down?
- bnj 5y agoI hadn't thought about that angle before. Is the public interest best served by adding friction to copyright enforcement, or by adding friction to fair use?
- qayxc 5y agoGood question. I honestly don't know. We would have to re-evaluate the value of IP as whole first, I guess.
- lanstin 5y agoCertainly the right to copy is different when copying is so much cheaper than when the laws where written. We gave up:the right to copy when you couldn’t really copy. Now we can copy and maybe we want to retain that right to the people. Or at least rejigger so we either get more or give less via this exchange
- mellosouls 5y agoThe notice is here: https://github.com/github/dmca/blob/master/2021/04/2021-04-20-pearson.md https://github.com/github/dmca/blob/master/2021/04/2021-04-2... Kevin Wayne (repo owner presumably) is listed as co-author on the book that is published by Pearson, who have issued the notice. https://www.pearson.com/us/higher-education/program/Sedgewick-Algorithms-Fourth-Edition-Deluxe-Book-and-24-Part-Lecture-Series/PGM2366556.html https://www.pearson.com/us/higher-education/program/Sedgewic...
- phoe-krk 5y agoYours sincerely, [private] Internet Investigator On behalf of: Pearson Education, Inc. ------------------------------------ Seems like another glorious victory of the overzealous online copyright strikers. Nothing out of the ordinary, will likely get sorted out in a day or two.
- st_goliath 5y ago> Internet Investigator Haha... that's a brilliant euphemism for "profession Google user". Or maybe we could call it "serial scrapist"? Of all the silly titles I tried to come up with, this sure hadn't crossed my mind yet. And it has a catchy ring to it too. "Internet Investigator" got to be somewhere up there between "Adult Software Engineer" (middle ground between "Junior" and "Senior") and "Server Captain/Commodore/Admiral".
- rzzzt 5y agoI'm a bit of Web Browsing Enthusiast myself.
- BenFeldman1930 5y agoCheck out our recent job offers at 221B Baker Street, London.
- Operyl 5y agoIt literally is the name of the company, haha. https://theinternetinvestigators.com/ https://theinternetinvestigators.com/
- kevinventullo 5y agoI give a lot of credit to Sedgewick’s course for helping me break into the software industry, having come from an adjacent field in academia. I hope they can get the repo back up.
- vanderZwan 5y agoIn my case it didn't help me break into the software industry - somehow I already managed to do that despite some well-deserved imposter syndrome at the time, but it definitely leveled up the quality of my code
- caymanjim 5y agoThis isn't worth getting your feathers ruffled over. Some algorithm or barely-paid human automaton flagged it, someone higher up will notice it, it'll be added to an exception list, problem solved. It's not ideal, but it's not malicious or nefarious or even worth talking about every time something like this happens. It's a simple mistake that has a simple solution.
- syshum 5y agoI 100% disagree, the fact that as a society we are accepting of this kind of thing as "just the way it is", is the problem "Ohh its just some bot" should never be an acceptable response, and people creating these bots should have massive fines attached to their false positives. Otherwise there is no incentive to actually validate or improve what these automation systems are doing, the incentive today is the flag everything and worry about it later... That is not something society should accept
- 2pEXgD0fZ5cF 5y ago> or even worth talking about every time something like this happens I think it is, because the core problem is not the wrongful takedown of a single repository itself, it's the outlook that the only way to get these "mistakes" reversed is the hope that it will reach the news, Youtube/Google shows us what the latestage of this nightmare looks like. Github is obviously not nearly this bad (yet?). This case will obviously be restored, it's a popular author, but what about people that don't have this reach? I don't see how talking about this situation can be a bad thing.
- tzs 5y ago> I think it is, because the core problem is not the wrongful takedown of a single repository itself, it's the outlook that the only way to get these "mistakes" reversed is the hope that it will reach the news, Youtube/Google shows us what the latestage of this nightmare looks like. YouTube/Google is bad because they use their own system instead of or in addition to the system required to get the DMCA safe harbor. For sites that simply follow DMCA, such as GitHub, there is no need for the issue to reach the news. There is no need to be popular or have any reach. The most obscure poster with no audience can easily get their content restored. Here is what happens at such sites: 1. Someone claiming to represent the copyright owner files a take down request. 2. The site temporarily takes the content down, and notifies whoever posted the content. 3. The poster files a counter-notice saying that they have the legal right to post the content. 4. The site put the content back up, and notifies the party that filed the take down notice, telling them who filed the counter-notice. The site is then out of the loop. The content is back up, and the party claiming copyright violation cannot sue the site over this. The site is now in the DMCA safe harbor. If the party claiming violation wants to take it farther, they need to go to court and sue the party that posted the content. The counter-notice is trivial to fill out and file. Here's the one for this case [1] if you want to see how simple the form is. [1] https://github.com/github/dmca/blob/master/2021/04/2021-04-29-pearson-counternotice.md https://github.com/github/dmca/blob/master/2021/04/2021-04-2...
- truth_ 5y agoNot totally unrelated: Pearson is a piece of trash company. They sell very overpriced textbooks and always use DRM in online copies. They lobby for standardized tests in states and get them mandated. Students are put through them. And guess who provides the certificate required by teachers to evaluate those exam copies? Yes, Pearson. And they make a ton of money. Please watch this talk for more context (even if you don't like TED talks): https://youtu.be/BnC6IABJXOI https://youtu.be/BnC6IABJXOI
- 2pEXgD0fZ5cF 5y ago> very overpriced textbooks Not just way overpriced, the few Pearson books I happen to own are of an alarmingly low quality. The pages of a phone book look and feel like premium paper compared to my copy of Blitzer's College Algebra.
- vector_spaces 5y agoNot to mention some of them have wonderful anti-features, such as physical books that have a handful of online-only chapters that you can only access if you buy the textbook new, thereby reducing the value of used copies.
- generationP 5y agoYou'll then appreciate the "Pearson New International Editions" (with a chapter and the preface removed to tank the resale value).
- vector_spaces 5y agoOh yeah, those are all a real hoot with the lovely almost see-through paper and shoddy binding and printing errors like upside down pages.
- shakaijin 5y agoAlso, they release new versions of books all the time and force their use in classes, making it impossible to use 2nd hand textbooks in a lot of cases. *I mean, come on. How of ten do fundamental physics change... Pearson is a company that should not exist.
- DerNuntius 5y agoHere is the counternotice: https://github.com/github/dmca/blob/master/2021/04/2021-04-29-pearson-counternotice.md https://github.com/github/dmca/blob/master/2021/04/2021-04-2...
- mirthflat83 5y agoWhat a shitshow
- fartcannon 5y agoIs it time for a sci-hub like event to replace github? It would be better for the repository of nearly all modern public programming creativity be something more like Wikipedia, and less like LinkedIn.
- easton 5y agoGiven the technical nature of the work, developers who can't jive with GitHub's policies can throw up a GitLab/Gitea server in an hour and evade censorship. And of course, as we learned with youtube-dl, the issue isn't the code/commits (since everyone gets that with a git clone), it's the issues and PR history. Any public file host has to deal with the DMCA, and (for now, barring any evil brought on by Microsoft), I'd bet on GitHub siding with the developer over the lawyers. They had lots of incentives to go the other way in the youtube-dl case and they didn't, and then they put a bunch of money in a fund if someone wants needs a lawyer to get their project back online.
- trasz 5y agoShould we perhaps invent a standard convention for keeping issues and PR history along with the usual source code in the repo itself?
- yissp 5y agohttps://fossil-scm.org/home/doc/trunk/www/index.wiki https://fossil-scm.org/home/doc/trunk/www/index.wiki from the sqlite guy.
- trasz 5y agoBut this is a whole another SCM instead of Git, isn’t it? What I’ve meant was just a convention within Git, same way we have a convention for README.md and “doc/“.
- account42 5y agoThere is git-bug [0]. Haven't had the time to evaluate it yet, so no idea how useful it is. [0] https://github.com/MichaelMure/git-bug https://github.com/MichaelMure/git-bug
- airstrike 5y agoIf you're just looking for the algorithms, there's always this C# port (from 7 years ago): https://github.com/angellaa/algs4 https://github.com/angellaa/algs4 Interestingly, the README includes an e-mail exchange with Kevin, who noted the code is GPL
- gravypod 5y agoIt looks like the original code is located here: https://algs4.cs.princeton.edu/code/ https://algs4.cs.princeton.edu/code/ Found this in the link you provided. As you can see from the code there is a GPL license footer on each file.
- valkum 5y agoThis site even links to the now taken down Git repo.
- IceWreck 5y agoWe should have a penalty/fine for fake DCMA requests. Big corps keep sending automated DCMAs with little consequence and often the little guys do have the resources to fight back.
- rvz 5y agoAnother reason to self-host. Rather than sit on someone else's platform.
- acdha 5y agoThis is in the law, but it’s rarely enforced. Stepping up education for judges and support for people filing lawsuits against major companies would be a relatively easy way to start making the automated bots riskier to run.
- Semaphor 5y agoIs it? I always thought there is only a fine for misrepresenting that you are representing the copyright owner. Not for falsely claiming a piece of work is indeed infringing.
- acdha 5y agoYes — see page 12 here: https://www.copyright.gov/legislation/dmca.pdf https://www.copyright.gov/legislation/dmca.pdf “Penalties are provided for knowing material misrepresentations in either a notice or a counter notice. Any person who knowingly materially misrepresents that material is infringing, or that it was removed or blocked through mistake or misidentification, is liable for any resulting damages (including costs and attorneys’ fees) incurred by the alleged infringer, the copyright owner or its licensee, or the service provider.” Here's the actual text: https://www.law.cornell.edu/uscode/text/17/512#f https://www.law.cornell.edu/uscode/text/17/512#f I'm not a lawyer but my understanding is that by now the courts have confirmed that it's not just enough to say that you own the copyright on the material: you also have to confirm that the person you're sending the claim to doesn't have a right to use the material as well and that includes fair-use. If you were, say, using a Disney clip in a film studies context they would likely be liable if a takedown bot sent a DMCA claim unless they could show that you did something like posting the entire film claiming it was for “study” purposes. In this case, that's highly relevant since this repository was apparently released under an open source license so even if Pearson held the copyright they wouldn't be able to take away the permission granted by the open source license unless they could prove that it was never legally approved for release under that license.
- st_goliath 5y agoWhen I was an undergrad student, we had an algorithms and data structures lecture based on the Java version of the book. We also got an archive with the samples to work with that was IIRC GPL licensed. Did this repo contain the code or the entire book as well? In the former case, if I remember correctly, this DMCA claim would clearly be bogus. Yet another case where sending out DMCA claims having penalties for the sender if they are fraudulent, but putting a burden on the recipient who might suffer if they don't immediately act.
- eaa 5y agoSo, distributed git has become sort of centralised. Maybe people depend too much on github? I wonder, what would happen to IT over the world if github will be down for a week or a month?
- csunbird 5y agoThere is literally nothing preventing someone to host they code on gitlab/sourcehut/their own gitea server. Most people use github because it is free, but the git protocol is far from centralized.
- judge2020 5y agoThankfully `git clone kevin-wayne/algs4` doesn't default to asking github.com like some package managers [[docker]].
- noir_lord 5y agoUtter chaos for a week, then a week or two of swearing at gitlab then the world would continue - since while we use github we don't really use github - it's just a shared git repository with a reasonably nice interface for reviewing code. That's for my company, for many open source projects the damage would be more severe since they often heavily use issues and GH project management features.
- kelnos 5y agoThe problem is that many people (including companies) use GH for more than just source code hosting. Bug tracking, CI/CD, artifact hosting, etc. That's not something most companies can replace in a week, unless GitLab (or others) have a seamless import tool. And even then, if GH is down or your account has been suspended, what do you import from?
- majkinetor 5y agoDont be naive. People depend on github for packages. Literary countless projects would fail to build.
- neatze 5y agoI don't understand this; does DMCA imply no one can use algorithms in a book within open source projects ?
- meepmorp 5y agoThe book has code samples, and this repository is where they are hosted.
- projectileboy 5y agoThe story I don’t see being discussed is that Github is now engaged in the same kind of rotten behavior as YouTube.
- IshKebab 5y agoGitHub are legally required to do this.
- dragonwriter 5y agoNo, they aren’t. If there is no merit to a copyright claim, there is no vicarious liability for Github to be insulated from by the DMCA safe harbor. It’s cheaper and less risky not to evaluate the merits of DMCA notices and just to blindly execute them, but it is erroneous to say that thet are legally required to act in that manner.
- mon77 5y agoYes, they are, and yes, there is. The claiming and counterclaiming process is independent of the merit of copyright under DMCA. Budget and risk are not part of the equation. Evaluating copyright applicability cannot be part of the equation; it’s not your content. The law specifies almost exactly what you must do with some vague concepts for interpretation (like expedient). If you don’t do those things even on an obviously bogus complaint that is otherwise properly formed, you stop having safe harbor, because those disputes are intended by the law to be handled in court and not your legal department. It’s genuinely not complicated (read OCILLA) and you should be annoyed with the law for this situation, not its subjects. You are sharing an opinion disguised as fact. It’s a wrong one, but I get why you’d conclude it. Source: Write DMCA policy for UGC.
- dragonwriter 5y ago> The claiming and counterclaiming process is independent of the merit of copyright under DMCA. The DMCA process isn’t a mandatory process, it is a process to receive safe harbor from whatever liability would otherwise exist. If there would be no liability independent of the DMCA safe harbor, there is no mandate to follow the safe harbor process. (If you disagree with this, here’s what you need to do, identify—with citation to supporting law—the available legal remedy that can be imposed on a provider who declines to adhere to the DMCA safe harbor process where there is no underlying copyright liability.) Service providers make a choice to follow the notice process blindly as a risk management measure, not because it is legally required to do so. (Which is also why the counternotice process is often less fully implemented, or why, e.g., Youtube has its own hyper-aggressive policy for certain content that goes beyond the DMCA process and lacks counternotice opportunity—the counternotice process is just as much part of the DMCA process, but because providers are able to structure their relationship with users in a way which avoids having any liability for takedowns which would benefit from a safe harbor, they can safely ignore that process.) Source: the actual text of the DMCA safe harbor provision, and the reason it is called a “safe harbor”. > You are sharing an opinion disguised as fact No, you are engaging in standard industry blameshifting by misrepresenting risk management strategy adopted responding to incentives created by a law with an actual legal mandate.
- rurban 5y agoIt's the Sedgewick/Wayne Algorithms 4 in Java part. The author had an arrangement with the Pearson editor to release the accompanying code for the Java update under the GPL, but apparently someone else at Pearson's had a different idea. Let's see if he has it in writing.
- orsenthil 5y agoIs there a clone of these not in github? Please share it so that we can keep them in a different repositories. Risky to keep these exclusively in Github.
- sabujp 5y agolet me get this straight, this was gplv3 licensed code, allowed by pearson, and then pearson auto dmca's it? Looks like some algos need to be fixed on pearson's side, in any case here you go : https://algs4.cs.princeton.edu/home/ https://algs4.cs.princeton.edu/home/
- avipars 5y agoWayback machine has a version http://web.archive.org/web/20201016112929/https://github.com/kevin-wayne/algs4 http://web.archive.org/web/20201016112929/https://github.com...
- emayljames 5y agosource folder is not backed up
- hintymad 5y agoSpeaking of Robert Sedgewick, I deeply enjoyed his course Analysis of Algorithms, as Sedgewick beautifully showed the power of generating functions. Generating functions, or analytical combinatorics in general, make it order-of-magnitude easier to analyze algorithm complexities compared to what we learn in an introductory course. It's mind boggling to experience how higher-level constructs of math reveal simplicity and deep insights.
- dogman144 5y agoIt’s a great algo book. I learned a ton of math-oriented CS thinking and application without having a particularly strong math background.
- generationP 5y agoFor anyone looking for the code: archive.org doesn't have the source, but it has the list of forks ( https://web.archive.org/web/20200908104954/https://github.com/kevin-wayne/algs4/network/members https://web.archive.org/web/20200908104954/https://github.co... ). The forks themselves are still online. For instance, https://github.com/FGM-148/algs4 https://github.com/FGM-148/algs4 . As usual with forks, use caution. Then there is the website of the book https://algs4.cs.princeton.edu/code/ https://algs4.cs.princeton.edu/code/ . This might be a less useful format, though.
- StreamBright 5y agoDMCA needs a song like YMCA.
- bo1024 5y agoSo is whoever filed the notice going down for perjury?
- orsenthil 5y agoLooks like we have a mirror that is outside of github. https://git.ahlawat.com/sharad/algs4 https://git.ahlawat.com/sharad/algs4