16 ms·
Disabled at 22 million commits
- flusensieb 3y agohttps://web.archive.org/web/20230702211636/https://programming.dev/post/355447 https://web.archive.org/web/20230702211636/https://programmi...
- DocTomoe 3y ago[flagged]
- phs2501 3y agoWhat? That's not how Git works, even ignoring packfiles. (And yes I have no idea what Github's backend does, but I can't imagine "duplicate all data for every commit" would be a feasible implementation strategy.)
- olavfosse 3y agoWhile it's correct that each Git commit is a snapshot of the worktree, and not a diff, Git uses copy-on-write / structural sharing, which means that generally speaking adding a commit to your repository is a very cheap operation storage wise. Recommended reading if you're interested in Git's data model, it's pretty easy to understand compared to Git's UI: https://git-scm.com/book/en/v2/Git-Internals-Git-Objects https://git-scm.com/book/en/v2/Git-Internals-Git-Objects
- littlestymaar 3y agoThat's my favorite fun fact about git: the diff you see isn't really a diff under the hood, but in fact it kind of is if you look one more layer under.
- cod1r 3y agohopefully this is a sarcastic comment. T_T
- Hello71 3y agoas pointed out in the sibling comments, that's not right, but also it can't possibly be right: torvalds/linux has about 1.2M commits, and according to https://cdn.kernel.org/pub/linux/kernel/v6.x/ https://cdn.kernel.org/pub/linux/kernel/v6.x/, the current Linux kernel source size is 131M after compression, so (assuming the kernel size has grown linearly with each commit) each maintainer would need to have 75 TB of storage on their personal machine?
- xen0 3y agoThat's not how git works. Unless every commit changes every single file in the repo (without ever reverting to a previous state). Most repos don't work that way. 22 million commits is a lot though; the author states that the repo was over 8GB "last they checked", and the entire thing looked to be an experiment to determine "Github's or git's breaking point" - I guess they found it.
- deleted 3y ago[deleted]
- orf 3y agoMany people have mentioned that this is incorrect, but I wanted to ask why you thought this? The Linux kernel repo contains 1.1 million commits and 80k files on master. That would mean a naive total of 88,000,000,000 files being stored as a “full copy of the repository”. Does this pass the smell test?
- orf 3y agoMany people have mentioned that this is incorrect, but I wanted to ask why you thought this, and why you thought it so confidently? The Linux kernel repo contains 1.1 million commits and 80k files on master. That would mean a naive total of 88,000,000,000 files being stored as a “full copy of the repository”. Does this pass the smell test?
- tedunangst 3y ago"Git stores a copy of the repo" is not an entirely unreasonable inference to draw from "git does not store deltas".
- orf 3y agoI guess that’s my question: is it? Does that inference pass any kind of smell test with basically any non-toy repo?
- tedunangst 3y agoI think one can (and should) look at some real repos to conclude it's not that simple, but if you're simply told it doesn't store diffs, what else would you think it stores?
- orf 3y agoI would think that either it stores a volume of data that would make it impossible to use… or it it doesn’t. And if it isn’t impossible to use then I wouldn’t post about how it is.
- renewiltord 3y ago[flagged]
- asadawadia 3y agoi remember reading that too https://willi.am/blog/2014/10/14/for-the-last-time-git-stores-snapshots-not-diffs/ https://willi.am/blog/2014/10/14/for-the-last-time-git-store...
- mtone 3y agoAssuming it's correct, I think this answer explains it well: https://stackoverflow.com/a/31996121/283879 https://stackoverflow.com/a/31996121/283879 Basically, yes it may write snapshots per-file (never per-commit) locally but there is a separate routine to transparently repack the whole thing with deltas.
- iBotPeaches 3y ago> I decided to see how many commits GitHub (and git) could take before acting kind of wonky. At ~19 million commits (and counting) to master: it’s wonky. This just doesn't seem right to me. Why? Its obvious at some point you'll harm the service. If the goal was to test it, why not try locally with git.
- xigency 3y agoA good lesson to learn - If you as a service owner aren’t testing the limits to the point of failure and enforcing sensible guardrails around that, then some random user eventually will.
- lucb1e 3y ago> why not try locally with git. Because you can't. GitHub is not open source, you'd need to steal the source code to try it locally. This comment is for educational purposes only, not trying to give OP ideas!!1 But you're right in spirit of course. Would be more interesting to install Forgejo/Gitea, GitLab, GitWeb, gitolite, TortoiseGit, etc., test them on various limits, and write that up in a nice blog post for magic internet points.
- ronsor 3y agoYou can download GitHub Enterprise Server for free.
- js2 3y ago> "GitHub (and git)" The "(and git)" portion can of course be tested locally. What OP will find out is that there is no more inherent limit on the number of commits in a repo than there is an inherent limit in the number of nodes in a linked list. You can go on forever till you run out of disk space. Possibly repacking will eventually require more than available memory.
- RexM 3y agogit runs outside of GitHub, which is what the comment you responded to was saying. Test the behavior of git locally, without testing GitHub.
- throwuxiytayq 3y agoThis feels slightly malicious, but I can’t help but admire the curiosity that takes someone to actually see what happens if. That said, now we know, so nobody else needs to bother GitHub engineers by doing this again, hopefully.
- lucb1e 3y agoNot like it would take a lot of code to check on push if the repository has more than 10k commits per day since its creation date or something, to stop such abuse. Doesn't thwart existing repositories with millions of commits (Linux is at ~2M) and gives time to formulate a long-term plan for what's allowed and what's paid or just disallowed. So even if people were to try, I don't see that being a big bother. Not that it's not malicious to do this now
- thewataccount 3y agoI don't think you can limit pure commit counts though, because you can push many commits/massive history changes in one go. Monorepo's in particular could be impacted
- lucb1e 3y agoGood point, perhaps (age_in_years+1)×1M would be a better limit. Anyone wanting to import more than 1M commits could get a paid tier or beg support. At any rate, not that hard to implement is what I would expect
- mwint 3y agoGit commit timestamps are 100% fudgeable. You could implement this based on GitHub repository age, but the assumptions would break for imported repos. (Understanding we’re waaaay off in edge case territory here and this is all basically academic.)
- MBCook 3y agoSo the author was purposefully trying to do the most extreme thing they could to see how git/GitHub act/break. I don’t blame GH at all. Source: https://web.archive.org/web/20230702215522/https://sh.itjust.works/post/580838 https://web.archive.org/web/20230702215522/https://sh.itjust...
- iLoveOncall 3y ago> I don’t blame GH at all. I don't really see anyone blaming GitHub, not even the original post, I'm not sure why all the responses here are insinuating that?
- housemusicfan 3y agoSo basically this: https://www.youtube.com/watch?v=1kzb6uf0U0k https://www.youtube.com/watch?v=1kzb6uf0U0k
- lucb1e 3y agoVideo showing Simpsons going to an 'all you can eat' and staying until after closing time still gobbling down food without end, owner has homer thrown out.
- deleted 3y ago[deleted]
- aleph_minus_one 3y agoSee also https://www.youtube.com/watch?v=Q6g8x0CPl2A https://www.youtube.com/watch?v=Q6g8x0CPl2A
- pragmatick 3y agoThis is the most blatant case of false advertising since my suit against the movie The Neverending Story.
- hayd 3y agoAnd used Github actions to do the (infinite loop?) compute.
- ranting-moth 3y agoMore correct title would be "GitHub stopped abuse after 22M commits". There is absolutely nothing wrong with GH stopping that and it's very wrong to insinuate otherwise like OP is doing. Wouldn't be surprised if GH would permaban him.
- deleted 3y ago[deleted]
- eyelidlessness 3y agoI don’t think the author is trying to insinuate that GitHub is in the wrong in any way. They explicitly say they understand the decision, and anticipated that it would happen. I don’t want to quibble with the term “abuse”, because I think in this scenario it depends on whether intent is a factor and whether we should trust their stated intent. But depending on how you look at it, GitHub would be just as likely to benefit from hiring the author as they would from banning.
- rat9988 3y agoIt is malicious as he knows he will harm the service to be able to draw whatever conclusion. This is not a case where the end justifies the means.
- eyelidlessness 3y agoThe first time I wrote and shared any kind of interactive code, it took approximately five minutes for someone to XSS it. At the time, I was pretty miffed too. After a polite explanation that the “abuse” was curiosity about defensive measures I’d taken, I understood pretty suddenly that there was a whole scope of programming I hadn’t even considered. More than 20 years later, I still remember the enormous benefit that little bit of malice has bestowed on me and my career. And every time I’ve been on the receiving end of such an exploratory exploit since has been exponentially more appreciated.
- eyelidlessness 3y ago
- Kwpolska 3y ago> I’ve also asked if they can re-enable it so I can give one more commit to say the final results on the readme then (public) archive it. Entitled much? The author should be happy GitHub didn't just ban them for violating the ToS and intentionally trying to break things.
- eyelidlessness 3y agoThey asked. They didn’t demand, and they seem prepared to accept whatever GitHub decides. If I were fielding that request, I’d certainly grant it—on the condition that any deviation from the stated intent would indeed result in a ban—purely on the basis that it’s a ~free QA contribution and postmortem.
- Kwpolska 3y agoKeeping the repository, even as a public archive, would still require a lot of resources on GitHub's side. The only fair thing to do here would be to apologize and ask for the repo to be deleted.
- rafark 3y agoAnd could be seen as a reward or an encouragement for other people to abuse the service.
- eyelidlessness 3y agoThey already do incentivize white hat exploit efforts[1]. The author seems to have run afoul of one of their rules[2] by impacting other users, but I don’t think that impact could be knowable without trying. GitHub could trivially honor the request without changing the incentives or even taking any defensive implementation action, by specifically citing this experiment in the rules and maybe adding some more specific wording to the TOS. 1: https://bounty.github.com/ https://bounty.github.com/ 2: https://bounty.github.com/#rules https://bounty.github.com/#rules
- chris_wot 3y agoSo basically, this guy is trying a DoS of GitHub. To hell with him.
- blowski 3y agoI’m surprised at the reaction in these comments. Somebody curiously pushing the limits of a service to see what would happen is very much in the spirit of all hackers. Meanwhile, GitHub responded appropriately, and his write up agrees.
- Mystery-Machine 3y agoSomeone potentially taking the service down for everyone, you know, just out of curiosity. Which part of this curiosity you need GitHub for? I'm curious how well GitHub handles DDoS attacks, what's their limit. Let's DDoS and find out, it will be fun!
- SparkyMcUnicorn 3y ago> Someone potentially taking the service down for everyone, you know, just out of curiosity. I think this is exactly why it's great, and it's basically turned into a GitHub advertisement. Either GitHub is simply unable to handle weird abuse methods and/or the abuse prevention is improved. As an enterprise, wouldn't it be a bit concerning if your git host was unable to function (or respond appropriately) when presented with a random script kiddie? This person didn't have bad intentions, but other people out there most definitely do.
- Mystery-Machine 3y agoGitHub is very much able to handle one person doing this. Doesn't matter if you had bad intentions or you were just ignorant to bad side effects.
- ryaneager 3y agoYou really think one lone repo could take down all of GitHub? If GitHub doesn’t have stops in place to prevent that then they honestly deserve it.
- Mystery-Machine 3y agoSo doing a DoS attacks from a single machine is fine, because "your servers can handle that"? Really? Of course GitHub can handle this, but if the sole purpose is to see where's the limit, you're stressing our servers and wasting our resources for nothing. I'd ban you no questions asked. Go test perf/scalability issues on someone else's live site.
- TacticalCoder 3y agoDid the dude bring GH down at some point?
- deleted 3y ago[deleted]
- BeefySwain 3y agoSidestepping all of the ethical questions of embarking on this "research", I'm surprised the number was that low. Linux[0] itself has about 1.2 million commits, so apparently Linux is within an order of magnitude of bringing GitHub to it's knees? [0] https://github.com/torvalds/linux https://github.com/torvalds/linux
- tikhonj 3y agoThere's a rough rule of thumb that you should expect to redesign your system to handle each order of magnitude increase in scale, and I figure it applies here too—gracefully handling that size of repo would require substantial engineering work, and they have plenty of time to handle it before human-oriented open source repos get even close to the current limit.
- lucb1e 3y agoI'm not sure redesigns were necessary between going 1 to 10, from 10 to 100, from 100 to 1000, from 1000 to 10'000, from 10'000 to 100'000, or from 100'000 to 1000'000 which we're now at. It sounds like a sensible engineering rule, but I'm not sure it translates to software, or at least not in this case. I don't know of any design changes made to Git since it was first created, there's no v1 and v2 repositories for example.
- aloer 3y ago> there's no v1 and v2 repositories for example We wouldn’t know. GitHub is probably running something very different to normal local git including optimizations for performance and cost. They must only ensure API/protocol compatibility and could have already replaced everything else many times over.
- spiralx 3y agoIt depends on how quickly you pass through each order of magnitude milestone. I remember reading about how MySpace grew something like five orders of magnitude in less than a year, and no matter how scalable your architecture is you're going to hit a point during that where you need to rearchitect your whole system. Slower growth allows for forward planning and incremental architectural changes.
- deleted 3y ago[deleted]
- medellin 3y agoThe mindset around programming and exploration in general is in a sad state. I don’t understand why we have so much hate here for things like this. Better that someone like this find it then someone who noticies it and spins up 1000s of repos to do the exact same thing. I think the sentiment here shows the current state that software engineering has devolved into. It’s a 9-5 where you put in minimal work and get mad when someone breaks your system because you might have to do an hour of work to fix it on your weekend.
- Arainach 3y ago"Devolved" implies a negative connotation, but this is a positive evolution.
- mabbo 3y agoThe author used up so much of github's resources that it impacted other users. 22 million commits is probably enough that something started to hit a linear or n-log-n scaling function, setting off an alarm on some metric. Yeah, you get in trouble for that. I'm reminded of a time in high school where my friend almost got himself banned from the school computers. At home he had dial-up internet (it was 2003 and he lived in a very rural area). But at school he had megabits of bandwidth he could (ab)use. So he started pirating everything on the internet using a computer nobody ever used in a side-room of the library. It ran 24/7 downloading his long list of desires: games, movies, tv series, etc. He stored his spoils on his network drive, which had no limits on how much it could hold (until he got caught). He'd occasionally bring in a hard drive, copy everything that fit on it and bring it home with him on the school bus. But all good things must end. The network admin for the school board eventually came by and sat him down. He showed my friend a pie chart where, as he described it to me, "my name was on the portion that took up more than 2/3 of the pie". After a conversation, all the data got deleted, my friend got a stern warning, and somehow didn't get into any worse trouble than that.
- Panzer04 3y ago"somehow" I don't get this attitude. Shit happens, we talk about it, we don't do it again. Not everything needs to have dire consequences.
- flutas 3y agoI think he means "somehow" in the meaning of "somehow, none of the copyright holders asked the school for his information."
- justinclift 3y agoSounds like the network admin and surrounding people had their heads screwed on properly. :)
- layer8 3y ago> The author used up so much of github's resources that it impacted other users. Note that the message only said “the potential to affect other users”. I would expect a professional service to catch such things before it actually affects other users.
- GMoromisato 3y agoA long time ago, the math column in Scientific American decided to run a contest. It asked readers to send a post card with the biggest number they could think of. Whoever came up with the biggest number would win $1 million--divided by the winning number. The editor of the magazine almost stopped the contest because he worried that someone might actually win real money and the magazine would be on the hook. But the author reassured him: human nature being what it is, the winning number is going to be not only larger than 1 million, but much larger than you can imagine. And so it was. The winning number was (IIRC) some tower of exponentials that would take most of the universe to write out as decimal digits. The SciAm budget was safe. If readers had coordinated somehow, they could have won a million dollars from SciAm and divided it among themselves. They might have made a hundred dollars each. But the author knew that such coordination would be impossible. Human nature would not allow it. Someone, somewhere, was going to send in a ridiculously large number to win. Classic Prisoner's Dilemma. The GitHub case is the same. Human nature being what it is, someone, somewhere is always going to try to push the limits. As the developer of a SaaS development platform, this is something I'm taking to heart.
- LouisSayers 3y agoThe biggest number I can think of is 0.001 :D They could have been in quite some trouble!
- pritambaral 3y agoIn the famous words of Calvin Coolidge, "you lose". https://clintonwhitehouse3.archives.gov/WH/glimpse/presidents/html/cc30.html https://clintonwhitehouse3.archives.gov/WH/glimpse/president...
- antimora 3y agoHere is another GH abuser I found recently: https://github.com/eemailme https://github.com/eemailme This account basically subscribes to thousands of repositories and monitors all activities. I am suspecting this account is harvesting user activities. I am not sure why GitHub allows this type of data harvesting.
- masterugway 3y ago[flagged]
- stiwari 3y agoI'm surprised that a lot of users here are telling OP that he was wrong. OP was well within his rights to do this, as his intention was to stop when any impact is observed, not continue with it. It is within their rights to test the system they want to use to make sure their requirements are met. To be honest, this is why companies also should not discourage this. Imagine if a malicious group did it with multiple users at the same time. At least now they will have pro active alarms for it.
- kjs3 3y agoWhen I was an junior admin in college, there was always at least one kid a trimester who 'experimented' with a fork-bomb on one of the shared Unix servers, and was shocked to learn that there are things you can do that you really shouldn't do. Same thing.