6 ms·
DéjàVu: a map of code duplicates on GitHub
- lcfcjs2 9y agoHigh level of code reusability indicates a more mature and stable architecture, in my opinion. People shouldn't be reinventing the wheel with every new project.
- inetknght 9y agoNow predicting automatic software that looks at duplicated code, flags it for violating license agreements, and sues for money. Welcome to the future of copyright trolls.
- zbentley 9y agoWow, GitHub could save a lot of storage space if they dedup'd across projects/files explicitly, rather than storing Git repos, which is what I'm assuming they do. Even with a good deduping/compressing filesystem, the way git history is stored means that they're probably missing out on a ton of savings here. Eh, it's probably not worth the complexity/deviation from standard Git tooling.
- dfox 9y agoGithub uses their own storage backend which I believe shares objects across all of projects regardless of whether they are explicit forks or not.
- zbentley 9y agoThat's really neat. Is there any documentation/discussion available on that technology? It sounds like something that would be fascinating to learn about.
- Edmond 9y agoI am not sure if the parent's claim is true, ie that Github is storing objects and sharing them across forked repos. If they are, then it likely just a direct implementation of git the technology. you can see how git stores data here: https://git-scm.com/book/en/v2/Git-Internals-Git-Objects https://git-scm.com/book/en/v2/Git-Internals-Git-Objects
- aidos 9y agoNot an expert, but that would only work if there was a single repo, right?
- Edmond 9y agoNot an expert either :) There is the concept of submodules which allows for multiple repos while maintaining the checksum mechanics that allows sharing the same bit of information between branches and across commits: https://git-scm.com/book/en/v2/Git-Tools-Submodules https://git-scm.com/book/en/v2/Git-Tools-Submodules The trick is that git maintains an abstract file system (ie a graph) across commits. The graph consist of pointers to content without having to create a clone of the actual content for every new version of said graph....it gets a little dizzy to explain but it is really not too complicated :)
- allenz 9y agoNot necessarily. To share blobs, GitHub would need to replace the standard git filesystem blob storage with a distributed database of blobs. All Git repos would share this distributed database.
- deleted 9y ago[deleted]
- kunal88 9y agogood
- nihonium 9y agoIn order to prevent code duplication on a global scale, we need more frameworks, like leftpad. :sarc:
- hinkley 9y agoYou’re being sacastic, but sitting down and looking at what subject areas appear most frequently and talking about why would be useful for any language. Are they even getting it right or do they all have the same bugs? Are there no existing libraries? Are the downsides worse? Can we fix that? Should this functionality live in the core language (did we miss a feature).
- coding123 9y agoWhat really sucks is people committing node_modules, that's just plain wrong.
- bpicolo 9y ago> that's just plain wrong there was at least one point in time where that was the recommended strategy
- aalleavitch 9y agoI'm not sure I could understanding the reasoning behind that. Does it have to do with dependency versions or the assumption npm might not be available or what?
- bpicolo 9y agoThere were/are a few factors. NPM availability is definitely one - before caching, and without the overhead of running your own npm replica. It also didn't used to have things like lock files. Vendoring gives you a deterministic build and removes availability concerns. In that aspect, it's not the worst thing ever, mostly just leads to noisy diffs (and maybe c extension issues if your team works on a variety of OSs?) This is pretty much what the golang world does (though now there are some tools that do a better job).
- deleted 9y ago[deleted]
- yoz-y 9y agoBoth actually. Even with yarn and lock files, npm servers can (and eventually will) pull a rug from under you. I have already been in a situation where a dependency version that I was locked to was simply removed from the official registry. The right solution is to have your own registry or backup the archives of dependencies that you are using. I think it is better to commit the archives rather than the whole node_modules as it does not produce a mess.
- az0 9y agoVery interesting from a security perspective. So much potentially dangerous code copy-pasted and most of it is probably never updated too. I've personally found some C vulnerabilities in code that I easily found used in many projects by Googling the vulnerable line... Usually not so much to do about it too.
- jlarocco 9y agoTrying to frame this as a security problem is a stretch, IMO. My impression is that most public projects on GitHub are only of interest to the author, and maybe a small handful of people. I, for example, have over 100 non-forked public repos and, except for 3 or 4 projects, nobody even looks at most of them, much less clones them and uses them. Even the ~4 that do get attention, it's usually not because they're using the code itself - it's because they're doing something similar and want to see how I did it. On the other hand, I only have anecdotal evidence to back up that claim, so who knows.
- hawski 9y agoI always like to see how some API is used in real projects. Sadly GitHub search is mostly useless for this, because of the number of duplicates. Google code search was great. It even supported regexps. Then the was koders.com, now there's also something from ohloh and it's better than GitHub AFAIR. EDIT: ohloh became openhub and now the code search is discontinued. So there is the nonfunctional GitHub search and an open niche for other projects...
- stickydink 9y agoI've used searchcode.com for a while, I don't think it has regex though.
- sdesol 9y agoDisclaimer: I'm the founder of GitSense (https://gitsense.com https://gitsense.com) that indexes code and Git history. Indexing and retrieving code at scale is actually a really challenging problem due to the fact that there is a lot of code, on a lot of branches, in a lot forked repos. With GitSense, it doesn't even try to determine the authoritative source (repo/branch), since I personally think this is a lost cause, given current AI technology. With GitSense, everything is context driven, which is how you can reasonably remove duplication. To search, you have to define what branches/repositories to consider, which can a be a few to a few thousand. Note, once a search context has been defined, it can be reused, so this isn't something you have to create every time, if you want to search. I sort of envision a Yahoo type (the first incarnation) approach to searching for code. The basic idea is provide a curated search experience, where domain experts can share what they believe to be relevant branches to consider, for a given problem. Without some human intervention, I think duplication is a given and as you point out, can lead to useless results.
- j_s 9y agoThere are a few things that might be worth checking out: https://hn.algolia.com/?query=code%20search https://hn.algolia.com/?query=code%20search
- neurotrace 9y agoThis is very interesting. I would have liked to see the results for JavaScript when you ignore the node_modules folder. If that's going to count for code duplication then pip dependencies should be included as well. This should definitely be taken as a lesson though: JS needs a better deployment solution. That, or better education on the current solution(s).
- k__ 9y agoDo people check in their node_modules?!
- neurotrace 9y agoApparently some people do. They really shouldn't. I can only imagine this is in some places ignorance and others out of fear for another left-pad scenario.
- k__ 9y agoThen shouldn't that article account for it? I would have thought that JS has fewer dupes because of NPM
- neurotrace 9y agoThat's exactly what I'm saying. The author even states that the node_modules folder makes up 70% of the files in the JS section. Seems like a poor way to measure.
- jlangemeier 9y agoThey need a way better deployment method. Pip dependencies are usually just enumerated in a file (much like the json for NPM), but I think there's fundamental differences between how the Python Foundation and NPM Inc. handle their repositories. And if something isn't a nicely bundled wheel, I can still go out and install it (and any dependencies) the old fashioned way. With some of the dependency chains for various js modules, you're really forced to use a package manager of some sort for anything beyond your basics with minimal dependencies; or you'll be pulling your hair out and looking for that virgin goat to sacrifice. FWIW, and I know it's not much, I really don't use Node unless I have to (or javascript outside the basics, JQuery & LoDash for that matter); I was turned off from it when I was told to download and install Node via a copy-paste from their website of some short command-line wget script. That's shoddy at best; so the current state of affairs can be linked back to early practices. It's nice that Node has been cleaning up their act, but it's still kinda a crap fest; and now that is the standard that they've provided for their community. To wrap up this meandering train of thought. The paper actually addresses this nicely, because when 70% of the fluff and cruft in JS repos is node_modules, you end up with hidden dependencies, which is how things like left-pad happen. With pip I know exactly what all of my dependencies are (explicit dependencies); with Node, you're required to dig into the node_modules for every known dependency of that initial dependency (implicit dependencies).
- Tommakx 9y agoWould be more interesting to see an analysis of almost equal files - to detect reimplementations of the same thing
- hultner 9y agoWould love to see a follow up where we would see how much duplication existed if we controlled for common dependencies and autogenerated code in conjunction with data on how many repositories are fully cloned (i.e. all code is near identical to another repository).