13 ms·
Operation Rosehub – patching thousands of open-source projects
- markcerqueira 10y ago"Patches were sent to many projects, avoiding threats to public security for years to come." Are these pull requests that the project would still need to approve/merge or were they just pushed in?
- edutechnion 10y agoThey were PRs that required approval and merge from the Github project maintainers. Here is a search to see some of their work: https://github.com/search?q=%22Upgrade+Apache+Commons+Collections%22&type=Issues&utf8=%E2%9C%93 https://github.com/search?q=%22Upgrade+Apache+Commons+Collec...
- deleted 10y ago[deleted]
- fudged71 10y agoIt's actually incredibly interesting to read how the developers individually responded to each of these PRs. It would have been great to see a count of how many PRs have been accepted.
- cpeterso 10y agoedutechnion's link says 1108 open PRs and 999 closed. Interesting that 2100 of the PRs are "Upgrade Apache Commons Collections to v3.2.2" and just 7 were "Upgrade Apache Commons Collections to v4.1".
- therealdrag0 10y agoProbably v3.2.2 was lower hanging fruit for most projects. Instead of having to make code changes.
- unityByFreedom 10y agoI'm sure they're requests.. They'd need to be deployed to production too. Still pretty awesome.
- hinkley 10y agoFor me, if a project had a bunch of open PRs for security issues, it would discourage me from using it for new work. It would also help break a tie in my head between keeping an old library and replacing it with something that has legs. So even if they don't get merged, they still serve a purpose, even if it's just for a few people who behave like I do.
- fhoffa 10y agoThis is one of the most impactful projects I've seen built using the GitHub source on BigQuery dataset (since we published it). If you want to see other use cases - I've collected plenty of other stories from multiple parties at: - https://medium.com/google-cloud/github-on-bigquery-analyze-all-the-code-b3576fd2b150 https://medium.com/google-cloud/github-on-bigquery-analyze-a... Disclosure: I'm Felipe Hoffa and I work for Google Cloud (https://twitter.com/felipehoffa https://twitter.com/felipehoffa)
- fhoffa 10y ago<meta> Title change by mods -- I submitted this post as "Googlers used BigQuery and GitHub to patch thousands of vulnerable projects". After it got to #1 on the front page, mods silently changed the title to "Operation Rosehub – patching thousands of open-source projects" I wish HN had a more transparent way to show that the mods changed a title and why. Since HN does not, the least I can do is add this info for transparency. (related https://news.ycombinator.com/item?id=6572466 https://news.ycombinator.com/item?id=6572466 https://news.ycombinator.com/item?id=4102013 https://news.ycombinator.com/item?id=4102013) </meta>
- mikekchar 10y agoI got a bit confused by your post. I gather that you are simply recording the fact that the title was changed in accordance with the posting guidelines, not that you are complaining about it. Just in case other people were similarly confused...
- ec109685 10y agoI think he wishes that here was an "edited" annotation next to changed titles.
- rak00n 10y agoIMHO the title coming from mods sounds better. Your title seemed like you're explicitly advertising for BigQuery.
- jayfk 10y agoI've built something like this for Python projects. You add your repo and a bot is constantly checking for insecure and/or outdated packages and sends you a pull request if you need to update. It's free for open source projects at https://pyup.io https://pyup.io
- po 10y agoWe've been using pyup for a few months now and it's awesome. Thank you. I was initially concerned that constant PR's for dependencies would be too noisy for our team but it turned out that it's configurable enough and gracefully handles us ignoring or closing PR's that we evaluate and decide to wait on. It's a great service that all python developers should be using. (I say that because I want as many pyup users as possible so it never goes away)
- nathancahill 10y agoSimilar to Greenkeeper.io for npm?
- luhn 10y agoBookmarking this. I recently was alerted to a CRLF injection vulnerability on the service I manage because an outdated dependency was vulnerable. This could have nipped that in the bud.
- lumpypua 10y agoJust signed up and discovered a vulnerability. Thanks! A quick suggestion, consider adding a full sample configuration along with your config docs: https://pyup.io/docs/configuration/ https://pyup.io/docs/configuration/ It's a lot easier to see how it all sits together with a sample. CircleCI has a great example: https://circleci.com/docs/1.0/config-sample/ https://circleci.com/docs/1.0/config-sample/
- notheguyouthink 10y agoIs there something like this for all languages / os / etc? Seems like an amazing tool to poll against.
- nkuttler 10y ago
- codelion 10y agowe have been doing thus for a while now : https://www.sourceclear.com/blog/millions-of-program-builds-vulnerable-to-man-in-the-middle-attacks/ https://www.sourceclear.com/blog/millions-of-program-builds-...
- hcs 10y agoGot a 404 on the above, looks like it should be: https://www.sourceclear.com/blog/millions-of-program-builds-vulnerable/ https://www.sourceclear.com/blog/millions-of-program-builds-... But I don't see where it discusses sending PRs to affected repos, only detecting them.
- codelion 10y agoAh yeah I had the old link, thanks for fixing. Actually we privately disclosure the problem to the developers and get it fixed following responsible disclosure and not post PRs directly.
- orf 10y agoIn their query they do: FROM (SELECT id,content FROM (SELECT id,content FROM [bigquery-public-data:github_repos.contents] WHERE NOT binary) WHERE content CONTAINS 'commons-collections<') Why the subquery? Why not WHERE NOT binary AND content CONTAINS...? is this a bigquery thing?
- deleted 10y ago[deleted]
- fhoffa 10y agoGood catch. It seems like an artifact of working and editing a query until they got what they wanted - certainly not a BigQuery restriction. (like this https://www.youtube.com/watch?v=cO1a1Ek-HD0 https://www.youtube.com/watch?v=cO1a1Ek-HD0) Note that the published query scans 2.25 TB of data. While impressive, for a better workflow and cost management I would split it into a 2 step process: - First extract all the files I'm interested in to a separate table (all pom.xmls?). - Then run whatever analysis you want over those files.
- vgt 10y agoThis appears to be "legacy SQL" in BigQuery, which did not have query optimization - entirely rule-based query planning. The query is a little inefficient indeed. BigQuery has since released ANSI 2011 "standard SQL", which would does have an optimizer and would push predicates down). (work on GCP and worked on BQ until recently)
- contravariant 10y agoMaybe they want to make sure Binary data is filtered out first? Not exactly sure how rigid the evaluation order is in SQL or BigQuery.
- havermeyer 10y agoIt should work fine either way. Using standard SQL in BigQuery, though, you can do: #standardSQL SELECT pop, repo_name, path FROM ( SELECT id, repo_name, path FROM `bigquery-public-data.github_repos.files` AS files WHERE path LIKE '%pom.xml' AND EXISTS ( SELECT 1 FROM `bigquery-public-data.github_repos.contents` WHERE NOT binary AND content LIKE '%commons-collections<%' AND content LIKE '%>3.2.1<%' AND id = files.id ) ) JOIN ( SELECT difference.new_sha1 AS id, ARRAY_LENGTH(repo_name) AS pop FROM `bigquery-public-data.github_repos.commits` CROSS JOIN UNNEST(difference) AS difference ) USING (id) ORDER BY pop DESC; Better yet, it runs faster than the legacy SQL query :) As a disclosure, I work on the project to support standard SQL in BigQuery.
- rrggrr 10y agoSo many questions... What does this say about Google's hiring, about its employee's values, about values across the tech community? I can remember a time when managements would have shut this down, when employees would have said, "not my problem", when entire industries would have buried their heads in the sand. Is it the lack of liability and regulation that clears the way for this kind of corporate citizenship? Is it cultural?
- ISL 10y agoIt may be, in part, Google's giant cash machine. It makes it possible to be altruistic. In a world of fierce/commoditized competition it is much harder to expend resources on 'side' projects. Google also thrives in a healthy internet.
- lloydde 10y ago"They were happy to see employees spontaneously self-organizing to put their 20% time to good use." Reminded me that I heard that Google's 20% time had generally like 10% participation and has had a number of conditions including manager approval for the last 4 years. Is this article breathing new life into the myth of 20% time or does this reflect the revival of the process? Either way, incredible accomplishment on the patch army!
- nostrademons 10y ago20% time has always meant different things depending on who you talk to. It's likely that you're just talking to different people. I took copious amounts of 20% time in my time at Google (2009-2014). Usually I'd start a 20% project with neither my manager's knowledge nor his approval; if it looked like it had legs, I'd let him know about it and ask what he thought. I never had a manager outright forbid me from working on a 20% project; responses ranged from "You should consider this your main project now; it's critically important that we understand this area" (along with a spot bonus for delivering on it) to "Well, you can work on it, but you are unlikely to get credit for it come promo time." In general, as long as I got my work done, my managers didn't care what else I was working on.
- tlrobinson 10y agoWow. I wonder how much a query that searches the content of all of Github costs (if you're not Google). This page says the dataset is 3TB+ https://cloud.google.com/bigquery/public-data/github https://cloud.google.com/bigquery/public-data/github and presumably most of that is content.
- zellyn 10y agoYou can do it yourself for free, I believe. https://github.com/blog/2298-github-data-ready-for-you-to-explore-with-bigquery https://github.com/blog/2298-github-data-ready-for-you-to-ex...
- Veratyr 10y agoThe contents table [0] which they ran the query on is 1.8TB. Assuming you only need to do a single pass (seems reasonable given that it's a simple regex), the price should be about $9 [1]. Free quota covers 1TB so the remainder would be $4. [0]: https://bigquery.cloud.google.com/table/bigquery-public-data:github_repos.contents?pli=1&tab=details https://bigquery.cloud.google.com/table/bigquery-public-data... [1]: https://cloud.google.com/bigquery/pricing https://cloud.google.com/bigquery/pricing
- skeletonjelly 10y agoThe data is stored in BigQuery and updated from GitHub weekly. I imagine they do a date range filter when they updated from GitHub
- deleted 10y ago[deleted]
- lolive 10y agoWouldn't a graph database be a more suitable tool for that kind of task?
- CydeWeys 10y agoWhy would it be a more suitable tool? What can a graph database tool do that BigQuery lacks? If you've already got the data conveniently preloaded into a SQL database for you, and all you need is a very simple SELECT statement with two WHERE clauses ... why would you use anything else? Spinning up an entire graph database unnecessarily seems like over-engineering.
- jart 10y agoAuthor here. There's some truth to what he's saying. One thing I've been meaning to do is get my hands on all the Maven pom.xml files that exist, so I can load them into a Guava Multimap (my graph database of choice) and figure out every single artifact that will transitively inherit vulnerable collections on the class path.
- saurik 10y agoI am extremely sad that this turns into an argument for making certain that all source code in the world is at least indirectly accessible specifically via GitHub (at which point people will find it there and expect the developers to respond and generally track everything going on there, even projects which are much happier using more open tools); like: it isn't sufficient that your code is "open", it actively has to be part of the unified GitHub empire.
- jevinskie 10y agoYour gitweb [0] has always worked perfectly fine for me! I agree, I don't see the need for Github. [0]: http://gitweb.saurik.com http://gitweb.saurik.com
- bla2 10y agoReally cool, kudos to people helping with this. I wonder if this could have been done in a way that non-Googlers could have pitched in too, given that this is for a public good -- but it's tricky with security issues.
- ploxiln 10y agoI think this is one good concrete example of why the npm style of private dependencies for each lib is not the greatest thing ever, while the non-recursive style in python (or C) is overall more manageable (if you are actually managing your dependencies instead of ignoring them).
- make3 10y agoI wish you could do the same thing with mental illness.. massively send pull request to correct everyone's bad brain code.. <sorry>
- nrdwavexe 10y agowell people have to learn to do things for themselves. very difficult when it takes a lot of work and, frankly, pain
- i336_ 10y agoThis is a genuinely cool idea. Seriously. It raises a lot of questions about what sort of transformative spectrum (excuse pun) would be applied here though. It's is incredibly abstract as presented. But even at the abstract level, the one thing I know would absolutely happen for sure is that the fixes that made the biggest difference would be hand-waved out of existence by infecting them with viruses, creating scare-campaigns, etc. Source: I've learned a lot about Big Pharma over the past 10 years as I've quietly found real solutions to my own mental health issues. I'm sadly too scared to share what I've found and I keep seeing products disappear off the market or suddenly attract customs/overseas shipping issues. Suffice it to say that the medical industry is opposed to anything they can't patent - and that, as an industry, it must ensure its own survival. Interpret that any way you see fit.
- tombh 10y agoIs https://libraries.io https://libraries.io not a more comprehensive and community-focused response to the same problem? libraries.io did make it to the front page a few months ago, but I think its underlying vision might not have been driven home from just glancing at its home page. It supports 33 package managers (not just Java, though I'm sure Rosehub doesn't just do that either) and Github/Gitlab/Bitbucket, not just Github. And it provides both email notifications and auto PRs. But that's just the overlap with Rosehub. On top of that it offers the means to discover libraries based on a Dependency Rank (think Page Rank but using dependencies instead of hyperlinks). Which in turn allows it to surface projects with a high "Bus Factor" -- projects maintained by few committers, but depended on by many (so they'd be more affected by said committers getting run over by a bus). AND it mines the licenses for a project, notifying if any of the dependent licenses are incompatible with the parent license. What's more it's a non-profit organisation receiving enough funding to employ 2 full time devs. I think libraries.io is Rosehub and more, to quote the about page; Our goal is to raise the quality of all software, by raising the quality and frequency of contributions to free and open source software; the services, frameworks, plugins and tools we collectively refer to as libraries. To take the liberty of extrapolating from the libraries.io vision: open source security isn't just about fixing patches, but about supporting the environment, people, conditions and tools that contribute to open source software.
- gpawl 10y agoI see nothing on the libraries.io website that explains how it would be used to solve the problem described in the OP.
- tombh 10y agoOP here, yeah I agree, like I said the home page could be more explicit. Here's some links: https://libraries.io/about https://libraries.io/about https://libraries.io/bus-factor https://libraries.io/bus-factor https://github.com/librariesio/lib2issues https://github.com/librariesio/lib2issues
- snambi 10y agoWhat is in it for google?
- idlewords 10y agoEmployee satisfaction.
- theDoug 10y agoAnd a healthier internet
- CydeWeys 10y agoDisclaimer: I work for Google and Justine (this blog post's author) was an immediate coworker at the time of Operation Rosehub. I think you're asking the wrong question. This wasn't some top-down directive from some VP trying to come up with ways to make Google look good in the open source community. This was a bottom-up effort, that happened simply because Justine wanted to do it, and the easiest way to get it done was to recruit other like-minded engineers to help her rather than having to do it all by herself. She would have done it at any other company that allowed her (though, knowing her, she would've done it regardless). As engineers we have agency. The decisions that I make in my day-to-day work are mine, not my employer's. I can directly impact and affect lots of things, and the only motive you need inquire about to explain it is mine.
- luhn 10y agoAs scary as Google's massive size and power is, it's pretty awesome that they're incentivized to do things like this to help the internet because they are the internet.
- nrdwavexe 10y agowhat a disaster. learning all of the wrong lessons. it sends the message that the infrastructure isn't secure, which is true, but the response is to ignore the opportunity to develop something that provides a reasonable proof that your system is secure.
- tropo 10y agoIf I understand it right, this bug involves code pulling in old buggy libraries, sometimes indirectly via other libraries. It seems that there is a reference to a specific bad version, not the actual inclusion of cut-and-paste code. Eh, why not just get rid of the bad version? Alternately, release a bug-fixed copy with the same version number. Any breakage is a case of "oh well, you're safe now". Leaving the security hole is probably worse breakage.
- richardwhiuk 10y agoIn theory someone could be relying on the bug, and not be vulnerable - e.g. if they don't allow any external access to the system. If you just re-publish the old version, it's difficult to know whether you've taken the change. If you are going to reissue the same version number - why bother having version numbers at all?
- benmmurphy 10y agoThe bug is not including this library. This library is 100% secure [i mean if this is a vuln in this library then a large proportion of libraries are insecure because they could be leaked to untrusted code and used to break the JVM trust model]. It just so happens this library used to contain a really-really useful gadget for exploiting another security problem. However, removing this gadget doesn't mean the security problem is fixed. There are other fun libraries. In fact classes similar to the Mad Gadget have been used in the JDK to escape the sandbox in the past. Yes, stuff like this exists or has existed in the JDK [https://github.com/jenkinsci/jenkins/blob/96a9fba82b850267506e50e11f56f05359fa5594/test/src/test/java/jenkins/security/security218/ysoserial/payloads/Jdk7u21.java https://github.com/jenkinsci/jenkins/blob/96a9fba82b85026750...]. And this work is very useful in so far as I'm sure the benefits it provides is going to massively outweigh the cost. However, if you have a naked ObjectInputStream#readObject in your code then you probably still have an exploitable security issue. Have a look at how well Jenkins strategy was to fixing this issue which was basically the same strategy as Operation Roshub. ie: removing the ability to access classes that were known to be used in gadget chains. Surprise, surprise it didn't last very long and people just found new gadgets. And if you read this blog post then you might be mistaken into thinking that removing commons-collections from your classpath or upgrading commons-collection to the 'safe' version would make object deserialization safe but this is not the case. if you have a naked ObjectInputStream#read in your code then you are vulnerable to remote code execution.
- cypherpunks01 10y agoNice! That's some good citizenry. Interesting fact: Justine was the founder of occupywallst.org, which was the highest-trafficked publisher/web hub for the Occupy Wall Street movement before she worked for Google.
- vog 10y agoI like the "bank teller" analogy used in the article. > it would be like hiring a bank teller who was trained to hand over all the money in the vault if asked to do so politely, and then entrusting that teller with the key. The only thing that would keep a bank safe in such a circumstance is that most people wouldn’t consider asking such a question. This does not only work for deserialization issues. It is a great analogy for a huge class of IT security issues! Maybe we should use that one when communicating with the media. This this works much better than the usual burglary analogy. I like how it points out that this is about stupid and/or malicious behaviour (code), where the attacker (hacker) just needs curiosity, and may find this out even by accident. The attacker did not have to break something, and did not damage anything, to get into something. In particular, this makes clear that this is caused by irresponsibile behaviour of the organization and/or other entities to whom they delegate trust. Even for more complicated scenarios, I like the bank teller analogy more than the classic burglary analogy. In that case, the attacker observes multuple bank tellers, and notices e.g. that if you ask the first teller for form A and put in certain words, another bank teller will accept it and give you a stamped form B, which you can show to a third teller in another branch office who will look a bit confused, but finally accept it and hand over all money to you. We need to get over blaming the messengers[1], buying zerodays and declaring cyberwar. What we really need to do is to finally make our[2] computer systems secure and trustworthy, at least up to a certain minimum-level of sanity: no exec, no injection (i.e. typing/tagging), no overflows (i.e. static analysis), input validation, testing, fuzzing, you name it. And this cannot work by just adding more and more complex security measures outside, but more importantly simplifying and cleaning up inside. Although rewriting software from scratch is very risky, radical refactoring is not! And every good software engineering course tells you how to do it correctly. [1] security researchers, but also "amateur" hackers, or just someone running into it by accident because the security issue became so large it finally had to be noticed by someone. [2] in the sense of: everyones!
- reacweb 10y agoI know it is not correct to ad "me too" comments here, but I went here for the same quote as you. It is the best quote I have never seen about security because it does not depict a evil hacker that breaks a not secure enough wall. It depicts a clever client that goes through a stupid security hole. That's a better analogy for 99% of security hacks.
- joelthelion 10y ago> But unlike big businesses, open source projects don’t have people on staff To read that from Google is frankly disappointing. While this is true of many open-source projects, it doesn't have to be that way. Red Hat (and Google!) are brilliant proofs of this.
- vog 10y agoMore generally, if a company uses software X (open-source or not), they need to: a) make a contract with a company that takes responsibility for X, or b) hire somebody who takes responsibility for X, or c) take responsibility for X on your own It doesn't help to "buy" closed-source software X from another company if you can't count on them in case of emergency, i.e. if they vanish, go bankrupt or put their lawyers onto you. Then, better take open-source software where you can take responsibility on your own, for which it may help to hire one or more of the lead developers.
- mirekrusin 10y agoIt's interesting that this type of initiative, which is admirable, will spike up some java "popularity" metrics on GitHub.
- hokkos 10y agoHow does it work for transitive depandancies ? If you use a package that use a vulnerable Apache common? Does a pr is sent to update the package when it is updated?
- 11928311 10y agoSo, Google does ... something and is showered with praise. Thousands of volunteers work in the saltmines and get nothing. Business as usual. Myths like "Google sponsored Python!!!" propagate when they do nothing at all. Disgusting.
- muzster 10y agoOperation Rosebud
- hawski 10y agoI was thinking about doing something similar with bigquery and github data to search for uses of strncpy in C code. But I am not that good with the query language and also bigquery didn’t support multiple users properly (this adds friction). I still think it’s a good idea. It would be even better to search for a few C pitfalls more, but strncpy is probably the easiest to search for.
- rburhum 10y agoMad thank-yous to Google for this!
- lvlds 10y agoAwesome! Contgrats to the team!
- mrgrowth 10y agoI read so many of these kinds of articles out of curiosity and rarely understand them. Thank you for adding in the part about the bank teller. For reference: "it would be like hiring a bank teller who was trained to hand over all the money in the vault if asked to do so politely, and then entrusting that teller with the key."
- Dem0stheneS 10y agoThat's outstanding news. Hats off to the volunteers doing the work on this.