8 ms·
Into the Borg – SSRF inside Google production network
- how2cflags 8y agoWebarchive link to the article above, as I found myself unable due to the system load on the page. https://web.archive.org/web/20180720170255/https://opnsec.com/2018/07/into-the-borg-ssrf-inside-google-production-network/ https://web.archive.org/web/20180720170255/https://opnsec.co...
- tptacek 8y agoThis is a great find. SSRF is a really unappreciated vulnerability; it is usually game-over. That it came from a Caja audit adds some tasty irony. A friendly word of advice: when you find flaws like this (you, the reader, not you, the guy who wrote this post), think carefully before disclosing internal network details you discover like this writer did. The internal details of a target network don't become public domain simply because you found a vulnerability. There are firms that get extremely itchy about this kind of stuff getting published, and I can't blame them.
- outworlder 8y agoI'm still amazed that this individual published as much data as he did. I don't see where it says they have permission to do so. Specially because: > I hope they won’t beat me with a stick for disclosing any of this This tells me that this wasn't cleared with them. Doesn't sound like a smart move. If your gut is telling you it may be a bad idea, it could be because it is...
- bowmessage 8y agoYeah, while it might be interesting to know what Gmail's system user is, it's not really related at all to the vuln. I wouldn't be surprised if payment of the bounty was contingent on some kind of ToS acceptance about responsible disclosure...
- londons_explore 8y agoThe page he accessed probably contained thousands of Google borg jobs. Many about secret or future projects for example. Many giving implementation details for secret sauce algorithms (eg. oh - they precompute all possible misspellings for their spell corrector via this mapreduce!). Simply knowing the fleet wide CPU, network and RAM usage for Gmail would give a competitor a lot of knowledge into the probable running costs of the service for example. In this case he exposed a tiny proportion of what he accessed. Sounds fairly reasonable to me.
- asfasgasg 8y agoUh, how does that follow? In what other domain do you get a pass for doing a little unnecessary harm because you could have done much more?
- elvinyung 8y ago> I should mention that Borg, like Kubernetes, relies on containers like Docker Hmm, containers _like_ Docker, or Docker? I thought Google used lmctfy since long before Docker?
- powera 8y agocgroups were written at Google, and have been used internally for a very long time; they provide "container"-like limits on resource usage for a group of processes. I assume that Google isn't using Docker internally for production services, but don't know for sure (and I assume anyone who does know for sure can't tell you).
- elvinyung 8y agoYeah, so to clarify, I know from the Borg paper that Google basically implemented the first cgroups and the first cgroups-based containers. I'm pretty sure that lmctfy was the open-sourcing of this work, but it's also been deprecated and last I heard, that code was moving to libcontainer/runc. If (as the article implies) Google uses Docker internally, that would be a surprising and interesting bit of news.
- puzzle 8y agoThey mentioned that Kubernetes runs some workloads and they are probably using Docker for that, like GKE does (17.03, I think). No way they would use it for Borg.
- beering 8y agoIt also just doesn't make sense for Google to use Docker (or even Kubernetes) for their core infra. They've been the foremost leader in distributing containerized applications across data centers. Whatever they've already built is almost certainly more battle-tested and more customized to their needs than anything public that's based on their concepts.
- elvinyung 8y agoSure it's battle tested, but it's quite possible that the advantage of having a software with more eyeballs on it is even more valuable. Last I heard, the biggest reason that Google still uses Borg instead of Kubernetes is mostly because of switching costs.
- hkr_mag 8y agoFor those who is new to the world of SSRF vulnerabilities, check the SSRF Bible (full disclaimer: I'm with Wallarm): https://docs.google.com/document/d/1v1TkWZtrhzRLy0bYXBcdLUedXGb9njTNIJXa3u9akHM/edit https://docs.google.com/document/d/1v1TkWZtrhzRLy0bYXBcdLUed...
- luhn 8y agoNot quite relevant to the article: Since it seems to be tricky to properly sanitize URLs for SSRF, I had an idea for safely calling user-defined URLs: Set up an unprivileged non-VPC Lambda function that calls a URL and call all user-defined URLs through the Lambda function. I think it should be bulletproof, anything I'm overlooking?
- bpicolo 8y agoGCP does it by requiring a header on requests to metadata systems. Require a header on your internal services and make sure you never send that header with user-requested URLs and you can guarantee safety there.
- kerng 8y agoStill allows tons of attacks and recon, like port scans, or even crashes of processes- since internal things might not go through the same fuzzing scrutiny as external endpoints.
- puzzle 8y ago> Google is still relying on Borg for its internal production infrastructure, but I can tell you it’s not because of the design of Borg interfaces! No matter how spartan, the Borg status pages are more helpful than most Kubernetes UIs out there when it comes to debugging a problem in depth, i.e. past CPU and memory graphs. Part of that is made possible by applications exposing debugging endpoints and telling Borg about them.
- GauntletWizard 8y agoYeah, I would kill for the k8s pod contract to have something like the borg status line.
- heavenlyblue 8y agoAs a matter of post-SLA curiousity, what is in the status line?
- jbeda 8y agoIt's something we've talked about on and off. Never seems to make it happen.
- threeseed 8y agoCan't you just use CAdvisor and route metrics to Prometheus/Grafana ?
- puzzle 8y agoThey're sometimes interactive HTTP endpoints, not metrics, although metrics are served by one of them (/varz, described in the SRE book). They're informally called z pages and are the inspiration behind the /debug handlers that Go's net/trace package installs: https://pbs.twimg.com/media/CebGW64W4AAJ_sc?format=jpg https://pbs.twimg.com/media/CebGW64W4AAJ_sc?format=jpg OpenCensus has a bunch, too: https://opencensus.io/zpages/ https://opencensus.io/zpages/
- jamesblonde 8y agoYarn apps typically expose a similar http endpoint for app status reports.
- dilyevsky 8y agoHa! I spent so much time squinting at that borglet page. Such nostalgia
- paulddraper 8y agotl;dr SSRF is server side request forgery, where you can gain access to private resources by convincing privledged servers to make requests for you. If you are using network access for security, either don't, it blacklist private IPs (and use public DNS) for untrusted URLs.
- justicezyx 8y agoVery impressive findings. Disclaimer: I am part of Borg team.
- londons_explore 8y agoTime to make everything accessible via stubby only, none of this http nonsense...
- asfasgasg 8y agoIntegrity and privacy for all streams.
- menage 8y ago> it’s not because of the design of Borg interfaces As the original creator of the Borglet status page, I think it's not really accurate to describe it as "designed". :-) It was more a case of gradual semi-random evolution, with people (mostly me, at least in the old days - I've no idea how much it might have changed in the last eight years) adding things that seemed like they could help Google engineers trying to track down issues with their Borg jobs. User-friendliness for random SSRF hackers was definitely not high on the priority list.
- kevincox 8y agoAs a user of the borglet pages they aren't pretty but they are quite functional and to-the-point. I can appreciate the no-frills design.
- dandigangi 8y agoI found this really cool to read into. Some of it went over my head but a great read none the less. Thanks for sharing.
- dandigangi 8y agoWhy was this downvoted? I really don't understand the users of this site sometimes.
- heipei 8y agoborglet status page in full resolution: https://opnsec.com/wp-content/uploads/2018/07/borg2.png https://opnsec.com/wp-content/uploads/2018/07/borg2.png
- pjjw 8y agoi can't help but very lol @ kubernetes as the successor to borg :0