7 ms·
Answering the security question specifically: v8 is a runtime and not a security boundary. Escaping it isn't trivial, but it is common [1]. You should still wra
by dsl 4y ago
Answering the security question specifically: v8 is a runtime and not a security boundary. Escaping it isn't trivial, but it is common [1]. You should still wrap it in a proper security boundary like gVisor [2].
The security claims from Cloudflare are quite bold for a piece of software that was not written to do what they are using it for. To steal an old saying, anyone can develop a system so secure they can't think of a way to defeat it.
1. https://www.cvedetails.com/vulnerability-list/vendor_id-1224/product_id-17734/Google-V8.html https://www.cvedetails.com/vulnerability-list/vendor_id-1224...
2. https://gvisor.dev/ https://gvisor.dev/
- kodah 4y agoPersonally, I think I'd recommend SELinux and SecComp before I use something like gVisor. There's significant performance impact with application kernels.
- opportune 4y agoThose are not mutually exclusive, but if you use a syscall blocklist only you will have to refuse certain workloads
- azakai 4y agoTo be fair to Clouldflare, they do a lot of very thoughtful work on sandboxing at all levels, including at the process level as well as above and below that. A lot of interesting details here: https://blog.cloudflare.com/mitigating-spectre-and-other-security-threats-the-cloudflare-workers-security-model/ https://blog.cloudflare.com/mitigating-spectre-and-other-sec... The article focuses on Spectre, but it also mentions general stuff. edit: It is true they are using V8 for a novel purpose, but a lot of their techniques mirror how Chrome sandboxes V8 (process sandboxing, an outer OS sandbox, etc.).
- jeremyjh 4y agoI have a lot of respect for Kenton, but the facts are: v8 security team recommends running untrusted code in separate processes [1], and as described in your link, Cloudflare doesn't do that. From what Kenton said earlier on this[2], it sounded like they had already committed to their architecture before that v8 recommendation was made. They thought through a lot of possible attacks and created some clever mitigations, and ultimately decided "its probably ok". [1] https://v8.dev/docs/untrusted-code-mitigations#sandbox-untrusted-execution-in-a-separate-process https://v8.dev/docs/untrusted-code-mitigations#sandbox-untru... [2] https://news.ycombinator.com/item?id=18280061 https://news.ycombinator.com/item?id=18280061 eta: Also, Chrome does use process isolation - every browser tab runs in a separate process. The whole point of Cloudflare isolates is to avoid the overhead of a fork() for each worker request.
- hinkley 4y agoAlso Cloudflare has already broken the Internet once by doing unorthodox things. Sometimes the guy missing a finger is exactly the one you want to take safety advice from. Sometimes it's just a matter of time before he's missing three fingers.
- skybrian 4y agoI don't think what Chrome is doing is quite that simple? When you click on a link in a tab, it might take you to a different site. Also, there can be iframes from different websites in the same browser tab. Also see "limitations" on this page: https://www.chromium.org/Home/chromium-security/site-isolation/ https://www.chromium.org/Home/chromium-security/site-isolati...
- jeremyjh 4y agoThat link tells you this: > Cross-site documents are always put into a different process, whether the navigation is in the current tab, a new tab, or an iframe (i.e., one web page embedded inside another). Note that only a subset of sites are isolated on Android, to reduce overhead.
- kentonv 4y agoChrome's use of strict process isolation is a fairly new thing. For most of its history, it was trivially easy to open any other site in the same process as your own site (via window.open(), iframe, etc.), and V8's sandboxing was the thing protecting those sites from each other. So V8 was, in fact, designed for this. When Spectre hit, Chrome concluded that defending the web platform from Spectre in userspace was too hard, so they decided to go all-in on process isolation so that Spectre was the kernel's problem instead. This is a great defense-in-depth strategy if you can spare the resources. On most platforms where Chrome runs, the additional overhead was deemed worth it. (Last I heard there were still some lower-powered Android phones where the overhead was too high but that was a while ago, maybe they've moved past that.) That doesn't mean V8 has decided their own security doesn't matter anymore. Securely sandboxing the renderer process as a whole isn't trivial, especially as it interacts with the GPU and whatnot. So it's still very much important to Chrome that V8's sandbox remains tight, and the renderer sandbox provides secondary defense-in-depth. When it comes to Cloudflare Workers, the performance penalty for strict process isolation is much higher than it is for a browser, due to the finer-grained nature of our compute. (Imagine a browser that has 10,000 tabs open and all of them are regularly receiving events, rather than just one being in the foreground...) But we have some advantages: we were able to rethink the platform from the ground up with side channel defense in mind, and we are free to reset the state of any isolate's state any time we need to. That lets us implement a different set of defense-in-depth measures, including dynamic process isolation. More details in the blog post linked earlier [0]. We also did research and co-authored a paper with the TU Graz team (co-discoverers of Spectre) on this [1]. I may be biased, but to be perfectly honest, the security model I find most terrifying is the public clouds that run arbitrary native code in hardware VMs. The VM software may be a narrower attack surface than V8, but the hardware attack surface is gigantic. One misimplemented CPU instruction allowing, say, reading physical addresses without access control could blow up the whole industry overnight and take years to fix. Spectre was a near miss / grazing hit. M1 had a near miss with M1racles [2]. I think it's more likely than not that really disastrous bugs exist in every CPU, and so I feel much, much safer accepting customer code as JavaScript / Wasm rather than native code. [0] https://blog.cloudflare.com/mitigating-spectre-and-other-security-threats-the-cloudflare-workers-security-model/ https://blog.cloudflare.com/mitigating-spectre-and-other-sec... [1] https://blog.cloudflare.com/spectre-research-with-tu-graz/ https://blog.cloudflare.com/spectre-research-with-tu-graz/ [2] https://m1racles.com/ https://m1racles.com/
- marktangotango 4y agoI've been hoping a runtime would arise for this use case explicitly, and v8 is close. But the solution has to have multitenancy, limiting of memory and cpu, and security baked in from the beginning. The lua runtime seems promising, but again, not designed for that.
- hinkley 4y agoWe are going the long way around to figure out that Tannenbaum was right and we need microkernels. The IPC costs cause sticker shock, but at the end of the day perhaps they are table stakes for a provably correct system.
- doliveira 4y agoI was reading the Wikipedia page about this debate and I find this low-level stuff so fascinating. Wish I had done CS instead of Physics. At a low level and at a place of ignorance, I kind of wish monolithic kernels hadn't won, just like I wish the morbidly obese browsers hadn't either... They sound way more, well, sound in their theoretical basis
- hinkley 4y agoThere's a few clever tricks that seem to make a pretty big difference with microkernels, but I know just enough to be dangerous. I think one of the biggest ones I recall, and I think it was part of the L4 secret sauce, was the idea of throwing a number of small services into a single address space, so that a context switch requires resetting the read and write permissions, but not purging the lookup table, making context switches a couple times faster. With 64 bit addressing that gets easier, and allows other tricks. For instance recent concurrent garbage collectors in Java alias the same pages to 4(?) different addresses with different read/write permissions. Unfortunately this means your Java process can 'only' address 4 exabytes of memory instead of 16, but we will surely persevere, somehow. How that plays with Meltdown and Specter might put the kibosh on this though, and most processors don't actually support 64 bit addressing. Some only do 44, 48. That's pretty old news though. The more interesting one to me is whether there's a flavor of io_uring that could serve as the basis for a new microkernel, allowing for 'chattier' interprocess relationships to be workable by amortizing the call overhead across multiple interactions.
- hinkley 4y agoDocker went through an era of thinking they could built multitenant systems that amortized the cost of resources across customers. Java went through an era of thinking they could build multitenant VMs that amortized resources across customers. Unix went through an era of thinking they could build multitenant processes that amortized resources across customers. As did Windows. I point out that Docker is trying - and failing - to offer us the same feature set we were promised by protected memory, preemptive multitasking operating systems 30 years ago. I'm still waiting. I don't recall if there were old farts complaining about this in the early 90's, but I suspect they exist. Also Java had a bunch of facilities for protection and permissions that I've not heard anyone claim exist in V8. Not that feature equality is required (in fact that may seal your doom), but I don't see feature parity either. edit: some of facilities Java investigated do exist in cgroups, so one might argue that there is a union of v8 and cgroups that is a superset of Java, but unless I am very mistaken, isolates don't work that way, so isolates are still not the way forward there.
- treis 4y ago>cost of resources across customers. IMHO, this is all a solution chasing a problem. A server from Hetzner costs the equivalent of 30 minutes of my time a month. A VM on a shared machine can cost even less. There's just not cost there to save in order to justify the security risks and performance ghosts. It only makes sense for extremely bursty loads or background processing where latency isn't important. But those are pretty atypical scenarios and usually buying for max load is still cheaper than the engineering time to set up the system.
- ignoramous 4y ago> A server from Hetzner costs the equivalent of 30 minutes of my time a month. My toy code, which takes ~6ms (at p75) to exec, runs in 200+ locations, serves requests from 150 different countries (with users reporting end-to-end latencies in low 50ms). This costs well below $50/mo on Cloudflare, and $0 in devops. Make what you will of that.
- 4y ago
- kube-system 4y agoI find it funny how we've ended up here: "Let's use containers, that way we aren't running unnecessary redundant kernels in each VM!" later "Oh shit, our containers share a kernel, let's add another kernel to each container to fix this!" Maybe the next step is that we realize there's so many CPU bugs, that we really just need to give each container their own hardware :)
- wmf 4y agoThat's what cloud-native processors are for.
- com2kid 4y ago> Maybe the next step is that we realize there's so many CPU bugs, that we really just need to give each container their own hardware :) I am reasonably sure that most of the micro services I write would be very happy running on a 400mhz CPU with a couple hundred megs of RAM, if they were rewritten in native code, or even just compiled to native code instead of being ran on top of Node. Throw it all on a minimal OS that provides networking and some file IO. How much does it cost to manufacture 400mhz CPUs with onboard memory? Those must cost a few pennies each, throw in a 4GB SSD, preferably using SLC chips for reliability, and a NIC, and sell 4 "machine" clusters of them for ~$100 a pop.
- 0xbkt 4y ago> [...] if they were rewritten in native code, [...] Throw it all on a minimal OS that provides networking and some file IO. You may want to check out MirageOS[0]. It gives you a library OS with the primitives you say you need, and then all you have to do is import them in your application code as if you are writing your typical OCaml, build the virtual appliance and boot it up anywhere you want. [0] https://mirage.io/docs/overview-of-mirage https://mirage.io/docs/overview-of-mirage
- int_19h 4y agoThere's also https://github.com/includeos/IncludeOS https://github.com/includeos/IncludeOS However, how much is this kind of stuff is actually used at scale today?
- CJefferson 4y agoV8 is still very hard to escape -- an escape is a remote hole in Chrome isn't it? To be honest, I'm happy to base my security on "as safe as Chrome".
- zsims 4y agoIf the patch gap is small, yes. But are you patching V8? Node generally isn't.
- aboodman 4y agoAs a user of CF workers, the API is clearly designed to allow the runtime to redeploy e.g., for patches. Example: there is no notification when your worker is unloading. It just dies. This is sometimes annoying as a developer, but it's the kind of decision a platform makes when it wants (needs) to be able to shut down code arbitrarily and restart it later on a different version of the platform.
- esprehn 4y agoChrome runs each origin in a different process. Chrome also runs all the privileged code in separate processes, so like file, network, GPU, and many OS calls are not actually happening from within the same process as v8 executing web JS code. Process isolation is also necessary to mitigate CPU and OS level timing attacks between origins (and the privileged process). v8 is a pretty good security boundary, but to be as secure as Chrome you need the other layers too.
- aboodman 4y agoHi Elliot! I think it's more complex than that. Chrome (all web browsers) have a massive and ancient attack surface -- all of HTML, JS, CSS, every DOM and web platform API, every image codec, video, and on and on and on. CF workers have a comparatively minuscule attack surface that takes into account all the most recent lessons (high-res timers, for example). So yeah, if we call Chrome of 2008 a baseline, then Chrome's more recent process isolation increases security. But so to does CF's much reduced attack surface.
- moralestapia 4y agoI know V8 from back to front and I feel that it is quite safe, after all, there's not much you can do with a javascript interpreter ... Regarding the CVEs mentioned, the overwhelming majority of them have a "denial of service" effect, not sure how something like gVisor helps you mitigate that.
- helloooooooo 4y agoIf you know it back to front, then you are likely aware of the number of TurboFan vulnerabilities that have been found being exploited in the wild.
- moralestapia 4y agoSure, and most of them have been addressed properly. Honestly (as another commenter said), V8 being behind one of Google's largest revenue streams makes it feel "safe enough" for me, they have all the resources and incentive to make it as safe as realistically possible. Of course it's not perfect, but is there anything better at the moment?
- mkl 4y agoWith all their resources and incentive, Google have chosen to rely on process isolation in Chrome, not just V8.
- moralestapia 4y agoProcess isolation (security) belongs to the OS so there's nothing stopping one from doing the same, also, it shouldn't be much difficult to implement.