7 ms·
Server-side sandboxing: Containers and seccomp
- Elena120 3y ago[flagged]
- nosefrog 3y agoWe had seccomp containers at Dropbox, and I remember Max Serrano helping me set that up with ReactServer :) Talented engineer, though I do remember that the jails were kind of a maintenance nightmare for the security team.
- johnkoepi 3y ago"seccomp containers" sounds weird... like what is a container in Linux anyway :D
- minitoar 3y agoNice, love seccomp though as with all things security it can be very fiddly.
- rubenfiszel 3y agoIf you are looking to self-host a scalable backend that runs arbitrary code in python/typescript/bash/go with optional sandboxing using nsjail like figma, nsjail is what we use as isolation layer at https://windmill.dev https://windmill.dev (Open-source alternative to Retool/Airplane) (Our python nsjail config for instance: https://github.com/windmill-labs/windmill/blob/main/backend/windmill-worker/nsjail/run.python3.config.proto https://github.com/windmill-labs/windmill/blob/main/backend/...)
- jagrsw 3y agonsjail author here (the original one, as the tool is also maintained by others), good job! Irrelevant nit: .proto files are protobuf definition files (like this one: https://github.com/google/nsjail/blob/master/config.proto https://github.com/google/nsjail/blob/master/config.proto), a text representation of a specific protobuf contents is typically called (as per man clang-format): .textpb .pb.txt or .textproto - I use .config for examples distributed with nsjail, but it's licentia poetica :)
- rubenfiszel 3y agoThe wonders of HN strikes again. Thank you for this amazing piece of technology that is nsjail. Nsjail is very core to our security, our multitenant would be so slow without it and I think we're one of the applications that leverage it in a way that showcase nsjail to its full extent (as in, we beat containers/firecracker cold starts by a fair margin while keeping most of their benefits). That's one of the reason we're order of magniture more efficient than Airplane that uses fargate under the hood. I would love to chat if you had time, my email in my profile.
- xyzzy_plugh 3y agoOddly enough the canonical extension seems to now be none of those but .txtpb: https://protobuf.dev/reference/protobuf/textformat-spec/#text-format-files https://protobuf.dev/reference/protobuf/textformat-spec/#tex...
- shooshx 3y agoWhy not just use json?
- jagrsw 3y agoI may be mistaken, but does JSON offer the ability to define a schema with default values? Utilizing a single .proto file, I can tackle both the issues of default values and configuration structure, eliminating the need to manually check for missing mandatory sections. However, I presume there are now JSON extensions that provide similar functionality?
- freeney 3y agoRunning arbitrary user code inside a jail that doesn’t isolate networking might not be enough isolation. Also kernel mount namespace binds into the jailed env increases the attack surface. Great for some use-cases, but multi-tenant workloads might need a tighter setup? I'm definitely going to give Windmill a try. It looks really cool!
- remram 3y agoWow, this nsjail setup is now part of your opensource version? Last I tried Windmill there was no isolation mechanism for scripts on the free version.
- imiric 3y agoI haven't used seccomp, but have recently been playing around with the Linux pledge port[1]. It has a very friendly UI, but I still struggled with allowing some complex apps to run at all, because of the sheer amount of syscalls and devices they required. Digging through a mountain of strace output is tedious... Can someone with experience with both comment on how (the Linux port of) pledge compares to seccomp? Can it be considered a replacement at this point? It seems like it could handle the last scenario described in the article fine, since it allows setting granular rwcx permissions on individual paths. [1]: https://justine.lol/pledge/ https://justine.lol/pledge/
- ShowalkKama 3y ago>Digging through a mountain of strace output is tedious did you consider logging the syscalls invoked during normal usage with 'strace --output=/some/dir -f ...'? This + grep + uniq should make it really simple.
- eyberg 3y agoWe too ended up adding pledge and unveil to Nanos. Seccomp and seccomp-bpf are indeed entirely way too limiting. It wasn't really designed for end app developers who are, imo, the ones that should be dictating the policy. The whole lack of pointer deref'ing makes it really difficult for application level developers to make policies that are easier to create. The promises arg in pledge, https://man.openbsd.org/pledge.2 https://man.openbsd.org/pledge.2 , does a decent job of grouping related calls together but I think there is a ton of room to make all of this a lot better than it is today.
- Arch-TK 3y agoPledge works well if the software developers implement it on their own application. It also works well if the software developers document what syscalls they rely on and what permissions they need. When it comes to retrofitting something like pledge (or seccomp) into an existing application when you've not developed it and/or can't easily tell what syscalls are being called then it's always a nightmare. It doesn't really matter if it's pledge or seccomp at that point (although undoubtedly seccomp is far harder to make use of), if you're doing this kind of security by retroactive whitelist, you're going to have trouble making it work. It's going to take time and effort to implement.
- zxcvgm 3y agoIt's pretty easy to apply seccomp to a process using systemd by adding SystemCallFilter= in its unit file. There's a reasonable set of permitted syscalls for general system processes, aptly called `@system-service`, but you can tweak that to suit your needs [1]. I generally use this, among other settings, to further lock down system services [2]. [1] https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html https://www.freedesktop.org/software/systemd/man/latest/syst... [2] https://www.redhat.com/sysadmin/mastering-systemd https://www.redhat.com/sysadmin/mastering-systemd
- CAP_NET_ADMIN 3y agoYep, can recommend systemd in this case, really easy to apply basic hardening to services that just works.
- deleted 3y ago[deleted]
- declan_roberts 3y agoseccomp is heaps better than selinux, but still too overly complicated to be using in everyday production unless you're truly on the "refine and secure" path or dealing with high-stakes sandboxing.
- johnkoepi 3y agoway too different things, everyone using seccomp when they don't have AppArmor only profile. sometimes even do both.
- lmeyerov 3y agoGood intro. I'd be curious how they do the syscall tracing, eg, strace logs as part of CI? Funny enough, we've gone the reverse path for LLM AI-generated code sandboxing for louie.ai / Graphistry . We started with container isolation with careful network, volume, compute etc enablement first, and only now adding nsjail to the runners within the container as an extra defense layer. The negative space is interesting too. We initially explored alternatives like wasm (too slow and underpowered for our generated python GPU analytics workloads) and firecracker vm (too unwieldy and unportable for our small team). As we do more k8s and enable more interactive data viz customization + web-scale static serving, would love to revisit both. On which note, we have a bit of budget for someone to help harden the nsjail layer, if of interest!
- hhh 3y agoI have yet to find a firecracker-style thing for k8s that is simple to deploy. Firekube seemed interesting, but is archived... Liquid Metal from Weaveworks seems interesting but I don't even know where I would start.
- remram 3y agoKata just released a new version, it is the only thing that I've found easy to setup with k8s... though my experience running Docker-in-Kata hasn't been very good.
- lmeyerov 3y agoWe were looking at Kata as well, especially as an 'easier' firecracker, though I forgot why we didn't go further, and I've been curious why they seem to get little attention in practice I believe part is portability, as they may require nested virtualization features to be available, and maybe QEMU overhead. Maybe also something about use in China vs elsewhere? They (and QEMU) have been around a long time...
- flurie 3y agoVirtink[1] has been reasonably stable as long as you're okay with Cloud Hypervisor instead. [1] https://github.com/smartxworks/virtink https://github.com/smartxworks/virtink
- sargun 3y agoSeccomp BPF is great. There was some recent issues due to IO_uring and extensible syscalls, but I believe for now, those issues are avoidable. I believe the next generation looks something like landlock (https://docs.kernel.org/userspace-api/landlock.html https://docs.kernel.org/userspace-api/landlock.html).
- johnkoepi 3y ago+1
- johnkoepi 3y agoI love ideas behind Landlock but I don't fully see the struggle currently without taking into considerations issues with io_uring api. Seccomp nowadays with AppArmor|SElinux is enough even for Nested rootless containers. Nested even into std runc things. Both AppArmor and Seccomp profiles are stackable. If you don't need to generate unique profiles per each container you should be fine...
- bdahz 3y agoSo what's the difference between nsjail[1] and bubblewrap[2]? [1] https://github.com/google/nsjail https://github.com/google/nsjail [2] https://github.com/containers/bubblewrap https://github.com/containers/bubblewrap
- xyzzy_plugh 3y agobubblewrap aims to be reasonably secure by default but leaves sleeping soundly at night as an exercise to the reader. It's not exhaustive. It's more of a blast radius/convenience tool. Conversely nsjail aspires to facilitate sleeping soundly out of the box, with security as the primary motivating factor.
- ximm 3y agoI don't have extensive experience with nsjail, but from reading the docs it seems to me like nsjail covers namespaces, cgroups and virtual networking, while bwrap only covers namespaces. On the other hand, bwrap is deliberately kept simple because it is SUID.
- baggy_trough 3y agoCheck out systemd-nspawn. Built in and works great!
- IcyWindows 3y agoAre operating systems failing at their jobs if one can't run independent workloads on them anymore? It seems like something is broken, and we are all patching things up piecemeal.
- SeriousM 3y agoIs there an isolation method close to the functionality to nsjail but for .net code? I know I can protect my AppDomain but how to protect the system/network from rouge .net code?