4 ms·
Rewriting Every Syscall in a Linux Binary at Load Time
- CableNinja 6mo agoI assume this would break observability through existing methods, right? If you were to strace a process that has been patched, would you see regular syscall data (as if it wasnt patched) or would your syscall replacement appear along the way?
- amitlimaye 6mo agoGood question. I didn't cover this in the post — the binary doesn't run on the host kernel directly. It runs inside a lightweight KVM-based VM with no operating system. The shim is the only thing handling syscalls inside the guest. So strace on the host wouldn't see anything — no syscalls reach the host kernel from the guest. From the host side, the only visible activity is the hypervisor process making syscalls on behalf of the guest. Inside the guest, there's no kernel to attach strace to — the shim IS the syscall handler. But we do have full observability: every syscall that hits the shim is logged to a trace ring buffer with the syscall number, arguments, and TSC timestamp. It's more complete than strace in some ways — you see denied calls too, with the policy verdict, and there's no observer overhead because the logging is part of the dispatch path. So existing tools don't work, but you get something arguably better: a complete, tamper-proof record of every syscall the process attempted, including the ones that were denied before they could execute. I'll publish a follow-on tomorrow that details how we load and execute this rewritten binary and what the VMM architecture looks like.
- coppsilgold 6mo agoYou mentioned SECCOMP_RET_TRACE, but there is also SECCOMP_RET_TRAP[1] which appears to perform better. There is also KVM. Both of these are options for gVisor: <https://github.com/google/gvisor https://github.com/google/gvisor> [1] <https://github.com/google/gvisor/blob/master/pkg/sentry/platform/systrap/README.md https://github.com/google/gvisor/blob/master/pkg/sentry/plat...>
- monocasa 6mo agoThere's also SECCOMP_RET_USER_NOTIF, which is typically used by container runtimes for their sandboxing.
- coppsilgold 6mo agoSECCOMP_RET_USER_NOTIF seems to involve sending a struct over an fd on each syscall. Do they really use it? Performance ought to suffer. Also gVisor (aka runsc) is a container runtime as well. And it doesn't gatekeep syscalls but chooses to re-implement them in userland.
- xuhu 6mo agoSECCOMP_RET_USER_NOTIF appears to switch between the tracee and tracer processes for each syscall. Using SECCOMP_RET_TRAP to trigger a SIGSYS for every syscall in IO intensive apps introduces 5% overhead (and avoids a separate tracer). I wonder if there's any mechanism that works for intercepting static ELF's like Go programs and such.
- monocasa 6mo agoThey use a seccomp filter to decide which syscalls get sent to the other process for processing.
- foota 6mo agoHah, I've been looking into something amusingly similar to track mmap syscalls for a process :)
- pocksuppet 6mo agoWhy not just use ptrace?
- amitlimaye 6mo agoptrace is atleast 2 context switches that will make it pretty slow
- foota 6mo agoYeah this wasn't something like "I want to debug a program" but rather I wanted to be able to track mmaping for later cleanup. Fortunately libc doesn't mmap that much internally so I think I can get away alright with interposing lib's mmap call.
- jmillikin 6mo agoThis might be a very dumb question, but if the process is being run under KVM to catch `int 0x03` then couldn't you also use KVM to catch `syscall` and execute the original binary as-is? I don't understand what value the instruction rewriting is providing here.
- rep_lodsb 6mo agoYes, that seems unneccessary. The overhead of trapping and rewriting every syscall instruction once can't be (much) greater than that required for rewriting them at the start either. Even if you disallow executing anything outside of the .text section, you still need the syscall trap to protect against adversarial code which hides the instruction inside an immediate value: foo: mov eax, 0xc3050f ;return a perfectly harmless constant ret ... call foo+1 (this could be detected if the tracing went by control flow instead of linearly from the top, but what if it's called through a function pointer?)
- rep_lodsb 6mo agoThinking a bit more about it (and reading TFA more carefully), what's the point of rewriting the instructions anyway? I first assumed it was redirecting them to a library in user mode somehow, but actually the syscall is replaced with "int3", which also goes to the kernel. The whole reason why the "syscall" instruction was introduced in the first place was that it's faster than the old software interrupt mechanism which has to load segment descriptors. So why not simply use KVM to intercept syscall (as well as int 80h), and then emulate its effect directly, instead of replacing the opcode with something else? Should be both faster and also less obviously detectable.
- jacobgorm 6mo agoGood point, an int3 is not going to be faster than a syscall, and if they implement the sandboxing policy in guest userspace is seems it would be quite easy to disable.
- jacobgorm 6mo ago
- ozgrakkurt 6mo agoReally informative writing thank you. How secure does this make a binary? For example would you be able to run untrusted binary code inside a browser using a method like this? Then can websites just use C++ instead of javascript for example?
- lmz 6mo agoThey already can use C++ if they want to. Emscripten? Jslinux?
- ozgrakkurt 6mo agoI mean just distributing the regular compiled x86_64 binary and then running it as a normal executable on the client side but just using that syscall shim so it is safe.
- direwolf20 6mo agoIf you think about the fundamentals involved here, what you actually need is for the OS to refuse to implement any syscalls, and not share an address space. A process is already a hermetically sealed sandbox. Running untrusted code in a process is safe. But then the kernel comes along and pokes holes in your sandbox without your permission. On Linux you should be able to turn off the holes by using seccomp.
- amitlimaye 6mo agoseccomp is a very coarse filter and a very limited action set. think what you could do if you could see the payload of the syscall or change the output of a read syscall depending on agent identity.
- amitlimaye 6mo agoyes that is the goal though C++ is something i am not targetting in the short term. The idea is to be able to run untrusted binaries in a vm with no kernel. saves memory makes for faster loads and the the bin cannot escape the vm so it can never compromise your host.
- edf13 6mo ago[dead]
- im3w1l 6mo agoWhat about int 80h?
- jcalvinowens 6mo agoYeah, I had the same question. But I'd guess they probably disable IA32 completely.
- amitlimaye 6mo agoInt80 is a great idea but int3 is what i landed on when i was looking and at this point just trying to get something working. The good thing about int80 is a 2 byte instruction i believe rather than int3 + nop that i am doing right now
- deleted 6mo ago[deleted]
- im3w1l 6mo agoI think you misunderstand my question. int 80h is an alternative legacy way that a program can issue syscalls. So without handling that your system may miss some syscalls. Which may be fine, I'm sure they are not that common. But if someone were to try to sneak a syscall past your monitoring that might be something they might do? Edit: Or maybe since it's running in a vm the outcome might just be that it doesn't work at all which may be fine I suppose.
- JSR_FDED 6mo agoLove the detailed write up, thanks! This is the kind of foundation that I would feel comfortable running agents on. It’s not the whole solution of course (yes agent, you’re allowed to delete this email but not that email can’t be solved at this level)… let me know when you tackle that next :-)
- amitlimaye 6mo agoAMA i am the author of that blog i have some working code just not something i want to share right away. Right now i am chasing density but yes security is something i will get to eventually. the issue is what to implement first :). This is the first of a series of blogs i am writing. you can check my substack. the next step is to show a density,launch speed demo hopefully middle of next week
- hparadiz 6mo agoI've been thinking of making a kernel patch that disables eBPF for certain processes as a privacy tool. Everyone is using eBPF now.
- xelaboi 6mo agoYou either have a writing style that is uncannily similar to what an LLM generates, or this article was substantially written by an LLM. I don't know what it is about the style, but I just find it a bit exhausting, like an overfit on "engaging writing" that strips away sincerity.
- renewiltord 6mo agoIt’s clearly LLM written but the idea was interesting enough that I read it. I suspect based on username the writer is cleaning up their voice. I think the idea of sharing the raw prompt traces is good. Then I can feed that to an LLM and get the original information prior to expansion.
- nonameiguess 6mo agoName sounds very likely not an English speaker. And the one reply here to a top-level comment is extremely obvious. I think it's unfortunate that people who write English poorly feel the need to do it, but I get it at least. The person behind this probably has a real interest and knowledge in the space but feels they can't communicate it without assistance. It is too bad, though. People bad at English will themselves be reading this forever now and think this is the way real people write, speak, or are supposed to. It's many things. The relentless ethusiasm about everything. Prefacing any answer to a question with an affirmation that it was a good question first. And yes, sorry, pedants of the web who feel witch-hunted because you knew how to employ keyboard shortcuts and used em-dashes in 2015 and have the receipts to prove it -- you never used 17 in the span of a single page. I think that was the first I can remember using ever and I had to contrive a way to do it where a semi-colon wouldn't clearly work better.
- qbane 6mo agoThere is even a table copy-pasted into a paragraph without noticing. > What’s needed is something different: > Requirement ptrace seccomp eBPF Binary rewrite Low overhead per syscall No (~10-20µs) Yes Yes Yes [...]
- notepad0x90 6mo agoI think it's better to just adapt to this. A lot of people write the content their own way, and get AI to rewrite it so that it is more readable, and free from errors. Content over appearance and all. I think the problem is you consider this auto-completion tool insincere. many do as well, because they anthropomorphize LLMs, it feels like a different sentient entity wrote it than the person posting it. but in reality, that isn't the case; it's more like a spellchecker that helped the person communicate their idea. The purpose of language is to communicate meaning and intent, not to sound or feel a particular way, unless you're reading for entertainment or enjoyment. This is the second post I'm commenting on within a span of like 30 minutes where someone did some really good work and shared it, but the top comments are complaining about AI usage. Either LLM-assisted content needs to be banned entirely (might be), or complaining about it should be considered a breach of etiquette at sites like HN that are tech-centric.
- szmarczak 6mo ago> It can’t detect the interception What's stopping the process from reading its own memory and seeing that the syscall was patched?
- amitlimaye 6mo agoActually you are right nothing is stopping it from reading but that does not help it escape the kernel. If you are worried about something adversarial that tries to detect its in a sandbox but that is not what we are trying to protect from the idea is to follow the same model of a container with something that is more secure and has less surface area to protect or attack.
- twic 6mo ago[dead]
- Thaxll 6mo agoIt's pretty much what gVisor does. https://gvisor.dev/ https://gvisor.dev/
- Thaxll 6mo agoSo why not using it instead of re-implementing the exact same thing.
- deleted 6mo ago[deleted]
- compsciphd 6mo agothis has been done for ages with a simple kernel module that just wraps the real kernel syscall, no binary changes needed. example how we used it in early 2000s to implement pre linux namespace containerization. https://www.usenix.org/legacy/publications/library/proceedings/osdi02/tech/full_papers/osman/osman_html/index.html https://www.usenix.org/legacy/publications/library/proceedin... (note the shepherd and where kubernetes arguably got the pod name from). and security policies on top of it https://www.usenix.org/legacy/event/lisa07/tech/full_papers/potter/potter_html/index.html https://www.usenix.org/legacy/event/lisa07/tech/full_papers/...