6 ms·
Segfaults in scripting languages are remarkably common, especially if arbitrary bytecode can be loaded into the VM. One I ran into in the wild recently is that
by lpghatguy 7y ago
Segfaults in scripting languages are remarkably common, especially if arbitrary bytecode can be loaded into the VM.
One I ran into in the wild recently is that in older versions of Lua, exceptions in GC finalizers (the `__gc` metamethod) can trigger a segfault. In those same versions of Lua, the bytecode format is notoriously dangerous to load.
I wonder whether this will be a large component of newer scripting language implementations. Do these safety issues warrant use of memory safe languages like Rust, or use of existing sandboxed VM implementations like WebAssembly?
- danielheath 7y agoIMO the biggest reason to adopt the WebAssembly format is that a segfault inside the runtime doesn't affect the host process at all. It's a plausible approach to a fully-safe, near-native-speed plugin architecture.
- Thaxll 7y agoSegfault inside the runtime, What does it means exactly? A segfault is by definition at the OS level. WebAssembly koolaid is strong on HN, let's wait the first exploits that escapes the runtime to assess the "fully-safe" architecture.
- wizzwizz4 7y agoI think they're trying to say that an out of bounds memory access within the emulated WebAssembly machine can be caught by the WebAssembly runtime. (I don't know whether this is true; I hardly know anything about WebAssembly.) The way they said it, though, makes it sound like WebAssembly is implemented with full process sandboxing or something, which is patently false. It works that way in neither Chrome nor Firefox, and there are no other browsers right now.
- rini17 7y agoWebAssembly follows the unfortunate C paradigm that no checks are done at runtime, only these that programmer requests explicitly (and only if no undefined behavior is involved), to improve speed. The WASM sandbox can call only set of specified host functions, but I expect so much functionality snowballing inside sandboxes that we'll have to allow everything including unsafe ones, anyway.
- Dylan16807 7y agoWebAssembly uses 32 bit indexes to access memory. The default way of implementing memory safety for it is to put a few gigabytes of dead address space before and after the memory a WASM program uses. This makes it impossible for the memory access instructions to escape, despite having no runtime checks. So it can segfault all the way up to the machine level. But more importantly, it's safe for it to 'segfault' out of its allocated memory, because it can't reach any other memory with 32 bit numbers.
- rini17 7y agoThis is nothing new, memory space of a process in all modern OSes is protected so that segfaults are safe too. Combined with original UNIX idea of small tools that communicate by pipes it was fine. But now the processes are behemoths with gigabytes of dynamically linked libraries that are too hard to secure and to restrict system access, we just enable everything. This will happen to wasm. I'm sure there are WASM blobs configured with unfettered access to DOM in the wild already.
- Dylan16807 7y ago> This is nothing new, memory space of a process in all modern OSes is protected so that segfaults are safe too. Right, but look at what I'm replying to. "The way they said it, though, makes it sound like WebAssembly is implemented with full process sandboxing or something, which is patently false." Unlike with Javascript, the program hosting a WASM script is immune to corrupt pointers inside the script. That's equivalent to OS-level isolation, which is pretty good! > But now the processes are behemoths with gigabytes of dynamically linked libraries that are too hard to secure and to restrict system access, we just enable everything. This will happen to wasm. I'm sure there are WASM blobs configured with unfettered access to DOM in the wild already. There will inevitably be bugs in the code handling the html/css/dom/rendering. But WASM greatly reduces the attack surface of the scripting VM. The types of bugs most likely to still exist with WASM are the types where isolating it in a separate process that communicates by pipes wouldn't help.
- chubot 7y agoSeg faults are "safe" in the programming language sense: A program fragment is safe if it does not cause untrapped errors to occur. Languages where all program fragments are safe are called safe languages. Therefore, safe languages rule out the most insidious form of execution errors: the ones that may go unnoticed. ... It is useful to distinguish between two kinds of execution errors: the ones that cause the computation to stop immediately, and the ones that go unnoticed (for a while) and later cause arbitrary behavior. The former are called trapped errors, whereas the latter are untrapped errors. Type Systems, Luca Cardelli https://scholar.google.com/scholar?cluster=9044245776831751011&hl=en&as_sdt=0,5&sciodt=0,5 https://scholar.google.com/scholar?cluster=90442457768317510... And practically speaking seg faults are easy to debug and fix. I see a lot of abuse of the terms "safe" and "safe language" lately.
- gizmo686 7y agoThe problem is that segfaults are frequently the latter type of problem. You do some unsafe thing and corrupt your state, but keep on going without obvious issue. Then at some point in the future, your corrupted state causes an otherwise bug-free portion of you program to segfault.
- chubot 7y agoThe problem isn't the segfault -- it's the earlier unsafe behavior. That is, the untrapped error that eventually caused the segfault. If the language only has trapped errors, which are indicated by seg faults, it's a safe language. Every language has such errors. What does divide by zero do? What does blowing the stack / infinite recursion do in Rust? It seg faults. The seg fault is the safe behavior. If the stack overflow overwrote heap data structures or global data structures and the program kept running, that would be unsafe.
- deathanatos 7y ago> What does divide by zero do? Returns a value, of course! (And that said, JavaScript does have some trapped errors, such as (1/0).foo.foo (and yes you need the second .foo…)) IMO, the execution "error" here (in this thread) is accessing memory illegally. Sometimes the runtime traps it, but sometimes it does not, and "sometimes" isn't always, so its effectively untrapped as we cannot depend on the trap. (Especially in adversarial circumstances.) And further, the original quote talks about languages — the behavior of a language like C is that memory access is not necessarily trapped; the behavior is not well defined. Given the lack of a requirement in the C language for a trap, I think it is fair to call C "unsafe" given the above definition of safe/unsafe. > What does blowing the stack / infinite recursion do in Rust? It seg faults. Somewhat interestingly, it detects it and SIGABRTs, which technically isn't a segfault. And that's now some black magic that I'm curious about as I really thought it would have segfaulted.
- tus88 7y ago> Segfaults in scripting languages are remarkably common What about Javascript running on V8?
- saagarjha 7y agoV8 is not immune to memory corruption.
- mailslot 7y agoomfg. seriously?!!
- shakna 7y agoYes, even there. [0] [0] https://github.com/nodejs/node/pulls?utf8=%E2%9C%93&q=is%3Apr+segfault https://github.com/nodejs/node/pulls?utf8=%E2%9C%93&q=is%3Ap...
- snek 7y agonode dev here - none of the prs here deal with segfaults from js code, because that doesn't really happen (at least, I've never seen it happen). You might have some more success searching V8's bug tracker.
- shakna 7y agoThis [0] one had a segfault in a JS test file. [0] https://github.com/nodejs/node/pull/22273#issuecomment-415883392 https://github.com/nodejs/node/pull/22273#issuecomment-41588...
- snek 7y agoit was not caused by a failure of the js vm though, it was caused (from what i can tell) by a failure in c++ tracing code.
- schoen 7y agoIs anybody fuzzing Python bytecodes? This sounds like a super-great application for afl.
- xapata 7y agoI think people have, but it's not clear that searching for bugs by perverse input is a great use of time. Maybe better to focus on solving currently known issues.
- poizan42 7y agoIt has always been the position of the CPython developers that using python for sandboxing is unsupported. With that in mind it doesn't really matter if you can "exploit" python with weird bytecode because you are supposed to be on the other side of the airtight hatchway[0] anyways. I don't know what the stance of other python runtimes are, but you should probably just use a sandbox at OS level which is likely to be tested far more thoroughly. [0]: https://devblogs.microsoft.com/oldnewthing/20060508-22/?p=31283 https://devblogs.microsoft.com/oldnewthing/20060508-22/?p=31...
- quietbritishjim 7y agoComments for that link: https://web.archive.org/web/20190113115213/https://blogs.msdn.microsoft.com/oldnewthing/20060508-22/?p=31283/ https://web.archive.org/web/20190113115213/https://blogs.msd...
- ChrisSD 7y agoIt's worth mentioning the interesting failure of pysandbox: > I now think that putting a sandbox directly in Python cannot be secure. To build a secure sandbox, the whole Python process must be put in an external sandbox. https://mail.python.org/pipermail/python-dev/2013-November/130132.html https://mail.python.org/pipermail/python-dev/2013-November/1...
- __s 7y agoYou don't need a fuzzer. Python doesn't do bounds checking of stack manipulation. This is not considered a bug
- tom_mellior 7y ago> In those same versions of Lua, the bytecode format is notoriously dangerous to load. Could you expand on this? Are the dangers such that they could be avoided with a bytecode verifier somewhat like Java's? Things like checking that the stack can never underflow, and that at the merge points of branches the stack always has the same depth.
- lpghatguy 7y agoHere are two slide decks talking (roughly) about real-world attacks against Lua 5.1 and 5.2's bytecode: 5.1: https://www.lua.org/wshop11/Cawley.pdf https://www.lua.org/wshop11/Cawley.pdf 5.2: https://apocrypha.numin.it/talks/lua_bytecode_exploitation.pdf https://apocrypha.numin.it/talks/lua_bytecode_exploitation.p... Lua used to have a built-in bytecode verifier as I understand it, but it never reached the point to where it was enough to safeguard the VM.
- stevekemp 7y agoCan confirm. I did some fuzz-testing of random scripting languages a while back, and reported bugs in things like GNU Awk. Here's a trivial example: /usr/bin/gawk 'for (i = ) in steve kemp rocks' I found fuzz-testing like this very very useful when writing my own BASIC interpreter, and playing with scripting languages though. It's almost magical how quickly problems are found!