10 ms·
I like zig but this is taking a page out of rust book and exaggerating C and C++ clang and gcc will both tell you at runtime if you go out of bounds, have an i
by ArrayBoundCheck 4y ago
I like zig but this is taking a page out of rust book and exaggerating C and C++
clang and gcc will both tell you at runtime if you go out of bounds, have an integer overflow, use after free etc. You need to turn on the sanitizer. You can't have them all on at the same time because code will be unnecessarily slow (ex: having thread sanitizer on in a single threaded app is pointless)
- lijogdfljk 4y agoWhat is the cause of all those notorious C bugs then?
- CodeSgt 4y ago> at runtime
- deleted 4y ago[deleted]
- woodruffw 4y agoNeither Clang nor GCC has perfect bounds or lifetime analysis, since the language semantics forbid it: it's perfectly legal at compile time to address at some offset into a supplied pointer, because the compiler has no way of knowing that the memory there isn't owned and initialized. Sanitizers are great; I love sanitizers. But you can't run them in production without a significant performance hit, and that's where they're needed most. I don't believe this post blows that problem out of proportion, and is correct in noting that we can solve it without runtime instrumentation and overhead.
- AlotOfReading 4y agoState of the art sanitizing is pretty consistently in the <50% overhead range (e.g. SANRAZOR), with things like UBSAN coming in under 10%. If you can't afford even that, tools like ASAP have been around for 7-ish years now to make overhead arbitrarily low by trading off increased false-negatives in hot codepaths. Yes, the "just-enable-the-compiler-flags" approach can be expensive, but the tools exist to allow most people to be sanitizing most of the time. Devs simply don't know what's available to them.
- woodruffw 4y agoI'd consider even 10% to be a significant performance hit. People scream bloody murder when CPU-level mitigations cause even 1-2% regressions. The marginal cost of mitigations when memory safe code can run without them is infinite. But let's say, for the sake of argument, that I can tolerate programs that run twice as long in production. This doesn't improve much: * I'm not going to be deploying SoTA sanitizers (SANRAZOR is currently a research artifact; it's not available in mainline LLVM as far as I can tell.) * No sanitizer that I know of guarantees that execution corresponds to memory safety. ASan famously won't detect reads of uninitialized memory (MSan will, but you can't use both at the same time), and it similarly won't detect layout-adjacent overreads/writes. That's a lot of words to say that I think sanitizers are great, but they're not a meaningful alternative to actual memory safety. Not when I can have my cake and eat it too.
- AlotOfReading 4y agoI think we basically agree. Hypothetically ideal memory safety is strictly better, but sanitizers are better than nothing for code using fundamentally unsafe languages. My personal experience is that more people are dissuaded from sanitizer usage more by hypothetical (and manageable) issues like overhead than real implementation problems.
- KerrAvon 4y agoIf you can afford a 10-50% across-the-board performance reduction, why would you not use a higher-level, actually safe language like Ruby or Python? Remember that the context of this article is Zig vs other languages, so the assumption is you’re writing new code.
- AlotOfReading 4y agoI work in real time, often safety critical environments. High level interpreted languages aren't particularly useful there. The typical options are C/C++, hardware (e.g. FPGAs), or something more obscure like Ada/Spark. But in general, sanitizers are also something you can do to legacy code to bring it closer to safety and you can turn them off for production if you absolutely, definitely need those last few percent (which few people do). It's hard to overstate how valuable all of that is. A big part of the appeal of zig is its interoperability with C and the ability to introduce it gradually. Compare to the horrible contortions you have to do with CFFI to call Python from C.
- com2kid 4y ago> it's perfectly legal at compile time to address at some offset into a supplied pointer, because the compiler has no way of knowing that the memory there isn't owned and initialized. Embedded land, everything is a flat memory map, odds are malloc isn't used at all, memory is possibly 0'd on boot. It is perfectly valid to just start walking all over memory. You have a bunch of #defines with known memory addresses in them and you can just index from there. Fun fact: Microsoft Band writes crash dumps to a known location in SRAM and because SRAM doesn't instantly lose its contents on reboot, after a crash the runtime checks for crash dump data at that known address and if present would upload the crash dump to servers for analysis and then 0 out that memory.[1] Embedded rocks! [1] There is a bit more to it to ensure we aren't just reading random data after along power off, but I wasn't part of the design, I just benefited from a 256KB RAM wearable having crash dumps that we could download debugging symbols for.
- xedrac 4y ago> So I'm not covering tools like AddressSanitizer that are intended for testing and are not recommended for production use. How is it an exaggeration when he explicitly called this out?
- ArrayBoundCheck 4y agoASAN isn't just "for testing". A lot of people went straight to the chart (like me) and it reeks of bullshit. double free is the same as use after free, null pointer dereference is essentially the same as type confusion since a nullable pointer is confused with a non null pointer, invalid stack read/write is the same as array out of bounds (or invalid pointers), etc I also never heard of a data race existing without a race condition existing. That's a pointless metric like many of the above I mentioned
- hyperpape 4y agoCan you explain why, in spite of the fact that (according to you) C & C++ aren't that unsafe, critical projects like Chromium can't get this right? https://twitter.com/pcwalton/status/1539112080590217217 https://twitter.com/pcwalton/status/1539112080590217217 Is the Project Zero team just too lazy to remind Chromium to use sanitizers?
- uecker 4y agoI think the big question is, whether two teams writing software on a fixed budget using Rust or C using modern tools and best practices would end up with a safer product. I think this is not clear at all.
- uecker 4y ago(Ok, I should read the text before sending.)
- pcwalton 4y agoPeople have done just that with, for example, Firefox components and found that yes, Rust gives you a safer product.
- uecker 4y agoDo you have a pointer? I know they rewrote Firefox components, but I am not aware of a real study with a 1:1 comparison.
- lmm 4y agoI think it's very clear for anything other than a no-true-Scotsman definition of "modern tools and best practices" (which is sadly the only one that seems to exist).
- jerf 4y agoWhile I'm generally in favor of the proposition that C++ is an intrinsically dangerous language, pointing at one of the largest possible projects that uses it isn't the best argument. If I pushed a button and magically for free Chrome was suddenly in 100% pure immaculate Rust, I'm sure it would still have many issues and problems that few other projects would have, just due to its sheer scale. I would still consider it an open question/problem as to whether Rust can scale up to that size and still be something that humans can modify. I could make a solid case that the difficulty of working in Rust would very accurately reflect a true and essential difficulty of working at that scale in general, but it could still be a problem. (Also Rust defenders please note I'm not saying Rust can't work at that scale. I'm just saying, it's a very big scale and I think it's an open problem. My personal opinion and gut say yes, it shouldn't be any worse than it has to be because of the sheer size (that is, the essential complexity is pretty significant no matter what you do), but I don't know that.)
- wyldfire 4y agoOne interesting distinction is that it sounds as if - for Zig, this is a language feature and not a toolchain feature. Although if there's only one toolchain for zig maybe that's a distinction-without-a-difference. At least it's not opt-in, that's really nice. Believe it or not, there are lots of people who write and debug C/C++ code who don't know about sanitizers or they know about it and never decide to use them.
- throwawaymaths 4y agoI think it would be interesting to see zig move towards annotation-based compile time lifetime checking plugin (ideally in-toolchain, but alternatively as a library). You could choose to turn it on selectively for security-critical pathways, turn it off for "trust me" functions, or, do it on "not every recompilation", as desired.
- pjmlp 4y agoThe irony being that lint exists since 1979, and already using a static analyser would be a bing improvement in some source bases.
- kubanczyk 4y agoWhoa the username checks out perfectly.
- ArrayBoundCheck 4y agoHaha yes. I love knowing I'm in bounds but unfortunately saying anything about C++ (that isn't a criticism) is out of bounds and my comment got downvoted enough that I don't feel like saying more
- masklinn 4y ago> clang and gcc will both tell you at runtime if you go out of bounds [...] You can't have them all on at the same time because code will be unnecessarily slow Yeah, so clang and gcc don't actually tell you at runtime if you go out of bounds. How many program ship production binaries with asan or ubsan enabled, to say nothing of msan or tsan? Also you can't have them all on at the same time because they're not necessarily compatible with one another[0], you literally can't run with both asan and msan, or asan and tsan. [0] https://github.com/google/sanitizers/issues/1039 https://github.com/google/sanitizers/issues/1039
- pjmlp 4y agoQuite a few subsystems on Android, but that is about it. https://source.android.com/devices/tech/debug/hwasan https://source.android.com/devices/tech/debug/hwasan