9 ms·
Unsafe Zig Is Safer Than Unsafe Rust
- steveklabnik 9y agoTransmute is like, the most unsafe thing possible. It basically checks if the two things have the same size, and that's it. You're responsible for everything else. See all the warnings and suggested other ways to accomplish things with https://doc.rust-lang.org/stable/std/mem/fn.transmute.html https://doc.rust-lang.org/stable/std/mem/fn.transmute.html This is UB becuase `Foo` is not `#[repr(C)]`, in my understanding. I haven't checked if it works if you add the repr though. I don't think I'd expect it to.
- AndyKelley 9y agoI changed it to: let foo = &mut array[0] as *mut u8 as *mut Foo; (*foo).a += 1; and the IR has the same undefined behavior: https://godbolt.org/g/5Bv3FL https://godbolt.org/g/5Bv3FL
- steveklabnik 9y agoYeah I mean, to be clear, it's cool zig checks this stuff. Unsafe code is extremely dangerous, in a variety of ways. Luckily, outside of FFI, it's very rare to actually need to write it, though that does of course depend on what exactly you're doing. We hope, in the future, to basically have tooling here that can detect when you do something UB, and warn you. As we're still sorting out the memory model, etc, it's not here yet, but it's certainly on the agenda.
- bbatha 9y agoWouldn't you need to make `Foo` a union for this to be defined anyway?
- kzrdude 9y agoRust `#[repr(C)]` is about representation in memory and function call ABI for the type. C rules about casting through union are not relevant to Rust code.
- kibwen 9y ago> Transmute is like, the most unsafe thing possible. Yes, the first rule of auditing Rust unsafe blocks is that if you see someone using std::mem::transmute, you walk over and ask the author if they're really certain what they're doing. :) However, it should be noted that std::mem::transmute still has some guard rails; the real "most unsafe thing possible" is the variant of this function that does away with those guard rails: std::mem::transmute_copy. Required reading: https://doc.rust-lang.org/nightly/nomicon/transmutes.html https://doc.rust-lang.org/nightly/nomicon/transmutes.html
- kzrdude 9y agoRust is a language that offers you lots of compile time checks, and an escape hatch called unsafe that says “trust the programmer here.” Yes, it is possible—and easy—to make mistakes in the place where you have asked to be trusted, not checked. We have a big pedagogical task ahead of us in teaching safe practices for unsafe Rust, and defensive coding practices in unsafe Rust. We should also think of if we can improve unsafe Rust to be harder to misuse. There are improvements coming in compile time evaluation, and those can potentially make the compiler much stronger when it comes to detecting memory errors in unsafe code at compile time.
- dman 9y agoIt will be interesting to watch the proportion of safe/unsafe code in large Rust codebases over time.
- mcguire 9y agoAt a guess, it will increase until the necessary-but-not-currently-handled constructs are dealt with (arena-based memory management, I'm looking at you) and then decrease asymptotically. Already in Rust, if you adopt C++ STL idioms (and don't want to squeeze more performance out) and don't need to visit currently-unwrapped interfaces, you won't need unsafe at all. Rust is a very good C++ replacement.
- klodolph 9y agoThere's not only a pedagogical task here, but the Rust community must learn how to write code safely. The major difficulty here is that in general, unsafe pieces of code cannot be safely composed, even if the unsafe pieces of code are individually safe. This allows you to bypass runtime safety checks without unsafe code just by composing "safe" modules that internally use unsafe code in their implementation. This kind of problem comes up a lot. Composed atomic operations are not atomic. Composed correct threaded code is not always correct. Mixing Scheme control structures made with call/cc don't work as desired. Enabling different Haskell language extensions gets you off the deep end quickly, and some unsafe combinations are surprising (see GeneralizedNewtypeDeriving, which is considered unsafe even though it used to be safe).
- vbernat 9y agoThe x86 ABI enforces alignment of the stack to 16 bytes. Isn't that enough to make this particular problem go away?
- cesarb 9y agoNo. Nothing guarantees that the array is aligned within the stack frame, even if the stack frame is aligned. What if the compiler introduced a boolean flag (for instance, a drop flag) immediately before the array, in the same stack frame?
- mdip 9y agoGood point, here. As is often said, when the documentation says "undefined behavior", it means the compiler can do whatever it wants, including "work just fine"; and sometimes it'll cause time travel[0]. Hence the "nasal demons" lore. Often, it'll cause optimizations to be applied that would have otherwise been avoided resulting in a bug that appears to occur somewhere else and a programmer to look at the result of execution and ... if it actually continues executing ... swear a lot. These are especially fun because the problem frequently won't appear in debug builds. [0] https://blogs.msdn.microsoft.com/oldnewthing/20140627-00/?p=633/ https://blogs.msdn.microsoft.com/oldnewthing/20140627-00/?p=... - worth a read for some entertainment - basically what happens when the compiler assumes "undefined behavior" can't happen and optimizes accordingly.
- stochastic_monk 9y agoTake off every zig... for great justice!
- jfo 9y agoall your codebase are belong to us!
- viperscape 9y agoI'd like to see a C equivalent of this, just for comparisons sake to zig
- tiehuis 9y ago#include <stdint.h> #include <string.h> typedef struct { int32_t a; int32_t b; } Foo; int main(void) { uint8_t array[1024]; memset(array, 1, sizeof(array)); Foo *foo = (Foo*)(&array[0]); foo->a += 1; } Using clang 3.8.0-2. Compiling examples with `clang -S llvm-ir`. It appears that the array is aligned with the minimum ABI requirement 16 by default? May be a note of this in the standard, can't recall of the top of my head. %array = alloca [1024 x i8], align 16 ... %6 = load i32, i32* %5, align 4 ... store i32 %7, i32* %5, align 4 We can also explicitly specify the alignment required in C11. #include <stdalign.h> #include <stdint.h> #include <string.h> typedef struct { int32_t a; int32_t b; } Foo; int main(void) { uint8_t alignas(alignof(Foo)) array[1024]; memset(array, 1, sizeof(array)); Foo *foo = (Foo*)(&array[0]); foo->a += 1; } Results in the following IR. %array = alloca [1024 x i8], align 4 ... %6 = load i32, i32* %5, align 4 ... store i32 %7, i32* %5, align 4
- bjourne 9y agoOn 64bit Linux, stack frames are always aligned at 16 byte boundaries. The first 8 bytes of the frame contains the return address then there are 8 bytes of padding and then comes the stack allocations. I think the example is poorly constructed, because it is inconceivable that the address to the start of an array would not be aligned sizeof(int*) bytes.
- dbaupp 9y agoThe example is illustrative enough: all the array needs to be misaligned in practice is a small value on the stack near it, e.g. if the Rust code has `let x: u8 = 1;` inserted after the array (or, I imagine, `uint8_t x = 1;` in the C, etc.), then the array's address is odd.
- lossolo 9y agoZig looks very interesting. There is only TODO in memory section in documentation. From what I understand there is only manual memory management? I've seen there is a mention about custom allocators, any details? Any RAII like concept? or full manual memory management?
- tiehuis 9y agoMemory is manually managed, yes. We do have defer (as in go) for slightly easier resource management. Zig doesn't have a default memory allocator. Allocators instead are expected to be passed as an argument to functions as they need them. This makes it trivial to replace an allocator with something custom or use multiple different allocators within a small code block. A contrived example: const std = @import("std"); pub fn GiveMeAnInt(alloc: &std.mem.Allocator) -> %&u32 { return alloc.create(u32); } test "using two allocators" { const int1 = try GiveMeAnInt(std.heap.c_allocator); *int1 = 2; // Would usually store the allocator with the type on construction. defer std.heap.c_allocator.destroy(int1); const int2 = try GiveMeAnInt(std.debug.global_allocator); *int2 = 2; }
- jfo 9y ago"Zig does not support RAII or operator overloading because both make it very difficult to tell where function calls happen just by looking at a function body." more in the 0.1.1 release notes! http://ziglang.org/download/0.1.1/release-notes.html http://ziglang.org/download/0.1.1/release-notes.html "Zig's standard library is still very young, but the goal is for every feature that uses an allocator to accept an allocator at runtime, or possibly at either compile time or runtime." more in this wiki! https://github.com/zig-lang/zig/wiki/Why-Zig-When-There-is-Already-CPP%2C-D%2C-and-Rust%3F https://github.com/zig-lang/zig/wiki/Why-Zig-When-There-is-A...
- imtringued 9y ago>"Zig does not support RAII or operator overloading because both make it very difficult to tell where function calls happen just by looking at a function body." How about showing an error if you don't call the deconstructor manually?
- irundebian 9y ago> we are professionals, and so we do not accept undefined behavior Lol'd, tell that wannabe-elite-C-programmers.
- slaymaker1907 9y agoOne big glaring flaw IMO is that it is not really possible to just turn off certain checks as opposed to turning them all off. For instance, maybe I need to call an unsafe C api or something but could still use the borrow checker.
- burntsushi 9y agoIn Rust, the borrow checker is still enabled in unsafe blocks.
- dbaupp 9y agoAn `unsafe` block only enables extra features, it doesn't change existing behaviour of safe Rust. Specifically, it allows calling `unsafe` functions (FFI and pure Rust `unsafe` ones), dereferencing raw pointers and some minor other stuff (e.g., inline assembly, some manipulations of packed structs). The borrow checker still works on references, the trait system still enforces Send/Sync for concurrency, and the type system still requires things to have matching types. It's definitely true that having a one dimensional `unsafe` might seem unnecessarily powerful in some cases (e.g. an particular unsafe block might just need to do some pointer offsetting and dereferencing, but no FFI), but it isn't a "you're on your own" hammer.
- hashmal 9y agoThe Rust book clearly states: "It’s important to understand that unsafe doesn’t turn off the borrow checker or disable any other of Rust’s safety checks" [1] "unsafe" unlocks only 4 things: Dereferencing a raw pointer, Calling an unsafe function or method, Accessing or modifying a mutable static variable, Implementing an unsafe trait. [1] https://doc.rust-lang.org/book/second-edition/ch19-01-unsafe-rust.html https://doc.rust-lang.org/book/second-edition/ch19-01-unsafe...
- devit 9y agoI think both "as" and using "transmute" for non-exceptional circumstances are mistakes in Rust. There should instead be a bunch of type-specific cast operators that can check things like alignment and that what you intended to be a zero-extending integer cast is not in fact truncating to a smaller integer type, and so on. It's not too late to deprecate "as" and discourage using "transmute" in favor of those.
- elcritch 9y agoExactly what I was thinking while reading the OP. It seems like it'd be possible to add alignment checks either manually via cast-operators or automatically via the compiler. Rust could at a minimum display a warning "possible alignment errors" when emitting that kind of LLVM IR.
- edflsafoiewq 9y agoThis isn't about transmute or having a specific operator that checks alignment. The point is that the alignment is part of the type is zig and, to a lesser degree, it's about having the comptime machinery for zig to decide, when you offset a &align(4) u8 by an expression, whether the result should have type &align(1) u8, &align(2) u8, or &align(4) u8.
- deathanatos 9y agoI think the Rust is not how you should write such a code. Why not start with the struct, and cast to a void* or a char* when C code requires it? I.e., the buggy example becomes: #[derive(Copy, Clone, Debug)] #[repr(C)] struct Foo { a: i32, b: i32, } fn main() { let mut array = [Foo { a: 0x01010101i32, b: 0x01010101i32 }; 256]; let foo = &mut array[0]; foo.a += 1; } The unsafe section isn't even required, and the effect is the same. And I don't think this violates the spirit of his example, either. Consider the author's first link to a real-world occurrence of this: let size = mem::size_of::<FILE_NAME_INFO>(); let mut name_info_bytes = vec![0u8; size + MAX_PATH]; let res = GetFileInformationByHandleEx(handle, FileNameInfo, &mut *name_info_bytes as *mut _ as *mut c_void, name_info_bytes.len() as u32); This is again, IMO, the wrong way to do this. You should just cast a pointer to an instance of the FILE_NAME_INFO struct into a c_void; the structure will need to use #[repr(C)] and the code will still be unsafe due to the C FFI, but it will be correct (and a lot simpler). This is the same thing that you would do in C, were you to call this function: FILE_NAME_INFO file_name_info; GetFileInformationByHandleEx( handle, FileNameInfo, &file_name_info, sizeof(file_name_info), ) just in Rust.
- dbaupp 9y agoWhile the approach you suggest usually works well, it doesn't in this case: FILE_NAME_INFO[1] uses a "flexible array member"[2] (although not the C99 version of it), of requiring a dynamically sized character array in the struct's allocation, and writing directly to the memory after a struct instance. The 'WCHAR FileName[1];' field at the end of the struct is just a placeholder to allow easy access to that character array, the length 1 is a lie. [1]: https://msdn.microsoft.com/en-us/library/windows/desktop/aa364388(v=vs.85).aspx https://msdn.microsoft.com/en-us/library/windows/desktop/aa3... [2]: https://en.wikipedia.org/wiki/Flexible_array_member https://en.wikipedia.org/wiki/Flexible_array_member
- deathanatos 9y agoUgh. You're absolutely right. I never liked those even in C. So, it seems like this is relative easy to do on the stack, which is how the example does it presently. See the link below to my attempt; the stack allocation is still all safe code, still a single line. However, I presume that one will want to also create one on the heap, especially since in the example the author poses it would be a rather large stack allocation, and one might — quite reasonably — put that on the heap. My attempt is here: https://play.rust-lang.org/?gist=1c50b35941506316372da860caeceba3&version=nightly https://play.rust-lang.org/?gist=1c50b35941506316372da860cae... Couldn't avoid the unsafe for that, but, I was able to get rid of the transmute call, and transmute is a function where the warning on the tin is "this function is not just unsafe, it is radioactive". But the amount of code required still felt a bit lacking. It seems these are an area of active work[1][2] currently. I think there is still definitely a valid point that the author is hitting — that encoding more information into the program can allow the compiler to catch more classes of errors. (This is, after all, the very logic that gave us Rust.) [1]: https://github.com/rust-lang/rfcs/pull/1909 https://github.com/rust-lang/rfcs/pull/1909 [2]: https://github.com/rust-lang/rust/issues/18806 https://github.com/rust-lang/rust/issues/18806
- vivaan 9y agoClearly both Rust and Zig tackle tough problems and implement solutions that will have trade offs. I don't think the top answer to a post talking about Zig's advantages should defensively try to point out how things could be different in Rust - if only you knew exactly what to do - instead it would be nice to see more discussion about other areas where Rust is perhaps better suited than Zig. For instance, you Rust clearly handles memory/pointers better (?), while maybe Zig is easier to learn?
- didibus 9y agoHad never heard of zig. Does it also provide memory safety without a GC like Rust?
- masklinn 9y ago"Unsafe Zig" encompasses all of Zig, much like e.g. C.
- audunw 9y agoNo, not to the same extent. It attempts to make C-style memory management as safe as possible, and also make it easy to use different memory allocators, but does not attempt advanced techniques like borrow checkers. There's also a pretty good metaprogramming system, so it may be possible to implement some smart memory-management libraries. Zig is about simplicity. It's a C (and partly C++) replacement, not a Rust replacement. Think of it this way: I could easily imagine a TCC-like, dirt-simple, super-fast compiler for Zig. I'm not sure we'll ever see the same for Rust. That's nothing against Rust, just saying they have very different goals.
- anfilt 9y agoI find it so funny people are so fixed on bounds checking. A minimal run time environment is good. It's easier to port and runs faster. Further, there are more issues than bounds checking. Also a big part of it is companies don't really pay for quality software. They just care about software that works mostly made to cost. I don't see rust reducing this cost much except. First, one still has to interact with hardware, that does not fit rust's/zig's/(insert safe language) run time model. Secondly, soon as you start interacting with software out side of that model same issues apply.
- audunw 9y ago> I find it so funny people are so fixed on bounds checking. A minimal run time environment is good. It's easier to port and runs faster. Further, there are more issues than bounds checking. Bounds checking on arrays is a compile-time check in Zig. Other forms of bounds-checking can be disabled in release-mode. I don't see a single compelling reason why you wouldn't at least want bounds checking in debug mode. If you're out of bounds, something is wrong, and it's always better to get an early and precise error about it. In Zig you can take slices of arrays or pointers, which contain a pointer and a length. This is not just about safety, it's also a convenience. There's a lot of usecases where you want to pass around both a pointer and a length. Considering how many extremely serious bugs have resulted from a lack of bounds-checking, and considering the relatively low run-time overhead of doing it (especially with some decent optimizations from the compiler), I don't find it funny at all.
- humanrebar 9y ago> I don't see rust reducing this cost much except. At scale, a language with a module system will reduce cost substantially.
- cryptos 9y agoIf something is specified as "unsafe", it is implemented correctly if it is unsafe - ask Intel ;-)
- eggy 9y agoI’m finding Zig easier to learn and hold in my head so that also helps me right correct code and safe code. Zig is pretty much one man’s work and is very impressive. I’m still playing with Rust but I am using Zig as my C replacement right now.