29 ms·
Low-Level Optimization with Zig
- flohofwoe 1y ago> I love Zig for it's verbosity. I love Zig too, but this just sounds wrong :) For instance, C is clearly too sloppy in many corners, but Zig might (currently) swing the pendulum a bit too far into the opposite direction and require too much 'annotation noise', especially when it comes to explicit integer casting in math expressions (I wrote about that a bit here: https://floooh.github.io/2024/08/24/zig-and-emulators.html https://floooh.github.io/2024/08/24/zig-and-emulators.html). When it comes to performance: IME when Zig code is faster than similar C code then it is usually because of Zig's more aggressive LLVM optimization settings (e.g. Zig compiles with -march=native and does whole-program-optimization by default, since all Zig code in a project is compiled as a single compilation unit). Pretty much all 'tricks' like using unreachable as optimization hints are also possible in C, although sometimes only via non-standard language extensions. C compilers (especially Clang) are also very aggressive about constant folding, and can reduce large swaths of constant-foldable code even with deep callstacks, so that in the end there often isn't much of a difference to Zig's comptime when it comes to codegen (the good thing about comptime is of course that it will not silently fall back to runtime code - and non-comptime code is still of course subject to the same constant-folding optimizations as in C - e.g. if a "pure" non-comptime function is called with constant args, the compiler will still replace the function call with its result). TL;DR: if your C code runs slower than your Zig code, check your C compiler settings. After all, the optimization heavylifting all happens down in LLVM :)
- skywal_l 1y agoMaybe with the new x86 backend we might see some performance differences between C and Zig that could definitely be attributed solely to the Zig project.
- saagarjha 1y agoI would be (pleasantly) surprised if Zig could beat LLVM's codegen.
- Zambyte 1y agoSo would the Zig team. AFAIK, they don't plan to (and have said this in interviews). The plan is for super fast compilation and incremental compilation. I think the homegrown backend is mainly for debug builds.
- Cloudef 1y agoThe backends do already have some simple optimizations. Of course focus is debug builds and speed, but long term goal is for them to be competitive as well.
- Retro_Dev 1y agoAhh perhaps I need to clarify: I don't love the noise of Zig, but I love the ability to clearly express my intent and the detail of my code in Zig. As for arithmetic, I agree that it is a bit too verbose at the moment. Hopefully some variant of https://github.com/ziglang/zig/issues/3806 https://github.com/ziglang/zig/issues/3806 will fix this. I fully agree with your TL;DR there, but would emphasize that gaining the same optimizations is easier in Zig due to how builtins and unreachable are built into the language, rather than needing gcc and llvm intrinsics like __builtin_unreachable() - https://gcc.gnu.org/onlinedocs/gcc-4.5.0/gcc/Other-Builtins.html#Other-Builtins https://gcc.gnu.org/onlinedocs/gcc-4.5.0/gcc/Other-Builtins.... It's my dream that LLVM will improve to the point that we don't need further annotation to enable positive optimization transformations. At that point though, is there really a purpose to using a low level language?
- flohofwoe 1y agoYeah indeed. Having access to all those 'low-level tweaks' without having to deal with non-standard language extensions which are different in each C compiler (if supported at all) is definitely a good reason to use Zig. One thing I was wondering, since most of Zig's builtins seem to map directly to LLVM features, if and how this will affect the future 'LLVM divorce'.
- Retro_Dev 1y agoGood question! The TL;DR as I understand it is that it won't matter too much. For example, the self-hosted x86_64 backend (which is coincidentally becoming default for debugging on linux right now - https://github.com/ziglang/zig/pull/24072 https://github.com/ziglang/zig/pull/24072) has full support for most (all?) builtins. I don't think that we need to worry about that. It's an interesting question about how Zig will handle additional builtins and data representations. The current way I understand it is that there's an additional opt-in translation layer that converts unsupported/complicated IR to IR which the backend can handle. This is referred to as the compiler's "Legalize" stage. It should help to reduce this issue, and perhaps even make backends like https://github.com/xoreaxeaxeax/movfuscator https://github.com/xoreaxeaxeax/movfuscator possible :)
- 1y ago
- messe 1y agoWith regard to the casting example, you could always wrap the cast in a function: fn signExtendCast(comptime T: type, x: anytype) T { const ST = std.meta.Int(.signed, @bitSizeOf(T)); const SX = std.meta.Int(.signed, @bitSizeOf(@TypeOf(x))); return @bitCast(@as(ST, @as(SX, @bitCast(x)))); } export fn addi8(addr: u16, offset: u8) u16 { return addr +% signExtendCast(u16, offset); } This compiles to the same assembly, is reusable, and makes the intent clear.
- flohofwoe 1y agoYes, that's a good solution for this 'extreme' example. But in other cases I think the compiler should make better use of the available information to reduce 'redundant casting' when narrowing (like the fact that the result of `a & 15` is guaranteed to fit into an u4 etc...). But I know that the Zig team is aware of those issues, so I'm hopeful that this stuff will improve :)
- deleted 1y ago[deleted]
- hansvm 1y agoThis is something I used to agree with, but implicit narrowing is dangerous, enough so that I'd rather be more explicit most of the time nowadays. The core problem is that you're changing the semantics of that integer as you change types, and if that happens automatically then the compiler can't protect you from typos, vibe-coded defects, or any of the other ways kids are generating almost-correct code nowadays. You can mitigate that with other coding patterns (like requiring type parameters in any potentially unsafe arithmetic helper functions and banning builtins which aren't wrapped that way), but under the swiss cheese model of error handling it still massively increases your risky surface area. The issue is more obvious on the input side of that expression and with a different mask. E.g.: const a: u64 = 42314; const even_mask: u4 = 0b0101; a & even_mask; Should `a` be lowered to a u4 for the computation, or `even_mask` promoted, or however we handle the internals have the result lowered sometimes to a u4? Arguably not. The mask is designed to extract even bit indices, but we're definitely going to only extract the low bits. The only safe instance of implicit conversion in this pattern is when you intend to only extract the low bits for some purpose. What if `even_mask` is instead a comptime_int? You still have the same issue. That was a poor use of comptime ints since now that implicit conversion will always happen, and you lost your compiler errors when you misuse that constant. Back to your proposal of something that should always be safe: implicitly lowering `a & 15` to a u4. The danger is in using it outside its intended context, and given that we're working with primitive integers you'll likely have a lot of functions floating around capable of handling the result incorrectly, so you really want to at least use the _right_ integer type to have a little type safety for the problem. For a concrete example, code like that (able to be implicitly lowered because of information obvious to the compiler) is often used in fixed-point libraries. The fixed-point library though does those sorts of operations with the express purpose of having zeroed bits in a wide type to be able to execute operations without loss of precision (the choice of what to do for the final coalescing of those operations when precision is lost being a meaningful design choice, but it's irrelevant right this second). If you're about to do any nontrivial arithmetic on the result of that masking, you don't want to accidentally put it in a helper function with a u4 argument, but with implicit lowering that's something that has no guardrails. It requires the programmer to make zero mistakes. That example might seem a little contrived, and this isn't something you'll run into every day, but every nontrivial project I've worked on has had _something_ like that, where implicit narrowing is extremely dangerous and also extremely easy to accidentally do. What about the verbosity? IMO the point of verbosity is to draw your attention to code that you should be paying attention to. If you're in a module where implicit casting would be totally fine, then make a local helper function with a short name to do the thing you want. Having an unsafe thing be noisy by default feels about right though.
- knighthack 1y agoI'm not sure why allowances are made for Zig's verbosity, but not Go's. What's good for the goose should be good for the gander.
- nurbl 1y agoI think a better word may be "explicitness". Zig is sometimes verbose because you have to spell things out. Can't say much about Go, but it seems it has more going on under the hood.
- ummonk 1y agoZig's verbosity goes hand in hand with a strong type system and a closeness to the hardware. You don't get any such benefits from Go's verbosity.
- Zambyte 1y agoFWIW Zig has error handling that is nearly semantically identical to Go (errors as return values, the big semantic difference being tagged unions instead of multiple return values for errors), but wraps the `if err != nil { return err}` pattern in a single `try` keyword. That's the verbosity that I see people usually complaining about in Go, and Zig addresses it.
- kbolino 1y agoThe way Zig addresses it also discards all of the runtime variability too. In Go, an error can say something like unmarshaling struct type Foo: in field Bar int: failed to parse value "abc" as integer Whereas in Zig, an error can only say something that's known at compile time, like IntParse, and you will have to use another mechanism (e.g. logging) to actually trace the error.
- metaltyphoon 1y agoYep. Errors carry no context whatsoever and you have no idea where they came from.
- Zambyte 1y agoRegarding the explicit integer casting, it seems like there is some cleanup that will be coming soon: https://ziggit.dev/t/short-math-notation-casting-clarity-of-math-expressions/10008/19?u=zambyte https://ziggit.dev/t/short-math-notation-casting-clarity-of-...
- titzer 1y agoZig has some interesting ideas, and I thought the article was going to be more on the low-level optimizations, but it turned out to be "comptime and whole program compilation are great". And I agree. Virgil has had the full language available at compile time, plus whole program compilation since 2006. But Virgil doesn't target LLVM, so speed comparisons end up being a comparison between two compiler backends. Virgil leans heavily into the reachability and specialization optimizations that are made possible by the compilation model. For example it will aggressively devirtualize method calls, remove unreachable fields/objects, constant-promote through fields and heap objects, and completely monomorphize polymorphic code.
- int_19h 1y agoI rather suspect that the pendulum will swing rather strongly towards more verbose and explicit languages in general in the upcoming years solely because it makes things easier for AI. (Note that this is orthogonal to whether and to what extent use of AI for coding is a good idea. Even if you believe that it's not, the fact is that many devs believe otherwise, and so languages will strive to accommodate them.)
- KingOfCoders 1y agoI do love the allocator model of Zig, I would wish I could use something like an request allocator in Go instead of GC.
- usrnm 1y agoCustom allocators and arenas are possible in go and even do exist, but they ara just very unergonomic and hard to use properly. The language itself lacks any way to express and enforce ownership rules, you just end up writing C with a slightly different syntax and hoping for the best. Even C++ is much safer than go without GC
- KingOfCoders 1y agoThey are not integrated in all libraries, so for me they don't exist.
- saagarjha 1y ago> As an example, consider the following JavaScript code…The generated bytecode for this JavaScript (under V8) is pretty bloated. I don't think this is a good comparison. You're telling the compiler for Zig and Rust to pick something very modern to target, while I don't think V8 does the same. Optimizing JITs do actually know how to vectorize if the circumstances permit it. Also, fwiw, most modern languages will do the same optimization you do with strings. Here's C++ for example: https://godbolt.org/z/TM5qdbTqh https://godbolt.org/z/TM5qdbTqh
- Retro_Dev 1y agoYou can change the `target` in those two linked godbolt examples for Rust and Zig to an older CPU. I'm sorry I didn't think about the limitations of the JS target for that example. As for your link, It's a good example of what clang can do for C++ - although I think that the generated assembly may be sub-par, even if you factor in zig compiling for a specific CPU here. I would be very interested to see a C++ port of https://github.com/RetroDev256/comptime_suffix_automaton https://github.com/RetroDev256/comptime_suffix_automaton though. It is a use of comptime that can't be cleanly guessed by a C++ compiler.
- saagarjha 1y agoI just skimmed your code but I think C++ can probably constexpr its way through. I understand that's a little unfair though because C++ is one of the only other languages with a serious focus on compile-time evaluation.
- vanderZwan 1y agoIn general it's a bit of an apples to fruit salad comparison, albeit one that is appropriate to highlight the different use-cases of JS and Zig. The Zig example uses an array with a known type of fixed size, the JS code is "generic" at run time (x and y can be any object). Which, fair enough, is something you'd have to pay the cost for in JS. Ironically though in this particular example one actually would be able to do much better when it comes to communicating type information to the JIT: ensure that you always call this function with Float64Arrays of equal size, and the JIT will know this and produce a faster loop (not vectorized, but still a lot better). Now, one rarely uses typed arrays in practice because they're pretty heavy to initialize so only worth it if one allocates a large typed array one once and reuses them a lot aster that, so again, fair enough! One other detail does annoy me a little bit: the article says the example JS code is pretty bloated, but I bet that a big part of that is that the JS JIT can't guarantee that 65536 equals the length of the two arrays so will likely insert a guard. But nobody would write a for loop that way anyway, they'd write it as i < x.length, for which the JIT does optimize at least one array check away. I admit that this is nitpicking though.
- uecker 1y agoYou don't really need comptime to be able to inline and unroll a string comparison. This also works in C: https://godbolt.org/z/6edWbqnfT https://godbolt.org/z/6edWbqnfT (edit: fixed typo)
- Retro_Dev 1y agoYep, you are correct! The first example was a bit too simplistic. A better one would be https://github.com/RetroDev256/comptime_suffix_automaton https://github.com/RetroDev256/comptime_suffix_automaton Do note that your linked godbolt code actually demonstrates one of the two sub-par examples though.
- uecker 1y agoI haven't looked at the more complex example, but the second issue is not too difficult to fix: https://godbolt.org/z/48T44PvzK https://godbolt.org/z/48T44PvzK For complicated things, I haven't really understood the advantage compared to simply running a program at build time.
- Cloudef 1y agoTo be honest your snippet isn't really C anymore by using a compiler builtin. I'm also annoyed by things like `foo(int N, const char x[N])` which compilation vary wildly between compilers (most ignore them, gcc will actually try to check if the invariants if they are compile time known) > I haven't really understood the advantage compared to simply running a program at build time. Since both comptime and runtime code can be mixed, this gives you a lot of safety and control. The comptime in zig emulates the target architecture, this makes things like cross-compilation simply work. For program that generates code, you have to run that generator on the system that's compiling and the generator program itself has to be aware the target it's generating code for.
- uecker 1y agoIt also works with memcpy from the library: https://godbolt.org/z/Mc6M9dK4M https://godbolt.org/z/Mc6M9dK4M I just didn't feel like burdening godbolt with an inlclude. I do not understand your criticism of [N]. This gives compiler more information and catches errors. This is a good thing! Who could be annoyed by this: https://godbolt.org/z/EeadKhrE8 https://godbolt.org/z/EeadKhrE8 (of course, nowadays you could also define a descent span type in C) The cross-compilation argument has some merit, but not enough to warrant the additional complexity IMHO. Compile-time computation will also have annoying limitations and makes programs more difficult to understand. I feel sorry for everybody who needs to maintain complex compile time code generation. Zig certainly does it better than C++ but still..
- justmarc 1y agoOptimization matters, in a huge way. Its effects are compounded by time.
- sgt 1y agoOnly if the software ends up being used.
- el_pollo_diablo 1y ago> In fact, even state-of-art compilers will break language specifications (Clang assumes that all loops without side effects will terminate). I don't doubt that compilers occasionally break language specs, but in that case Clang is correct, at least for C11 and later. From C11: > An iteration statement whose controlling expression is not a constant expression, that performs no input/output operations, does not access volatile objects, and performs no synchronization or atomic operations in its body, controlling expression, or (in the case of a for statement) its expression-3, may be assumed by the implementation to terminate.
- tialaramex 1y agoC++ says (until the future C++ 26 is published) all loops, but as you noted C itself does not do this, only those "whose controlling expression is not a constant expression". Thus in C the trivial infinite loop for (;;); is supposed to actually compile to an infinite loop, as it should with Rust's less opaque loop {} -- however LLVM is built by people who don't always remember they're not writing a C++ compiler, so Rust ran into places where they're like "infinite loop please" and LLVM says "Aha, C++ says those never happen, optimising accordingly" but er... that's the wrong language.
- el_pollo_diablo 1y agoSure, that sort of language-specific idiosyncrasy must be dealt with in the compiler's front-end. In TFA's C example, consider that their loop while (i <= x) { // ... } just needs a slight transformation to while (1) { if (i > x) break; // ... } and C11's special permission does not apply any more since the controlling expression has become constant. Analyzes and optimizations in compiler backends often normalize those two loops to a common representation (e.g. control-flow graph) at some point, so whatever treatment that sees them differently must happen early on.
- pjmlp 1y agoIn theory, in practice it depends on the compiler. It is no accident that there is ongoing discussion that clang should get its own IR, just like it happens with the other frontends, instead of spewing LLVM IR directly into the next phase.
- sidjduij 1y ago[flagged]
- dustbunny 1y agoWhat interests me most by zig is the ease of the build system, cross compilation, and the goal of high iteration speed. I'm a gamedev, so I have performance requirements but I think most languages have sufficient performance for most of my requirements so it's not the #1 consideration for language choice for me. I feel like I can write powerful code in any language, but the goal is to write code for a framework that is most future proof, so that you can maintain modular stuff for decades. C/C++ has been the default answer for its omnipresent support. It feels like zig will be able to match that.
- FlyingSnake 1y agoI recently, for fun, tried running zig on an ancient kindle device running stripped down Linux 4.1.15. It was an interesting experience and I was pleasantly surprised by the maturity of Zig. Many things worked out of the box and I could even debug a strange bug using ancient GDB. Like you, I’m sold on Zig too. I wrote about it here: https://news.ycombinator.com/item?id=44211041 https://news.ycombinator.com/item?id=44211041
- osigurdson 1y agoI've dabbled in Rust, liked it, heard it was bad so kind of paused. Now trying it again and still like it. I don't really get why people hate it so much. Ugly generics - same thing in C# and Typescript. Borrow checker - makes sense if you have done low level stuff before.
- int_19h 1y agoIf you don't happen to come across some task that implies a data model that Rust is actively hostile towards (e.g. trees with backlinks, or more generally any kind of graph with cycles in it), borrow checker is not much of a hassle. But the moment you hit something like that, it becomes a massive pain, and requires either "unsafe" (which is strictly more dangerous than even C, never mind Zig) or patterns like using indices instead of pointers which are counter to high performance and effectively only serve to work around the borrow checker to shut it up.
- curtisszmania 1y ago[dead]
- timewizard 1y agoThat for loop syntax is horrendous. So I have two lists, side by side, and the position of items in one list matches positions of items in the other? That just makes my eyes hurt. I think modern languages took a wrong turn by adding all this "magic" in the parser and all these little sigils dotted all around the code. This is not something I would want to look at for hours at a time.
- int_19h 1y agoSuch arrays are an extremely common pattern in low-level code regardless of language, and so is iterating them in parallel, so it's natural for Zig to provide a convenient syntax to do exactly that in a way that makes it clear what's going on (which IMO it does very well). Why does it make your eyes hurt?
- timewizard 1y agoIt looks to me like: for (one, two, three) |uno, dos, tres| { ... } My eyes have to bounce back and forth between the two lists. When the identifiers are longer than this example it increases eye strain. Maybe it's better when you wrote it and understand it, but trying to grok someone else's code, it feels like an obstacle to me.
- csjh 1y ago> High level languages lack something that low level languages have in great adundance - intent. Is this line really true? I feel like expressing intent isn't really a factor in the high level / low level spectrum. If anything, more ways of expressing intent in more detail should contribute towards them being higher level.
- wk_end 1y agoI agree with you and would go further: the fundamental difference between high-level and low-level languages is that in high-level languages you express intent whereas in low-level languages you are stuck resorting to expressing underlying mechanisms.
- jeroenhd 1y agoI think this isn't referring to intent as in "calculate the tax rate for this purchase" but rather "shift this byte three positions to the left". Less about what you're trying to accomplish, and more about what you're trying to make the machine do. Something like purchase.calculate_tax().await.map_err(|e| TaxCalculationError { source: e })?; is full of intent, but you have no idea what kind of machine code you're going to end up with.
- csjh 1y agoMaybe, but from the author's description, it seems like the interpretation of intent that they want is to generally give the most information possible to the compiler, so it can do its thing. I don't see why the right high level language couldn't give the compiler plenty of leeway to optimize.
- raincole 1y agoIn other words, high-level languages express high-level intents, while low-level languages express low-level intents. In yet other words, tautology.
- 9d 1y ago> People will still mistakenly say "C is faster than Python", when the language isn't what they are benchmarking. Yeah but some language features are disproportionately more difficult to optimize. It can be done, but with the right language, the right concept is expressed very quickly and elegantly, both by the programmer and the compiler.
- WalterBright 1y ago> Rust's memory model allows the compiler to always assume that function arguments never alias. You must manually specify this in Zig. I've avoided such manual specification of aliasing because: 1. few people understand it 2. using it erroneously can result in baffling bugs in your code
- WalterBright 1y ago> The flexibility of Zig's comptime has resulted in some rather nice improvements in other programming languages. Compile time function execution and functions with constant arguments were introduced in D in 2007, and resulted in many other languages adopting something similar. https://dlang.org/spec/function.html#interpretation https://dlang.org/spec/function.html#interpretation
- kamma4434 1y agoI know nothing of Zig, but I worked long enough in lisp to know that the best macros are the ones you don’t write. They are wonderful but they have just as many drawbacks, and don’t compose nicely.