24 ms·
Zlib-rs is faster than C
- IshKebab 2y agoIt's barely faster. I would say it's more accurate to say it's as fast as C, which is still a great achievement.
- throwaway48476 2y agoBut it is faster. The closer to theoretical maximum the smaller the gains become.
- mananaysiempre 2y agoZlib-ng is between a couple and multiple times away from the state of the art[1], it’s just that nobody has yet done the (hard) work of adjusting libdeflate[2] to a richer API than “complete buffer in, complete buffer out”. [1] https://github.com/zlib-ng/zlib-ng/issues/1486 https://github.com/zlib-ng/zlib-ng/issues/1486 [2] https://github.com/ebiggers/libdeflate https://github.com/ebiggers/libdeflate
- qweqwe14 2y ago"Barely" or not is completely irrelevant. The fact is that it's measurably faster than the C implementation with the more common parameters. So the point that you're trying to make isn't clear tbh. Also I'm pretty sure that the C implementation had more man hours put into it than the Rust one.
- bee_rider 2y agoI think that would be really hard to measure. In particular, for this sort of very optimized code, we’d want to separate out the time spent designing the algorithms (which the Rust version benefits from as well). Actually I don’t think that is possible at all (how will we separate out time spent coding experiments in C, then learning from them). Fortunately these “which language is best” SLOC measuring contests are just frivolous little things that only silly people take seriously.
- ajross 2y agoIt's... basically written in C. I'm no expert on zlib/deflate or related algorithms, but digging around https://github.com/trifectatechfoundation/zlib-rs/ https://github.com/trifectatechfoundation/zlib-rs/ almost every block with meaningful logic is marked unsafe. There's raw allocation management, raw slicing of arrays, etc... This code looks and smells like C, and very much not like rust. I don't know that this is a direct transcription of the C code, but if you were to try something like that this is sort of what it would look like. I think there's lots of value in wrapping a raw/unsafe implementation with a rust API, but that's not quite what most people think of when writing code "in rust".
- hermanradtke 2y ago> basically written in C Unsafe Rust still has to conform to many of Rust’s rules. It is meaningfully different than C.
- est31 2y agoIt has also way less tooling available than C to analyze its safety.
- nindalf 2y agoThe number of tools matters less than the quality of the tools. Rust’s inherent guarantees + miri + software verification tools mean that in practice Rust code, even with unsafe, ends up being higher quality.
- wyager 2y agoMiri is better than any C tool I'm aware of for runtime UB detection.
- est31 2y agoMiri is the closest to a UB specification for Rust that there is, coming in the form of a tool so you can run it. It's really cool but Valgrind, which is a C tool that also supports Rust, also supports Rust code that calls to C and that does I/O, both pretty common things for programs to do.
- johnisgood 2y ago"faster than C" almost always boils down to different designs, implementations, algorithms, etc. Perhaps it is faster than already-existing implementations, sure, but not "faster than C", and it is odd to make such claims.
- oneshtein 2y ago... because by "C" we mean handwritten inline assembler. Typical realworld C code uses \0 terminated strings and strlen() with O(len^2) complexity.
- qweqwe14 2y agoThe fact that it's faster than the C implementation that surely had more time and effort put into it doesn't look good for C here.
- johnisgood 2y agoIt says absolutely nothing about the programming language though.
- acdha 2y agoDoesn’t it say something if Rust programmers routinely feel more comfortable making aggressive optimizations and have more time to do so? We maintain code for longer than the time taken to write the first version and not having to pay as much ongoing overhead cost is worth something.
- jason-johnson 2y agoHow can it not? Experts in C taking longer to make a slower and less safe implementation than experts in Rust? It's not conclusive but it most certainly says something about the language.
- johnisgood 2y ago> Experts in C taking longer to make a slower and less safe implementation than experts in Rust? How do you know this exactly?
- kahlonel 2y agoYou mean the implementation is faster than the one in C. Because nothing is “faster than C”.
- arlort 2y agoTachyons?
- mkoubaa 2y agoC after an optimizing compiler has chewed through it is faster than C
- Jaxan 2y agoOf course many things can be faster than C, because C is very far from modern hardware. If you compile with optimisation flags, the generated machine code looks nothing like what you programmed in C.
- dijit 2y agoThe kind of code you can write in rust can indeed be faster than C, but someone will wax poetic about how anything is possible in C and they would be valid. The major reason that rust can be faster than C though, is because due to the way the compiler is constructed, you can lean on threading idiomatically. The same can be true for Go, coroutines vs no coroutines in some cases is going to be faster for the use case. You can write these things to be the same speed or even faster in C, but you won’t, because it’s hard and you will introduce more bugs per KLOC in C with concurrency vs Go or Rust.
- deleted 2y ago[deleted]
- kccqzy 2y agoWhile AI can certainly produce code that's faster or otherwise better than human-written code, I am utterly skeptical of LLMs doing that. My own experience with LLM is that humans can do everything they do, but they are faster than humans. I believe we should look at non-LLM AI technologies for going beyond what a skilled human programmer can expect to do. The most famous example of AI doing that is https://www.nature.com/articles/s41586-023-06004-9 https://www.nature.com/articles/s41586-023-06004-9 where no LLM is involved.
- YZF 2y agoI found out I already know Rust: unsafe { let x_tmp0 = _mm_clmulepi64_si128(xmm_crc0, crc_fold, 0x10); xmm_crc0 = _mm_clmulepi64_si128(xmm_crc0, crc_fold, 0x01); xmm_crc1 = _mm_xor_si128(xmm_crc1, x_tmp0); xmm_crc1 = _mm_xor_si128(xmm_crc1, xmm_crc0); Kidding aside, I thought the purpose of Rust was for safety but the keyword unsafe is sprinkled liberally throughout this library. At what point does it really stop mattering if this is C or Rust? Presumably with inline assembly both languages can emit what is effectively the same machine code. Is the Rust compiler a better optimizing compiler than C compilers?
- Filligree 2y agoThe usual answer is: You only need to verify the unsafe blocks, not every block. Though 'unsafe' in Rust is actually even less safe than regular C, if a bit more predictable, so there's a crossover point where you really shouldn't have bothered. The Rust compiler is indeed better than the C one, largely because of having more information and doing full-program optimisation. A `vec_foo = vec_foo.into_iter().map(...).collect::Vec<foo>`, for example, isn't going to do any bounds checks or allocate.
- johnisgood 2y agoI have been told that "unsafe" affects code outside of that block, but hopefully steveklabnik may explain it better (again). > isn't going to do any bounds checks or allocate. You need to add explicit bounds check or explicitly allocate in C though. It is not there if you do not add it yourself.
- LegionMammal978 2y ago> I have been told that "unsafe" affects code outside of that block, but hopefully stevelabnik may explain it better (again). Poorly-written unsafe code can have effects extending out into safe code. But correctly-written unsafe code does not have any effects on safe code w.r.t. memory safety. So to ensure memory safety, you just have to verify the correctness of the unsafe code (and any helper functions, etc., it depends on), rather than the entire codebase. Also, some forms of unsafe code are far less dangeous than others in practice. E.g., most of the SIMD functions are practically safe to call in every situation, but they all have 'unsafe' slapped on them due to being intrinsics. > You need to add explicit bounds check or explicitly allocate in C though. It is not there if you do not add it yourself. Unfortunately, you do need to allocate a new buffer in C if you change the type of the elements. The annoying side of strict aliasing is that every buffer has a single type that's set in stone for all time. (Unless you preemptively use unions for everything.)
- water9 2y ago[flagged]
- berkes 2y agoThis sounds embittered. But if it isn't that, what problem with rust do you see?
- qweqwe14 2y ago> Add a borrow checker to C++ and put rust to bed once and for all Ah yes, C++ is just one safety feature away from replacing Rust, surely, any moment now. The bizzare world C++ fanboys live in. Every single person that had been writing C++ for a while and isn't a victim of Stockholm syndrome would be happy when C++ is put to bed once and for all. It's a horrible language only genuinely enjoyed by bad programmers.
- cb321 2y agoI think this may not be a very high bar. zippy in Nim claims to be about 1.5x to 2.0x faster than zlib: https://github.com/guzba/zippy https://github.com/guzba/zippy I think there are also faster zlib's around in C than the standard install one, such as https://github.com/ebiggers/libdeflate https://github.com/ebiggers/libdeflate (EDIT: also mentioned elsethread https://news.ycombinator.com/item?id=43381768 https://news.ycombinator.com/item?id=43381768 by mananaysiempre) zlib itself seems pretty antiquated/outdated these days, but it does remain popular, even as a basis for newer parallel-friendly formats such as https://www.htslib.org/doc/bgzip.html https://www.htslib.org/doc/bgzip.html
- hinkley 2y agoZlib is unapologetically written to be portable rather than fast. It is absolutely no wonder that a Rust implementation would be faster. It runs on a pathetically small number of systems by contrast. This is not a dig at Rust, it’s an acknowledgement of how many systems exist out there, once you include embedded, automotive, aerospace, telecom, industrial control systems, and mainframes. Richard Hipp denounces claims that SQLite is the widest-used piece of code in the world and offers zlib as a candidate for that title, which I believe he is entirely correct about. I’ve been consciously using it for almost thirty years, and for a few years before that without knowing I was.
- maccard 2y agoExcept this comparison isn’t against zlib, it’s against zlib-ng [0]. The readme states: > The result is a better performing and easier to maintain zlib-ng. So they’re comparing a first pass rewrite against a variation of zlib designed for performance [0] https://github.com/zlib-ng/zlib-ng https://github.com/zlib-ng/zlib-ng
- lern_too_spel 2y agoThey're comparing against zlib-ng, not zlib. zlib-ng is more than twice as fast as zlib for decompression. https://github.com/zlib-ng/zlib-ng/discussions/871 https://github.com/zlib-ng/zlib-ng/discussions/871 libdeflate is not zlib compatible. It doesn't support streaming decompression.
- jrockway 2y agoChromium is kind of stuck with zlib because it's the algorithm that's in the standards, but if you're making your own protocol, you can do even better than this by picking a better algorithm. Zstandard is faster and compresses better. LZ4 is much faster, but not quite as small. Some reading: https://jolynch.github.io/posts/use_fast_data_algorithms/ https://jolynch.github.io/posts/use_fast_data_algorithms/ (As an aside, at my last job container pushes / pulls were in the development critical path for a lot of workflows. It turns out that sha256 and gzip are responsible for a lot of the time spent during container startup. Fortunately, Zstandard is allowed, and blake3 digests will be allowed soon.)
- jeffbee 2y agoYeah I just discovered this a few days ago. All the docker-era tools default to gzip but if using, say, bazel rules_oci instead of rules_docker you can turn on zstd for large speedups in push/pull time.
- jeroenhd 2y ago`Content-Encoding: zstd` was added to Chromium a while ago: https://chromestatus.com/feature/6186023867908096 https://chromestatus.com/feature/6186023867908096 You can still use deflate for compression, but Brotli and Zstd have been available in all modern browsers for quite some time.
- amaranth 2y agoSafari doesn't support zstd, that means if you want to use it you have to support multiple formats.
- j16sdiz 2y agoChromium supports brotli and zstd
- cesarb 2y ago> Zstandard is faster and compresses better. However, keep in mind that zstd also needs much more memory. IIRC, it uses by default 8 megabytes as its buffer size (and can be configured to use many times more than that), while zlib uses at most 32 kilobytes, allowing it to run even on small 16-bit processors.
- amorio2341 2y ago[flagged]
- akagusu 2y agoBravo. Now Rust has its existence justified.
- CyberDildonics 2y agoIf you're dealing with a compiled system language the language is going to make almost no difference in speed, especially if they are all being optimized by LLVM. An optimized version that controls allocations, has good memory access patterns, uses SIMD and uses multi-threading can easily be 100x faster or more. Better memory access alone can speed a program up 20x or more.
- 1vuio0pswjnm7 2y agoWhich library compiles faster. Which library has fewer dependencies. Is each library the same size. Which one is smaller.
- rnijveld 2y agoI would argue compile time changes don't matter much, as the amount of data going through zlib all across the world is so large, that any performance gain should more than compensate any additional compilation time (and zlib-rs compiles in a couple of seconds anyway on my laptop). As for dependencies: zlib, zlib-ng and zlib-rs all obviously need some access to OS APIs for filesystem access if compiled with that functionality. At least for zlib-rs: if you provide an allocator and don't need any of the file IO you can compile it without any dependencies (not even standard library or libc, just a couple of core types are needed). zlib-rs does have some testing dependencies though, but I think that is fair. All in: all of them use almost exactly the same external dependencies (i.e.: nothing aside from libc-like functionality). zlib-rs is a bit bigger by default (around 400KB), with some of the Rust machinery. But if you change some of that (i.e. panic=abort), use a nightly compiler (unfortunately still needed for the right flags) and add the right flags both libraries are virtually the same size, with zlib at about 119KB and zlib-rs at about 118KB.
- 1vuio0pswjnm7 2y agoOne of the things I like about C is I can download a statically-compiled native GCC for use on a computer with modest amounts of memory, storage and a relatively old, slow CPU. Total size uncompressed is 242.3MB. Using this I can statically compile a cross-compiler. Total size uncompressed 169.4MB. I use GCC to compille zlib and a wide variety of other software. I can build an operating system from the ground up. Perhaps someday during my lifetime it will be possible to compile programs written in Rust using inexpensive computers with modest amounts of memory, storage and relatively slow CPUs. Meanwhille, there is C.
- WalterGillman 2y ago> Which library has fewer dependencies. This is not insignificant. Remember xz? That could have been a disaster. That the language includes a package manager that fetches an assortment of libraries from who knows whom on demand doesn't exactly inspire confidence in the process to me. Alice's secure AES implementation might bring Eve's string padding function along for the ride. Rust(TM) the language might be (memory) safe in theory but I have serious issues (t)rusting (t)rust and anything built with it.
- throwaway2037 2y agoDoes this performance have anything to do with Rust itself, or is it just more optimized than the other C-language versions (more SIMD instructions / raw assembly code)? I ask because there is a canonical use case where C++ can consistently outperform C -- sorting, because the comparison operator in C++ allows for more compiler optimization compared to the C version: qsort(). I am wondering if there is something similar here for Rust vs C.
- anonymoushn 2y agothese are facts about the C and C++ stdlib sort functions which nobody should really use.
- up2isomorphism 2y agoRust folks love compare rust to C but C folks seldom compare C to rust.
- Narishma 2y agoNot that surprising, Rust folks are more likely to be familiar with C than the reverse.
- quotemstr 2y agoNew native code implementation of zlib faster than old native code version. So what? Rust has a lot of recommend it, but it's not automatically faster than C.
- brianpane 2y agoI contributed a number of performance patches to this release of zlib-rs. This was my first time doing perf work on a Rust project, so here are some things I learned: Even in a project that uses `unsafe` for SIMD and internal buffers, Rust still provided guardrails that made it easier to iterate on optimizations. Abstraction boundaries helped here: a common idiom in the codebase is to cast a raw buffer to a Rust slice for processing, to enable more compile-time checking of lifetimes and array bounds. The compiler pleasantly surprised me by doing optimizations I thought I’d have to do myself, such as optimizing away bounds checks for array accesses that could be proven correct at compile time. It also inlined functions aggressively, which enabled it to do common subexpression elimination across functions. Many times, I had an idea for a micro-optimization, but when I looked at the generated assembly I found the compiler had already done it. Some of the performance improvements came from better cache locality. I had to use C-style structure declarations in one place to force fields that were commonly used together to inhabit the same cache line. For the rare cases where this is needed, it was helpful that Rust enabled it. SIMD code is arch-specific and requires unsafe APIs. Hopefully this will get better in the future. Memory-safety in the language was a piece of the project’s overall solution for shipping correct code. Test coverage and auditing were two other critical pieces.
- Boereck 2y agoInteresting! I wonder if you have used PGO in the project? Forcing fields to be located next to each other kind of feels like something that PGO could do for you.
- brianpane 2y agoI basically did manual PGO because I was also reducing the size of several integer fields at the same time to pack more into each cache line. I’m excited to try out the rustc+LLVM PGO for future optimizations.
- ofek 2y agoA long-standing issue with that was just recently fixed: https://github.com/rust-lang/rust/pull/133250 https://github.com/rust-lang/rust/pull/133250
- miki123211 2y agoI think performance is an underappreciated benefit of safe languages that compile to machine code. If you're writing your program in C, you're afraid of shooting yourself in the foot and introducing security vulnerabilities, so you'll naturally tend to avoid significant refactorings or complicated multithreading unless necessary. If you have Rust's memory safety guarantees, Go's channels and lightweight goroutines, or the access to a test runner from either of those languages, that's suddenly a lot less of a problem. The compiler guarantees you get won't hurt either. Just to give a simple example, if your Rust function receives an immutable reference to a struct, it can rely on the fact that a member of that struct won't magically be mutated by a call to some random function through spooky action at a distance. It can just keep it on the stack / in a callee-saved register instead of fetching it from memory at every loop iteration, if that's more optimal. Then there's the easy access to package ecosystems and extensive standard libraries. If there's a super popular do_foo package, you can almost guarantee that it was a bottleneck for somebody at some point, so it's probably optimized to hell and back. It's certainly more optimized than your simple 10-line do_foo function that you would have written in C, because that's easier than dealing with yet another third-party library and whatever build system it uses.
- Georgelemental 2y ago> The C code is able to use switch implicit fallthroughs to generate very efficient code. Rust does not have an equivalent of this mechanism Rust very much can emulate this, with `break` + nested blocks. But not if you also add in `goto` to previous branches
- randomNumber7 2y agoFinally, now is the day - today - where rust is faster than C
- hackburg 2y ago[dead]