22 ms·
Rust std fs slower than Python? No, it's hardware
- Pop_- 3y agoDisclaimer: The title has been changed to "Rust std fs slower than Python!? No, it's hardware!" to avoid clickbait. However I'm not able to fix the title in HN.
- 3cats-in-a-coat 3y agoWhat's the TLDR on how... hardware performs differently on two software runtimes?
- lynndotpy 3y agoOne of the very first things in the article is a TLDR section that points you to the conclusion. > In conclusion, the issue isn't software-related. Python outperforms C/Rust due to an AMD CPU bug.
- j16sdiz 3y agoIt is software-related. Just the CPU perform badly on some software instruction.
- xuanwo 3y agoFSRM is a CPU feature embedded in the microcode (in this instance, amd-ucode) that software such as glibc cannot interact with. I refer to it as hardware because I consider microcode a part of the hardware.
- pornel 3y agoAMD's implementation of `rep movsb` instruction is surprisingly slow when addresses are page aligned. Python's allocator happens to add a 16-byte offset that avoids the hardware quirk/bug.
- sound1 3y agothank you, upvoted!
- sharperguy 3y ago"Works on contingency? No, money down!"
- deleted 3y ago[deleted]
- pvg 3y agoyou can mail hn@ycombinator.com and they can change it for you to whatever.
- royjacobs 3y agoI was prepared to read the article and scoff at the author's misuse of std::fs. However, the article is a delightful succession of rabbit holes and mysteries. Well written and very interesting!
- zerr 3y ago[flagged]
- marcosdumay 3y agoIt's about making misuse difficult. Rust doesn't actually restrict much. It would take looking at lots of details, but my impression is that it's less restrictive than C.
- qerti 3y agoThe rust compiler refuses to compile code which doesn’t adhere to a strict set of rules guaranteeing memory safety. Unless you intentionally call an unsafe block, misuse in this sense is impossible, not just difficult.
- db48x 3y agoBut other types of misuse are not. For example, a naive program might do many reads into a small buffer instead of one read into a large buffer.
- Yoric 3y agoJust stating the obvious: breaking memory safety is only one subset of all the possible misuse available in the space of manners of using an API, data structure, etc.
- bri3d 3y agoThis was such a good article! The debugging was smart (writing test programs to peel each layer off), the conclusion was fascinating and unexpected, and the writing was clear and easy to follow.
- sgift 3y agoEither the author changed the headline to something less clickbaity in the meantime or you edited it for clickbait Pop_- (in that case: shame on you) - current headline: "Rust std fs slower than Python!? No, it's hardware!"
- epage 3y agoBased on the /r/rust thread, the author seemed to change the headline based on feedback to make it less clickbait-y
- xuanwo 3y agoSorry for the clickbaity title, I have changed it based on others advice.
- thechao 3y agoI disagree that it's clickbait-y. Diving down from Python bindings to ucode is ... not how things usually go. Doubly so, since Python is a very mature runtime, and I'd be inclined to believe they've dug up file-reading Kung Fu not available to the Average Joe.
- jll29 3y agoThanks for this unexpected, thriller-like read. I'm impressed by your perseverance, how you follow through with your investigation to the lowest (hardware) level.
- Pop_- 3y agoThe author has updated the title and also contacted me. But unfortunately I'm no longer able to update it so.
- deleted 3y ago[deleted]
- Pesthuf 3y agoClickbait headline, but the article is great!
- joshfee 3y agoSurprisingly I think this usage of clickbait is totally reasonable because it matches the author's initial thoughts/experiences of "what?! this can't be right..."
- saghm 3y agoI think there might be a range of where people draw the line between reasonable headlines and clickbait, because I tend to think of clickbait as something where the "answer" to some question is intentionally left out to try to bait people into clicking. For this article, something I'd consider clickbait would be something like "Rust std fs is slower than Python?" without the answer after. More commonly, the headline isn't phrased directly as a question, but instead of saying something like "So-and-so musician loves burritos", it will leave out the main detail and say something like "The meal so-and-so eats before every concert", which is trying to get you to click and have to read through lots of extraneous prose just to find the word "burritos". Having a hook to get people to want to read the article is reasonable in my opinion; after all, if you could fit every detail in the size of a headline, you wouldn't need an article at all! Clickbait inverts this by _only_ having enough enough substance that you could get all the info in the headline, but instead it leaves out the one detail that's interesting and then pads it with fluff that you're forced to click and read through if you want the answer.
- Attummm 3y ago[flagged]
- insanitybit 3y agoIn this case it is a hardware bug and in no way attributable to Python being fast.
- meneer_oke 3y agoThe bug is the other way around :)
- agumonkey 3y agoreminds me of HP48 programming, there was sysrpl which was near asm speed, and then user rpl which was slooow
- MrJohz 3y agoIt's worth reading the article. In this case, it seems to have been a hardware issue - as such, not directly related to Rust, C, or Python, but triggered by an instruction that was only called by some file loading routines. It's a very cool deep dive into debugging these sorts of issues.
- Attummm 3y agoAlthough true that it's great article. It states that python is faster then c, that is not possible since python is build with c. There could be other reasons such libs or implementation. Also note that the issue he had was not resolved. The comment was about that python is seen as slow. But that is not always the case. Once a dev is able to understand the difference between the python and c parts. Python can be quite performant, and efficient with memory. But if one would actually create a application that does more then just read a file it will be slow again compared to c and rust.
- Pop_- 3y agoIt's not stating python is faster than c in general. This is just one very specific case where non-page-aligned memeory reading on AMD is involved.
- iampims 3y agoMost interesting article I've read this week. Excellent write-up.
- quietbritishjim 3y agoI'm a bit confused about the premise. This is not comparing pure Python code against some native (C or Rust) code. It's comparing one Python wrapper around native code (Python's file read method) against another Python wrapper around some native code (OpenDAL). OK it's still interesting that there's a difference in performance, but it's very odd to describe it as "slower than Python". Did they expect that the Python standard library is all written in pure Python? On the contrary, I would expect the implementations of functions in Python's standard library to be native and, individually, highly optimised. I'm not surprised the conclusion had something to do with the way that native code works. Admittedly I was surprised at the specific answer - still a very interesting article despite the confusing start. Edit: The conclusion also took me a couple of attempts to parse. There's a heading "C is slower than Python with specified offset". To me, as a native English speaker, this reads as "C is slower (than Python) with specified offset" i.e. it sounds like they took the C code, specified the same offset as Python, and then it's still slower than Python. But it's the opposite: once the offset from Python was also specified in the C code, the C code was then faster. Still very interesting once I got what they were saying though.
- xuanwo 3y agoThanks for the comments. I have fixed the headers :)
- crabbone 3y ago> individually, highly optimised. Now why would you expect that? What happened to OP is a pure chance. CPython's C code doesn't even care about const-consistency. It's flush with dynamic memory allocations, bunch of helper / convenience calls... Even stuff like arithmetic does dynamic memory allocation... Normally, you don't expect CPython to perform well, not if you have any experience working with it. Whenever you want to improve performance you want to sidestep all the functionality available there. Also, while Python doesn't have a standard library, since it doesn't have a standard... the library that's distributed with it is mostly written in Python. Of course, some of it comes written in C, but there's also a sizable fraction of that C code that's essentially Python code translated mechanically into C (a good example of this is Python's binary search implementation which was originally written in Python, and later translated into C using Python's C API). What one would expect is that functionality that is simple to map to operating system functionality has a relatively thin wrapper. I.e. reading files wouldn't require much in terms of binding code because, essentially, it goes straight into the system interface.
- drtgh 3y ago>Rust std fs slower than Python!? No, it's hardware! >... >Python features three memory domains, each representing different allocation strategies and optimized for various purposes. >... >Rust is slower than Python only on my machine. if one library performs wildly better than the other in the same test, on the same hardware, how can that not be a software-related problem? sounds like a contradiction. Maybe should be considered a coding issue and/or feature absent? IMHO it would be expected Rust's std library perform well without making all the users to circumvent the issue manually. The article is well investigated so I assume the author just want to show the problem existence without creating controversy because other way I can not understand.
- Pop_- 3y agoThe root cause is AMD's bad support for rep movsb (which is a hardware problem). However, python by default has a small offset when reading memories while lower level language (rust and c) does not, which is why python seems to perform better than c/rust. It "accidentally" avoided the hardware problem.
- CoastalCoder 3y agoI'm not sure it makes sense to pin this only on AMD. Whenever you're writing performance-critical software, you need to consider the relevant combinations of hardware + software + workload + configuration. Sometimes a problem can be created or fixed by adjusting any one / some subset of those details.
- hobofan 3y agoIf that's a bug that only happens with AMD CPUs, I think that's totally fair. If we start adding in exceptions at the top of the software stack for individuals failures of specific CPUs/vendors, that seems like a strong regression from where we are today in terms of ergonomics of writing performance-critical software. We can't be writing individual code for each N x M x O x P combination of hardware + software + workload + configuration (even if you can narrow down the "relevant" ones).
- exxos 3y agoIt's the hardware. Of course Rust remains the fastest and safest language and you must rewrite your applications in Rust.
- dang 3y agoYou've been posting like this so frequently as to cross into abusing the forum, so I've banned the account. If you don't want to be banned, you're welcome to email hn@ycombinator.com and give us reason to believe that you'll follow the rules in the future. They're here: https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html.
- Aissen 3y agoAssociated glibc bug (Zen 4 though): https://sourceware.org/bugzilla/show_bug.cgi?id=30994 https://sourceware.org/bugzilla/show_bug.cgi?id=30994
- Arnavion 3y agoThe bug is also about Zen 3, and even mentions the 5900X (the article author's CPU).
- nabakin 3y agoIf you read the bug tracker, a comment mentions this affects Zen 3 and Zen 4
- fweimer 3y agoAnd AMD is investigating: https://inbox.sourceware.org/libc-alpha/20231115190559.2911267-1-sajan.karumanchi@amd.com/ https://inbox.sourceware.org/libc-alpha/20231115190559.29112...
- explodingwaffle 3y agoAnyone else feeling the frequency illusion with rep movsb? (https://lock.cmpxchg8b.com/reptar.html https://lock.cmpxchg8b.com/reptar.html)
- deleted 3y ago[deleted]
- saagarjha 3y agoThis is unrelated.
- a1o 3y ago> Rust developers might consider switching to jemallocator for improved performance I am curious if this is something that everyone can do to get free performance or if there are caveats. Can C codebases benefit from this too? Is this performance that is simply left on table currently?
- nicoburns 3y agoI think it's pretty much free performance that's being left on the table. There's slight cost to binary size. And it may not perform better in absolutely all circumstances (but it will in almost all). Rust used to use jemalloc by default but switched as people found this surprising as the default.
- Pop_- 3y agoSwitching to non-default allocator does not always brings performance boost. It really depend on your workload, which requires profiling and benchmarking. But C/C++/Rust and other lower level languages should all at least be able to choose from these allocators. One caveat is binary size. Custom allocator does add more bytes to executable.
- vlovich123 3y agoI don’t know why people still look to jemalloc. Mimalloc outperforms the standard allocator on nearly every single benchmark. Glibc’s allocator & jemalloc both are long in the tooth & don’t actually perform as well as state of the art allocators. I wish Rust would switch to mimalloc or the latest tcmalloc (not the one in gperftools).
- masklinn 3y ago> I wish Rust would switch to mimalloc or the latest tcmalloc (not the one in gperftools). That's nonsensical. Rust uses the system allocators for reliability, compatibility, binary bloat, maintenance burden, ..., not because they're good (they were not when Rust switched away from jemalloc, and they aren't now). If you want to use mimalloc in your rust programs, you can just set it as global allocator same as jemalloc, that takes all of three lines: https://github.com/purpleprotocol/mimalloc_rust#usage https://github.com/purpleprotocol/mimalloc_rust#usage If you want the rust compiler to link against mimilloc rather than jemalloc, feel free to test it out and open an issue, but maybe take a gander at the previous attempt: https://github.com/rust-lang/rust/pull/103944 https://github.com/rust-lang/rust/pull/103944 which died for the exact same reason the the one before that (https://github.com/rust-lang/rust/pull/92249 https://github.com/rust-lang/rust/pull/92249) did: unacceptable regression of max-rss.
- fsniper 3y agoThe article itself is a great read and it has fascinating info related to this issue. However I am more interested/concerned about another part. How the issue is reported/recorded and how the communications are handled. Reporting is done over discord, which is a proprietary environment which is not indexed, or searchable. Will not be archived. Communications and deliberations are done over discord and telegram, which is probably worse than discord in this context. This blog post and the github repository is the lingering remains of them. If Xuanwo did not blog this. It would be lost in timeline. Isn't this fascinating?
- jll29 3y ago> Reporting is done over discord, which is a proprietary environment which is not indexed, or searchable. Will not be archived. That's why I don't accept the response "but there's Discord now" whenever I moan about USENET's demise. Back in the days before it, every post was nicely searchable by DejaNews (later Google). We need to get back to open standards for important communications (e.g. all open source projects that are important to the Internet/WWW stack and core programming and libraries).
- upsuper 3y agoYes, they are proprietary, which is not great. But I don't buy the allegation that they are not indexed or searchable. There are very few IMs that provide builtin publicly accessable log indexed or searchable by default. Does every IRC server come with public log? What about Matrix groups? How do discussion there not get lost in timeline? You can provide public log of them not because they are not proprietary, but that they have API to allow logging. Telegram also has such API, and FWIW our discussion group does have searchable log that you can access here: https://luoxu-web.vercel.app/#g=1264662201 https://luoxu-web.vercel.app/#g=1264662201 It is not indexable publicly more for privacy concern, again not because the platform is proprietary.
- fsniper 3y agoThis is not a way to have bug discussions, or record them. Do you really think I could find this information on a search for a similar issue? Only thing that makes this bug and the process of the debug visible is this blog post. Another point is I don't think IRC or any instant messaging app is the correct place for this kinds of discussions. Unless important points are logged to some bug reporting tool, or perhaps a mailing list, or to a blog post like this one, they are useless for historic purposes.
- amluto 3y agoI sent this to the right people.
- londons_explore 3y agoSo the obvious thing to do... Send a patch to change the "copy_user_generic" kernel method to use a different memory copying implementation when the CPU is detected to be a bad one and the memory alignment is one that triggers the slowness bug...
- p3n1s 3y agoNot obvious. Seems like if it can be corrected with microcode just have people use updated microcode rather than litter the kernel with fixes that are effectively patchable software problems. The accepted fix would not be trivial to anyone not already experienced with the kernel. But more important, it obviously isn’t obvious what is the right way to enable the workaround. The best way is to probably measure at boot time, otherwise how do you know which models and steppings are affected.
- londons_explore 3y agoI don't think AMD does microcode updates for performance issues do they? I thought it was strictly correctness or security issues. If the vendor won't patch it, then a workaround is the next best thing. There shouldn't be many - that's why all copying code is in just a handful of functions.
- p3n1s 3y agoA significant performance degradation due to normal use of the instruction (FSRM) not otherwise documented is a correctness problem. Especially considering that the workaround is to avoid using the CPU feature in many cases. People pay for this CPU feature now they need kernel tooling to warn them when they fallback to some slower workaround because of an alignment issue way up the stack.
- prirun 3y agoIf AMD has a performance issue and doesn't fix it, AMD should pay the negative publicity costs rather than kernel and library authors adding exceptions. IMHO.
- 3y ago
- pmontra 3y ago> However, mmap has other uses too. It's commonly used to allocate large regions of memory for applications. Slack is allocating 1132 GB of virtual memory on my laptop right now. I don't know if they are using mmap but that's 1100 GB more than the physical memory.
- Waterluvian 3y agoI’m not sure allocations mean anything practical anymore. I recall OSX allocating ridiculous amounts of virtual memory to stuff but never found OSX or the software to ever feel slow and pagey.
- dietrichepp 3y agoThe way I describe mmap these days is to say it allocates address space. This can sometimes be a clearer way of describing it, since the physical memory will only get allocated once you use the memory (maybe never).
- byteknight 3y agoBut is it not still limited by allocating the RAM + Page/Swap size?
- wbkang 3y agoI don't think so, but it's difficult to find an actual reference. For sure it does overcommit like crazy. Here's an output from my mac: % ps aux | sort -k5 -rh | head -1 xxxxxxxx 88273 1.2 0.9 1597482768 316064 ?? S 4:07PM 35:09.71 /Applications/Slack.app/Contents/Frameworks/Slack Helper (Renderer).app/... Since ps displays vsz column in KiB, 1597482768 corresponds to 1TB+.
- aseipp 3y agoMaybe I'm misunderstanding you but: no, you can allocate terabytes of address space on modern 64-bit Linux on a machine with only 8GB of RAM with overcommit. Try it; you can allocate 2^46 bytes of space (~= 100TB) today, with no problem. There is no limit to the allocation space in an overcommit system; there is only a limit to the actual working set, which is very different.
- mgaunard 3y ago[flagged]
- deleted 3y ago[deleted]
- comonoid 3y agojemalloc was Rust's default allocator till 2018. https://internals.rust-lang.org/t/jemalloc-was-just-removed-from-the-standard-library/8759 https://internals.rust-lang.org/t/jemalloc-was-just-removed-...
- exxos 3y ago[dead]
- titaniumtown 3y agoExtremely well written article! Very surprising outcome.
- diamondlovesyou 3y agoAMD's string store is not like Intel's. Generally, you don't want to use it until you are past the CPU's L2 size (L3 is a victim cache), making ~2k WAY too small. Once past that point, it's profitable to use string store, and should run at "DRAM speed". But it has a high startup cost, hence 256bit vector loads/stores should be used until that threshold is met.
- rasz 3y agoOr you leave it as is forcing AMD to fix their shit. "fast string mode" has been strongly hinted as _the_ optimal way over 30 years ago with Pentium Pro, further enforced over 10 years ago with ERMSB and FSRM 4 years ago. AMD get with the program.
- saagarjha 3y agorep movsb might have been fast at one point but it definitely was not for a few decades in the middle, where vector stores were the fastest way to implement memcpy. Intel decided that they should probably make it fast again and they have slowly made it competitive with the extensions you’ve mentioned. But for processors that don’t support it, using rep movsb is going to be slow and probably not something you’d want to pick unless you have weird constraints (binary size?)
- js2 3y agoIsn't the high startup cost what FSRM is intended to solve? > With the new Zen3 CPUs, Fast Short REP MOV (FSRM) is finally added to AMD’s CPU functions analog to Intel’s X86_FEATURE_FSRM. Intel had already introduced this in 2017 with the Ice Lake Client microarchitecture. But now AMD is obviously using this feature to increase the performance of REP MOVSB for short and very short operations. This improvement applies to Intel for string lengths between 1 and 128 bytes and one can assume that AMD’s implementation will look the same for compatibility reasons. https://www.igorslab.de/en/cracks-on-the-core-3-yet-the-5-ghz-sample-of-the-16-kernels-with-4-9-ghz-sample-emerged-on-the-implemented-further-x86-instructions-from-intel/%0A https://www.igorslab.de/en/cracks-on-the-core-3-yet-the-5-gh...
- forrestthewoods 3y agoDelightful article. Thank you author for sharing! I felt like I experienced every shock twist in surprise in your journey like I was right there with you all along.
- darkwater 3y agoTotally unrelated but: this post talks about the bug being first discovered in OpenDAL [1], which seems to be an Apache (Incubator) project to add an abstraction layer for storage over several types of storage backend. What's the point/use case of such an abstraction? Anybody using it? [1] https://opendal.apache.org/ https://opendal.apache.org/
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- the8472 3y agoThere are two dedicated CPU feature flags to indicate that REP STOS/MOV are fast and usable as short instruction sequence for memset/memcpy. Having to hand-roll optimized routines for each new CPU generation has been an ongoing pain for decades. And yet here we are again. Shouldn't this be part of some timing testsuite of CPU vendors by now?
- giancarlostoro 3y agoSo correct me if I am wrong but does this mean you need to compile two executables for a specific compile time build? Or is it just you need to compile it from specific hardware? Wondering what the fix would be, some sort of runtime check?
- immibis 3y agoglibc has the ability to dynamically link a different version of a function based on the CPU.
- dralley 3y agoGlibc supports runtime selection of different optimized paths, yes. There was a recent discussion about a security vulnerability in that feature (discussion https://news.ycombinator.com/item?id=37756357 https://news.ycombinator.com/item?id=37756357), but in essence this is exactly the kind of thing it's useful for.
- fweimer 3y agoThe exact nature of the fix is unclear at present. During dynamic linking, glibc picks a memcpy implementation which seems most appropriate for the current machine. We have about 13 different implementations just for x86-64. We could add another one for current(ish) AMD CPUs, select a different existing implementation for them, or change the default for a configurable cutover point in a parameterized implementation.
- saagarjha 3y agoThis code is in the kernel, so dynamic linking and glibc is not really relevant.
- lxe 3y agoI wonder what other things we can improve by removing spectre mitigations and tuning hugepage, syscall altency, and core affinity
- saagarjha 3y agoMitigations did not have a meaningful performance impact here.
- lxe 3y agoSo Python isn't affected by the bug because pymalloc performs better on buggy CPUs than jemalloc or malloc?
- js2 3y agoIt has nothing to do with pymalloc's performance per se. Rather, the performance issue only occurs when using `rep movsb` on AMD CPUs with certain page/data alignment. Pymalloc just happens to be using page/data alignment that makes `rep movsb` happy while Rust's default allocator is using alignments that just happen to make `rep movsb` sad.
- jokethrowaway 3y agoClickbait title but interesting article. This has nothing to do with python or rust
- codedokode 3y agoWhy is there need to move memory? Hardware cannot DMA data into non-page-aligned memory? Or Linux doesn't want to load non-aligned data?
- wmf 3y agoThe Linux page cache keeps data page-aligned so if you want the data to be unaligned Linux will copy it.
- codedokode 3y agoWhat if I don't want to use cache?
- tedunangst 3y agoPull out some RAM sticks.
- wmf 3y agoYou can use O_DIRECT although that also forces alignment IIRC.
- eigenform 3y agowould be lovely if ${cpu_vendor} would document exactly how FSRM/ERMS/etc are implemented and what the expected behavior is
- saagarjha 3y agoIt is documented; this is a performance bug.
- fulafel 3y agoA related thing from times when it was common that memory layout artifacts had high impact on sw performance: https://en.wikipedia.org/wiki/Cache_coloring https://en.wikipedia.org/wiki/Cache_coloring
- collinmanderson 3y agoBTW, I've always thought Python uses way too many syscalls when working with files. Simple code like this uses something like 9 syscalls (shown in the article): with open('myfile') as f: data = f.read() I'm not much of a C programmer myself. but I at least reported part of the issue to Python: https://bugs.python.org/issue45944 https://bugs.python.org/issue45944 This is the fastest way to read a file on python that I've found, using only 3-4 syscalls (though os.fstat() doesn't work for some special files kernel files like those in /proc/ and /dev/): def read_file(path: str, size=-1) -> bytes: fd = os.open(path, os.O_RDONLY) try: if size == -1: size = os.fstat(fd).st_size return os.read(fd, size) finally: os.close(fd)
- the8472 3y agoAs you say, the reported size is not necessarily correct so it should only be treated as a hint. And if os.read directly translates to a read syscall then you're also not handling short reads.
- collinmanderson 3y agoAhh ok so to be correct you have to keep reading until you get an empty read? Maybe I don’t need to query the file size at all?
- the8472 3y agoquerying the file size can be useful to choose the allocation size for a buffer. but yes, you have to keep reading until you get a zero-length read. https://man7.org/linux/man-pages/man2/read.2.html https://man7.org/linux/man-pages/man2/read.2.html > On success, the number of bytes read is returned (zero indicates end of file), [...] It is not an error if this number is smaller than the number of bytes requested