14 ms·
Curious lack of sprintf scaling
- diekhans 5y agoHas this been reported to apple? One would think they would pay attention to something that makes the M1 look slow it it reaches the right people. I18N was one of the few things POSIX didn't do well. Have an environment variable change the sort command causes lots grief.
- pjscott 5y agoIt’s on the front page of HN; I think it’s a safe bet that Radars have been filed — though I’m not sure what can be done about this. The fact that they were using os_unfair_lock indicates that someone has looked at the relevant code and tried to make it efficient; the default is a pthread_mutex_t.
- varenc 5y agoI went and filed this with Apple, radar #9930566, just referencing the blog post. I wouldn't assume the engineering team responsible for this code at Apple is reacting to HN posts so I figure it's still worth reporting. Nows a good time to rant about Apple's entirely opaque bug tracking process. Every bug you file is private. The only way to know if a bug has already been filed is to file a new issue and see if they close it as a duplicate or not.
- addaon 5y agoI wonder if the original export code that lead to this investigation was actually correct? It sounds like sprintf() was being called without an explicit locale. This can be fine if Blender does a top-level setlocale(), but can also be subtly and horrible unfine otherwise... Sounds like there's plenty of opportunity for library-level improvements here (as well as application-level workarounds), but certainly the sprintf_l(..., locale, ...) being slow is the most surprising to me, and likely the easiest to fix.
- markdog12 5y agoYou mean like this, from the article? > Given that this is an Apple operating system, we might know it has a snprintf_l function which takes an explicit locale, and hope that this would make it scale. Just pass NULL which means “use C locale”: And then the chart and discussion following it?
- addaon 5y agoRight. My point is that the change from sprintf() to sprintf(..., NULL, ...) is a semantic-modifying change. For the purpose of understanding what's going on, that's fine. For the purpose of optimizing software, that's scary. And even scarier is that it seems somewhat more likely that the version under test is more correct than the version as shipped.
- markdog12 5y ago> Technically, there are no bugs anywhere above - all the functions work correctly They were all correct. But yeah, scary how primitives can result in such poor performance. FWIW, they ended up using {fmt}: https://developer.blender.org/D13998 https://developer.blender.org/D13998
- rocqua 5y agoThat concerns the printf implementations. The point parent poster os making, is that previously, unless set_locale was being called by blender, the resulting export of a blender object was locale dependent. The change from sprintf(..) to sprintf(..., NULL), would then actually change the behavior (not just performance) of the program.
- EdSchouten 5y agoI think the only reasonable solution is to completely deprecate functions like setlocale(). Software should just use _l() functions if they want something localised. Global state considered harmful.
- sly010 5y agoOr in this case locale state considered harmful ;)
- worik 5y agoGlobal mutable state.....
- stingraycharles 5y agoWell if it wasn’t mutable, it wouldn’t be state but just a global constant variable. Which is much less harmful.
- hinkley 5y agoSnapshots are state that has very, very clear boundary conditions. There are lots of relatively sane systems that only allow modification in the interstitial. Of course as scale goes up, finding “before B but after A” breaks down, but we have whole systems running critical infrastructure based on Communicating Sequential Processes, which is the closest we’ve gotten to solving this problem.
- mort96 5y agoFWIW, the tests show that sprintf_l is no faster than sprintf, if you're passing around the same locale object in different threads. So it can help, but it's not an automatic win unless you make sure that each thread keeps a separate locale object.
- nicoburns 5y agoDo you understand why the locale object needs to be locked? I would have expected it to be immutable...
- MobiusHorizons 5y agoExcuse the lack of context here, I haven't dealt much with locale's in C, and I'm probably showing my ignorance. Why would you ever mutate a locale object? Is that the common way to change locales in C? Wouldn't it make more sense to have locale objects be roughly immutable? It doesn't seem like they should have any real reason to change very often in a typical use-case. I would think any given person only has a small (1-3 or so) number of locale's they use on any regular basis. Are locale objects being mutated really common enough that you need a mutex to protect against accidentally rendering something in the wrong locale?
- mort96 5y agoI would guess that there's just nothing in the standard which _prevents_ people from changing the locale. So if you want a conforming implementation, you need your implementation to work if the programmer changed the locale object directly. EDIT: Nevermind, sprintf_l isn't part of the C standard, so really they could be implemented however the authors chose.
- fyrn- 5y agoThis would really benefit from labeling the y Axis
- formerly_proven 5y agoms for formatting 2 million ints
- AdamH12113 5y agoAgreed. You should always label your axes. In this case, the Y axis is milliseconds on a logarithmic scale.
- formerly_proven 5y agoC/C++ locales are a trashfire. The path to enlightenment is to not use them and discard all libraries which think they can get away with calling setlocale (which a few do, but is more or less a given when we're talking about GUIs). > obviously you should not use sprintf, you should use C++ iostreams Friends don't let friends use iostreams.
- PandaPanda150 5y agoIostreams aside: You're absolutely right with the trashfire. I live in a country where the elders of the language decided that we don't format floating point numbers with a decimal point but use a decimal komma instead. So PI is not 3.14 but 3,14 instead. Just imagine what pain you have to go through if you want to parse a .csv file written by an application that "tried to do everything right and use locale".
- MaxBarraclough 5y ago> So PI is not 3.14 but 3,14 instead. How would you write 3,141.592?
- selfmodruntime 5y ago3.141,592 - Germany formats numbers that way.
- deleted 5y ago[deleted]
- vidarh 5y agoNorway as well
- pilif 5y agoHere in Switzerland it's even worse because it depends on where in Switzerland you're formatting the number (French parts have different rules from German parts) and whether you're formatting a number as a plain number or a monetary value. if it's money, you use the . as the decimal separator everywhere. if it's just a number, in the French parts, it's , in the German parts the . The thousands separator is ' everywhere
- garaetjjte 5y agoIf you want to read rant about setlocale, highly recommended: https://github.com/mpv-player/mpv/commit/1e70e82baa9193f6f027338b0fab0f5078971fbe https://github.com/mpv-player/mpv/commit/1e70e82baa9193f6f02...
- PandaPanda150 5y agoClassic use-case where a posix rwlock will outperform a mutex. How often does the locale change in practice? Almost never.
- hinkley 5y agoWith most locking logic, it’s a series of escalations from the most optimistic/polite to least polite solution, and the worst case behavior is when a sequence always goes to the worst case scenario. In these situations, assuming the worst up front, and jumping straight to it or something very similar saves a lot of bargaining that leads to cache pressure and branch prediction. ETA: It's also quite common in engineering blogs for languages, libraries or frameworks, an entry detailing how in the new version they have made a performance improvement by making the fast case faster, or the predictor more accurate, and then removed option 2 from the decision tree, so that we get a bigger benefit from the happy path and the average case, and as a benefit the system is now simpler as well.
- loeg 5y agorwlock still require contending on a shared, mutated cache line. Something like RCU would bypass that.
- greggman3 5y agowhy is locale stuff even in sprintf? What things get localized? Dates? Line endings? I'm probably dumb but one of things that bugs me with various libraries is when someone has made the decision to do something high level at a low level. For example, localizing inside sprintf, An another example might be an unzip library (read a zip file). Ideally the library should be small IMO. The simplest might you you pass it a bucket of bytes. If you want to make it flexible then you pass it some abstract interface (or 1-2 functions + void* userdata) so you can supply a "read(byteOffset, length)". You can then provide, outside of the library, streamed files, stream networking, etc... But, bad libraries (bad IMO) will instead provide like 12 overrides "unzip(void* bytes), unzip(const char* filename), unzip(socket), unzip(url)" and end up including the world in their library. This kind of "try to do everything" is extremely common in npm libraries :( I don't need your library to include command line parsing! If you want to make a tool, make a library, then make a separate tool that uses that library. Keep the 2 separated so users of the library don't need dependencies that only the tool needs. (probably the most common npm example but there are lots of others) Really surprised something as low-level as sprintf needs locale. Even streams I'd expect maybe a Date object would but not the stream itself.
- kazinator 5y agolocale stuff is even in strtod, so you can read numbers like 3,14159 in locales where comma is the decimal point.
- cyral 5y agoI get your point about bloated NPM libraries but I think it's also ironic that NPM is the only registry with some of the _smallest_ packages which are equally as annoying. left-pad, ansi-yellow, ansi-red (yes, packages for a single console color with over 200k monthly downloads), is-odd, is-even (which of course, depends on is-odd and inverts it), etc.
- haggy102 5y agoDon't forget about "is-thirteen" and its inverse lol https://www.npmjs.com/package/is-thirteen https://www.npmjs.com/package/is-thirteen https://www.npmjs.com/package/is-not-thirteen https://www.npmjs.com/package/is-not-thirteen
- markdog12 5y agoAnyone reading this will probably also be interested in his related post: https://aras-p.info/blog/2022/02/03/Speeding-up-Blender-.obj-export/ https://aras-p.info/blog/2022/02/03/Speeding-up-Blender-.obj...
- matheist 5y agoI ran into sprintf's dependence on locale recently when trying to use it in a WebAssembly module. Since I was compiling for the browser, I wanted to not depend on locale. Even setting aside the locale stuff, the wasi sdk still wanted to pull in file-related things like read/write/seek. I just wanted to do formatted print to a pre-existing buffer. I ended up using nanoprintf — it's a single header file and in the public domain. (https://github.com/charlesnicholson/nanoprintf https://github.com/charlesnicholson/nanoprintf)
- matheist 5y agoP.S.: musl's sprintf is the one that I noticed pulling in file APIs, via vfprintf. Maybe other libc's implement sprintf without using the file API.
- bla3 5y ago> And no, usual Internet advice of “MSVC sucks, use Clang” Given that this talks about a problem in the Microsoft standard library, wouldn't the usual internet advice be "use clang, and also llvm's libc++"? If clang just compiles the same slow code as msvc, it won't magically make it fast.
- kevingadd 5y agoIf you're building for windows, interop may require that you also use the msvc standard library.
- jcelerier 5y agoIt's honestly not really true on windows. Since for the longest time the msvc ABI was unstable people are used to interop through C APIs, or rebuilding everything. I personally build all my stuff with libc++ and it works fine.
- leni536 5y agoIf you also want to use libraries that are built against MS STL, then you can't really use libc++. Of course if you can build everything with libc++, that's fine. That's how Chromium is built, AFAIK.
- jcelerier 5y agoThat depends, on windows multiple standard libraries (even libc) can cohabit in the same process (but you have to be careful to, say, not free something in a different dll than the one which allocated it). If you are careful with that and aren't exchanging standard library types across library boundaries there won't be issued.
- fwsgonzo 5y agofmt is a must if you can get away with using it. On smaller systems you should try strf, as it produces very small memory footprint and doesn't blow up your binaries. It is 5x smaller than fmt in footprint.
- thaumasiotes 5y ago> So given all this knowledge, presumably, if each thread used a physically different locale object and snprintf_l, then it would scale fine. And it does: The author provides several graphs in which what appears to be the total execution time stays constant as the number of threads varies, and this is described as "good scaling". How do I know what's being graphed is the total execution time? > Converting two million numbers into strings takes 100 milliseconds when one CPU core is doing it. When all eight “performance” cores are doing it, it takes 1.8 seconds, or 18 times as long. This corresponds to a curve where "one core" takes the value 100 and "8 cores" takes the value 1866. But isn't constant execution time as we increase from one thread to eight threads terrible scaling? What's happening here?
- electroly 5y agoIf it takes 10 seconds to compute something once on one core, and it still takes 10 seconds to compute it eight times on eight cores, you have achieved perfect scaling. They're performing 8x as much work for the 8-core test. That is: they are converting 16 million numbers, two million per core. I think maybe you're thinking they're running the same amount of work for the 1- and 8-core tests, but that's not what this test is doing.
- rrdharan 5y agoIt is a slightly confusing way of presenting it though.. would’ve been a bit clearer IMO if the graphs were showing “time taken to complete 2 million conversions” with the graphs sloping downward as cores increase in the good cases (scaling) and going upward (pathological / lack of scaling) for the bad cases..
- rocqua 5y agoThis is much easier to read. Determining if a line is perfectly flat is easy. Determining if a line is perfectly 'c/x' (plotting time for same total work) is really hard. Plotting 1/runtime makes interpreting the actual meaning of a single point much harder. So that is also out of the question. I took this approach to mean "threads don't affect eachother's performance". Which is easily seen to be equivalent to perfect scaling.
- peapicker 5y agoCurious that strtol() etc wasn’t profiled as this concerns turning ints into strings. Snprintf has a lot of other overhead.
- eyelidlessness 5y agoParticulars aside, the pathology is exactly what I’d expect from low level manually managed memory languages. Oh you want a multithreaded abstraction without a VM or GC and you want it to just work? It’s going to do a lot of extra work or it’s gonna be the next “logging failboat” or both.
- rurban 5y agoI've recently implemented the secure variant without any locale support. Took me a day. Now just the secure scanf family is missing, and this needs locale support unfortunately. https://github.com/rurban/safeclib/blob/master/src/str/vsnprintf_s.c https://github.com/rurban/safeclib/blob/master/src/str/vsnpr...
- trulyme 5y ago> I only have an Ubuntu 20 install via WSL2 here to test, and using the default compilers there (clang 10 and gcc 9.3), things look pretty nice: WSL2 is not that bad, but installing a real Linux in a VM takes about 10 minutes, so this seems a weird thing to say.
- leni536 5y agoWSL2 is a real Linux in a VM.
- trulyme 5y agoWSL2, afaik, runs modified Linux kernel (maintained by MS) on a modified Hyper-V. I would hardly call that a VM in traditional sense because it might introduce differences in behavior. However it is true that WSL2 uses virtualization technologies, so in that sense it is a VM. But calling it "a real Linux in a VM" is a stretch imho.
- pjmlp 5y agoVery nice article, although the jabs at zero cost abstractions are out of place. Bjarne has always used the expression to mean it doesn't cost more than if the same code has been manually written by hand. An Assembly written version of iostreams, with the same architecture, would perform the same.
- kjgkjhfkjf 5y agoTL;DR sprintf has to acquire a lock to access some shared data so the multiple CPUs can't operate concurrently.
- bborud 5y agoLocale in library functions in general is an example of solving problems at the wrong abstraction level. It also breaks the fundamental idea of how shell tools are meant to be usable in unix. If they don't have stable outputs, then they become unusable. (We had to create a set of locale-resistant and consistent versions of shell tools for a system that made heavy use of shell tools to process large amounts of data across thousands of machines. All you needed was one misconfigured locale on one machine and the result would be chaos). You pay the cost every time rather than when you actually care about it. And when you really don't want locale to interfere it still comes back to haunt you if you don't pay special attention to it. (Remember how Python had locale-dependent XML-RPC that made sure two machines with different ways of formatting floats behaved?). Locale is bad design. Very bad design.
- nec4b 5y agoI don't think author understands what zero cost abstraction means. He does have a point about c++ standard library still being a hit or a miss regarding performance.
- sirwhinesalot 5y agoToday in why global variables are bad...
- Unit520 5y agoReading this article motivated me to take a closer look at printf/iostreams alternatives, and I have to say, the mentioned {fmt} library finally made me switch, so thanks for that!