13 ms·
Facebook's std::vector optimization
- deleted 12y ago[deleted]
- ajasmin 12y agoTLDR; The author of malloc and std::vector never talked to each other. We fixed that! ... also most types are memmovables
- darkpore 12y agoYou can get around a lot of these issues by reserving the size needed up front, or using a custom allocator with std::vector. Not as easy, but still doable. The reallocation issue isn't fixable this way however...
- thrownaway2424 12y agoYou can use a feedback directed optimization pass to choose the initial size.
- xroche 12y agoYep, this is my biggest issue with C++: you now have lambdas functions and an insane template spec, but you just can not "realloc" a new[] array. Guys, seriously ?
- bnegreve 12y agoIf you need to realloc a fixed size array, souldn't you use a std::vector instead?
- marksamman 12y agoYou probably should, but the problem is still there because std::vector implementations don't use realloc. They call new[] with the new size, copy over the data and delete[] the old chunk. This eliminates the possibility to grow the vector in-place.
- bnegreve 12y agoIt's the same with realloc: there is no guarantee that it will grow the chunk in place.
- xroche 12y agoNo. Modern realloc are efficient, when moving large memory blocks, because they rely on the kernel ability to quickly relocate memory regions without involving memcpy() (through mremap() on Linux). Edit: shamelessly citing my blog entry on this subject: http://blog.httrack.com/blog/2014/04/05/a-story-of-realloc-and-laziness/ http://blog.httrack.com/blog/2014/04/05/a-story-of-realloc-a...
- aktau 12y agoFantastic blog post! Now I don't have to start digging myself :). I've always thought realloc could do a few optimizations, glad to find out the details.
- davidtgoldblatt 12y agoThis isn't true for either of the common high performance mallocs. tcmalloc doesn't support mremap at all, and jemalloc disables it in its recommended build configuration. glibc will remap, but any performance gains that get realized from realloc tricks would probably be small compared to the benefits of switching to a better malloc implementation.
- marksamman 12y agoguarantee != possibility. There's no guarantee with realloc, but there's no possibility with new[], copy and delete[]. You can't grow the size you allocated with new[] in-place, and because you need to retain the existing data it's not safe to delete[] the old buffer, call new[] and hope that it points to the previous memory address and assume that the existing data remains intact. A realloc implementation can try to grow the buffer if there's contiguous space, and if it succeeds it doesn't need to deallocate or copy anything. I haven't had a look at realloc implementations so I don't know if that, or other optimizations are done in practice, but I assume that realloc's worst performance case is somewhere around new[], copy and delete[]'s best case. The copy mechanism in std::vector may also have a significant overhead over realloc if it has to call the copy constructor (or ideally move constructors in C++11) of every object in the vector, although I can imagine a C++ equivalent of realloc doing so too.
- pjmlp 12y agoWhy should C++persist in C insecure design decisons? Besides realloc can force a memory move anyway. Its behavior depends from implementation and memory state.
- xroche 12y agoHow realloc is insecure ? The only lacking construct is a "moved" operator, which would re-assign pointer members, for example, when moving this to another location.
- pjmlp 12y agoTypical security exploit with realloc is attacking code that assumes realloc always success and never moves, thus keeping all pointers around.
- huhtenberg 12y agoAssuming that realloc() always succeeds is one thing, I can believe there's code that makes it. But I'd very much like to look at the code that assumes that realloc() never moves the block. This sounds a remarkably elaborate assumption to make, hardly an oversight due to ignorance.
- pjmlp 12y agoLots of enterprise code with off-shoring provide entertainment of code quality.
- ck2 12y agoThen the teleporting chief would have to shoot the original As an aside, there was a great Star Trek novel where there was a long range transporter invented that accidentally cloned people. (I think it was "Spock Must Die")
- dalke 12y agoThere's also the ST:TNG episode "Second Chances". http://en.wikipedia.org/wiki/Second_Chances_%28Star_Trek:_The_Next_Generation%29 http://en.wikipedia.org/wiki/Second_Chances_%28Star_Trek:_Th... where Riker is duplicated. (There's also the good/bad Kirk in 'The Enemy Within', and the whole mirror universe concept, but those aren't duplicates.)
- userbinator 12y agoWhen the request for growth comes about, the vector (assuming no in-place resizing, see the appropriate section in this document) will allocate a chunk next to its current chunk This is assuming a "next-fit" allocator, which is not always the case. I think this is why the expansion factor of 2 was chosen - because it's an integer, and doesn't assume any behaviour of the underlying allocator. I'm mostly a C/Asm programmer, and dynamic allocation is one of the things that I very much avoid if I don't have to - I prefer constant-space algorithms. If it means a scan of the data first to find out the right size before allocating, then I'll do that - modern CPUs are very fast "going in a straight line", and realloc costs add up quickly. Another thing that I've done, which I'm not entirely sure would be possible in "pure C++", is to adjust the pointers pointing to the object if reallocation moves it (basically, add the difference between the old and new pointers to each reference to the object); in theory I believe this involves UB - so it might not be "100% standard C" either, but in practice, this works quite well.
- mlvljr 12y agoThe latter nicely addresses to the recent HN discussion about "Friendly C"; may I ask what kind of software you are building, incidentally? (...don't tell it's medical :O)
- coliveira 12y agostd::vector is supposed to work in the common cases, so you don't need to manually realloc for small/medium sized vectors. If you need to handle a lot of data, avoiding realloc is still an issue and you should pre-compute the length and use vector::resize (or the length-aware constructor).
- kabdib 12y agoI solved an "realloc is really costly" problem by ditching the memory-is-contiguous notion, paying a little more (really just a few cycles) for each access rather than spending tons of time shuffling stuff around in memory. This eliminated nearly all reallocations. The extra bit of computation was invisible in the face of cache misses. I'm guessing that most customers of std::vector don't really need contiguous memory, they just need something that has fast linear access time. In this sense, std::vector is a poor design.
- CJefferson 12y agoI'm glad to see this catch on and the C level primitives get greater use. This has been a well known problem in the C++ community for years, in particular Howard Hinnant put a lot of work into this problem. I believe the fundamental problem has always been that C++ implementations always use the underlying C implementations for malloc and friends, and the C standards committee could not be pursaded to add the necessary primitives. A few years ago I tried to get a reallic which did not move (instead returned fail) into glibc and jealloc and failed. Glad to see someone else has succeeded.
- cliff_r 12y agoThe bit about special 'fast' handling of relocatable types should be obviated by r-value references and move constructors in C++11/14, right? I.e. if we want fast push_back() behavior, we can use a compiler that knows to construct the element directly inside the vector's backing store rather that creating a temporary object and copying it into the vector.
- marksamman 12y agoemplace_back was added in C++11 which does just that: http://en.cppreference.com/w/cpp/container/vector/emplace_back http://en.cppreference.com/w/cpp/container/vector/emplace_ba...
- boomshoobop 12y agoIsn't Facebook itself an STD vector?
- general_failure 12y agoFunny :)
- general_failure 12y agoIn Qt, you can mark types as Q_MOVABLE_TYPE and get optimizations from a lot of containers
- kenperkins 12y ago> ... Rocket surgeon That's a new one. Usually it's rocket scientist or brain surgeon. What exactly does a rocket surgeon do? :)
- deleted 12y ago[deleted]
- krallja 12y agohttp://www.sensible.com/rsme.html http://www.sensible.com/rsme.html http://tvtropes.org/pmwiki/pmwiki.php/Main/ThisAintRocketSurgery http://tvtropes.org/pmwiki/pmwiki.php/Main/ThisAintRocketSur... http://www.urbandictionary.com/define.php?term=rocket%20surgery http://www.urbandictionary.com/define.php?term=rocket%20surg...
- 14113 12y agoIs it normal to notate powers using double caret notation? (i.e. ^^) I've only ever seen it using a single caret (^), in what presumably is meant to correspond to an ascii version of Knuth up arrow notation (https://en.wikipedia.org/wiki/Knuth's_up-arrow_notation https://en.wikipedia.org/wiki/Knuth's_up-arrow_notation). I found it a bit strange, and confusing in the article having powers denoted using ^^, and had to go back to make sure I wasn't missing anything.
- Sidnicious 12y agoIt's likely to distinguish it from XOR, which is the carat operator in many programming languages.
- kevin_thibedeau 12y agoIt is still weird. * * is the other established convention used for exponentiation in languages like Python, Ada, m4, and others.
- merraksh 12y agoBut * * has an entirely different meaning in C/C++.
- taejo 12y agoOnly as a prefix operator; * has a pointer-related meaning for prefix and an arithmetic meaning for infix, there's no reason * * shouldn't be the same.
- phs2501 12y agoThat'd be a syntactic ambiguity (in C-type languages) between "a * * b" being a to the power of b (i.e. "(a) * * (b)") or "a * * b" being a times the dereferencing of the pointer b (i.e. "(a) * (*b)"). You may be able to disambiguate this by prescedence, though it would be very ugly in the lexer (you could never have a STARSTAR token, it would have to be handled in the grammar) and would be terribly confusing.
- chickenandrice 12y agoGreetings Facebook, several decades ago welcomes you. Game programmers figured out the same and arguably better ways of doing this since each version of std::vector has been released. This is but a small reason most of us had in-house stl libraries for decades now. Most of the time if performance and allocation is so critical, you're better off not using a vector anyway. A fixed sized array is much more cache friendly, makes pooling quite easy, and eliminates other performance costs that suffer from std::vector's implementation. More to the point, who would use a c++ library from Facebook? Hopefully don't need to explain the reasons here.
- dbaupp 12y ago> More to the point, who would use a c++ library from Facebook? Hopefully don't need to explain the reasons here. Could you explain them for those of us not in the loop? Does Facebook have a bad reputation for C++?
- DonPellegrino 12y agoI would also like expanations, because Facebook actually has a good reputation when it comes to their compiled languages engineers. Their C++ and D engineers built HHVM/Hack, their OCaml engineers built some great analysis tools and much of the supporting code for the HHVM/Hack platform, etc., the list goes on, so I'd like to know why someone would want to avoid their C++ library based on the "Facebook" name only.
- chickenandrice 12y agoBecause Facebook also has a reputation of not playing nice with people, the rules, intellectual property, and so on. This is hardly a company anyone should support or trust and if you can't figure that out, I can't help you. As far as their work on HHVM, it was necessary due to failure by bad technology choices from the start. There's very little interesting about this work unless you somehow love PHP, want to make debugging your production applications more difficult, and refuse to address your real problems. I am 100% sure no one outside of the PHP community cares about anything Facebook has done in C++. Simply having a large company with lots of developers who might have even had good reputations elsewhere or even be smart doesn't mean much. Having worked in many places with lots of smart developers, I can tell you stories about too many geniuses in the room. Calling Facebook developers engineers is also about as apt as calling janitors sanitation engineers. We're programmers, or developers, or perhaps software architects at best depending on the position. I happen to have an EE and CS degree but given I do programming for a living, I'd hardly call myself an engineer. But we're way off topic :)
- pbw 12y agoAre there benchmarks, speedup? Seems strange to leave out that information or did I just miss it?
- shadytrees 12y agoYes! https://www.google.com/search?q=folly+facebook+benchmarks https://www.google.com/search?q=folly+facebook+benchmarks
- pbw 12y agoI don't see the results. Like a graph that shows std::vector vs. folly. I mean isn't that the entire point?
- shin_lao 12y agoI think the Folly small vector library is much more interesting and can yield better performance (if you hit the sweet spot). https://github.com/facebook/folly/blob/master/folly/docs/small_vector.md https://github.com/facebook/folly/blob/master/folly/docs/sma... From what I understand, using a "rvalue-reference ready" vector implementation with a good memory allocator must work at least as good as FBVector.
- johnwbyrd 12y agoIf you're spending a lot of time changing the size of a std::vector array, then maybe std::vector isn't the right type of structure to begin with...
- johnwbyrd 12y agoShow me a programmer who is trying to reoptimize the STL, and I'll show you a programmer who is about to be laid off. The guy who tried this at EA didn't last long there.
- xroche 12y agoThe STL is not optimized at all in this case, this is precisely the point. And like it or not, but Facebook has talented engineers to do that.
- richardwhiuk 12y agoThe STL isn't a library that can be optimized - it's a interface definition with expected complexity requirements. By it's nature (i.e. not tied to a platform) it doesn't have specific benchmark numbers. Specific implementations (e.g. MSVCRT, the GC++ implementation, the clang implementation) can be, and are.
- darkpore 12y agoGames and other high performance users generally used to stay away from the STL. Sony had their own version of STL which addressed some of these issues.
- maximilianburke 12y ago> The guy who tried this at EA didn't last long there. Paul Pedriana, the man responsible for the lions share of the EASTL work, started at Maxis before it was acqired by EA in 1997. EASTL has been in continuous development for more than 10 years.
- jeorgun 12y agoApparently the libstdc++ people aren't entirely convinced by the growth factor claims: https://gcc.gnu.org/ml/libstdc++/2013-03/msg00059.html https://gcc.gnu.org/ml/libstdc++/2013-03/msg00059.html
- thomasahle 12y agoThe factor-2 discussion is quite interesting. What if we could make the next allocated element always fit exactly in the space left over by the old elements? Solving the equations suggest a fibonacci like sequence, seeded by something like "2, 3, 4, 5". Continuing 9, 14, 23 etc.
- judk 12y agohe golden ratio and then rounding down. What's the point of putting 4 into the seed sequence?
- thomasahle 12y agoWithout 4, the sequence would be 2, 3, 5. Then the next value would be 9 by fibonacci. But that's bigger than the 5 we get from adding up all the unallocated pieces (2 and 3). We could use just 2, 3, 4 as a seed, but we can't use the fibonacci formula before the fourth element is added. Try some different seeds for yourself, it's trickier than you'd think.
- eps 12y agoThere was an academic paper in the 70s on Fibonacci-based allocation strategy... though I can't seem to find it now. So the idea is definitely not new :)
- judk 12y agoHow is it reasonable to expect that previously freed memory would be available later for the vector to move to?
- jlebar 12y agoIf you're interested in these sorts of micro-optimizations, you may find Mozilla's nsTArray (essentially std::vector) interesting. One of its unusual design decisions is that the array's length and capacity is stored next to the array elements themselves. This means that nsTArray stores just one pointer, which makes for more compact DOM objects and so on. To make this work requires some cooperation with Firefox's allocator (jemalloc, the same one that FB uses, although afaik FB uses a newer version). In particular, it would be a bummer if nsTArray decided to allocate space for e.g. 4kb worth of elements and then tacked on a header of size 8 bytes, because then we'd end up allocating 8kb from the OS (two pages) and wasting most of that second page. So nsTArray works with the allocator to figure out the right number of elements to allocate without wasting too much space. We don't want to allocate a new header for zero-length arrays. The natural thing to do would be to set nsTArray's pointer to NULL when it's empty, but then you'd have to incur a branch on every access to the array's size/capacity. So instead, empty nsTArrays are pointers to a globally-shared "empty header" that describes an array with capacity and length 0. Mozilla also has a class with some inline storage, like folly's fixed_array. What's interesting about Mozilla's version, called nsAutoTArray, is that it shares a structure with nsTArray, so you can cast it to a const nsTArray*. This lets you write a function which will take an const nsTArray& or const nsAutoTArray& without templates. Anyway, I won't pretend that the code is pretty, but there's a bunch of good stuff in there if you're willing to dig. http://mxr.mozilla.org/mozilla-central/source/xpcom/glue/nsTArray.h http://mxr.mozilla.org/mozilla-central/source/xpcom/glue/nsT...
- nly 12y ago> One of its unusual design decisions is that the array's length and capacity is stored next to the array elements itself. GNU stdlibc++ does this for std::string so you get prettier output in the debugger. The object itself only contains a char*.
- byuu 12y agoSeems like that would prevent small string optimization (a union of a small char array and the heap char pointer.) That lets me store about 85% of the strings in my assembler without any heap allocation, and is a huge win in my book.
- malkia 12y agoFor big vectors, if there is obvious way, I always hint vector with reserve() - for example knowing in advance how much would be copied, even if a bit less gets copied (or even if a bit more, at the cost of reallocation :().
- jheriko 12y agoi used to be a big fan of this sort of stuff, but the better solution for many of the problems described is to avoid array resizing. if std::vector is your bottleneck you have bigger problems i suspect. reminds me a bit of eastl as well... which is much more comprehensive: http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2007/n2271.html http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2007/n227...