3 ms·
No worries, too much work tends to do that to you, I recommend some beers to fix the state of mind. Regarding the implemention, I presumed you asked about the
by tuxalin 7y ago
No worries, too much work tends to do that to you, I recommend some beers to fix the state of mind.
Regarding the implemention, I presumed you asked about the lock-free logic, apart from m_currentIndex which is used with the m_commands vector, the LinearAllocator also has to be lock-free, see it's alloc function.
The allocator is used for storing the command packets (along with their command data) and also their auxiliary data (if they have that).
Another implementation detail of the allocator is that it has two modes, an auto-alignment mode which requires two atomic operations (load acquire and compare_exchange_weak) and will be slower but lower memory usage, therother mode uses the alignment of the CommandPacket class which will be faster (one fetch_add) but consume more memory.
And another detail is the CommandPacket uses a intrusive list approach, that's how you actually chain multiple commands together and do an indirect function call (instead of virtual or other blasphemies) with dispatch of the command.
Hope that clears some things up about it.
- dragontamer 7y agoI was getting ready to type something up, but you've basically answered my questions! > an auto-alignment mode which requires two atomic operations (load acquire and compare_exchange_weak)and will be slower but lower memory usage, ther other mode uses the alignment of the CommandPacket class which will be faster (one fetch_add) but consume more memory. That's the main thing I was wondering about, especially the relative speeds of the implementation.
- tuxalin 7y agoGlad to hear that! I dont have anymore the exact numbers, usually when you have around 1000 commands/calls you won't really feel a difference, in that case even using std::sort is fast enough and there's no point in using radix sort, however when you go over 10k commands/calls you'll feel it (can be up to a msec on the dispatch side). You can check this great post that provides numbers (see part 4): https://blog.molecular-matters.com/2014/11/06/stateless-layered-multi-threaded-rendering-part-1/ https://blog.molecular-matters.com/2014/11/06/stateless-laye... The performance should be very similar and it even goes further by adding thread local storage to reduce false sharing even more, which I didn't add but should be considered as it's easy to do.
- dragontamer 7y agoOkay, now that I've had some time to look at other things and come back to this, lemme offer some more detailed questions. COMMAND_QUAL::addCommand uses fetch_add(1, std::memory_order_relaxed); It seems like this is sufficient for memory-consistency for the list, but I'm worried about the eventual sort() command. In particular, your update in "addCommand" is as follows: const uint32_t currentIndex = m_currentIndex.fetch_add(1, boost::memory_order_relaxed); // asserts removed CommandPair& pair = m_commands[currentIndex]; pair.cmd = packet; pair.key = key; How does COMMAND_QUAL::sort() know that the other threads are "done adding" and that pair.key is a valid value for comparison? My race-condition would be "sort" has been called, while some thread is still setting pair.key=key, which may corrupt the command/key packet. ----------- EDIT: Oh, you answered this in your blogpost. Lol. You assume that the user waits and safely calls sort after waiting. Okay, that's perfectly acceptable in my eyes. I'll leave this post up for discussion purposes, even if I answered my own question.
- tuxalin 7y agoCorrect, you could also use a double buffering approach to avoid that, if possible, depending on your use case. Feel free to reach out if you have any other questions.