6 ms·
Constrained embedded systems cover a broad range of things up to and including your phone. There's few things I hate to see more than a flat profile from memory
by vvanders 6y ago
Constrained embedded systems cover a broad range of things up to and including your phone. There's few things I hate to see more than a flat profile from memory allocation or cache misses.
The good news is that flatbuffers[1] is a reasonable replacement for most of my use cases. In particular being able to mmap() them directly is a wonderful thing that you can't do with protobufs in addition to being very allocation sparse.
[1] https://google.github.io/flatbuffers/ https://google.github.io/flatbuffers/
- kentonv 6y ago> Constrained embedded systems cover a broad range of things up to and including your phone. No, modern phones are certainly not constrained in the way I meant, and I don't think you could call them "embedded" either. The common programming languages used on phones are very memory-allocation-friendly. > There's few things I hate to see more than a flat profile from memory allocation or cache misses. I think you may be arguing a different point, or a different level of extremity of the point. Reducing memory allocation to optimize performance is a fine thing that everyone does. The person I was replying to, though, seemed to be asserting that libraries should completely avoid allocating memory for themselves. > The good news is that flatbuffers[1] is a reasonable replacement for most of my use cases. In particular being able to mmap() them directly is a wonderful thing that you can't do with protobufs in addition to being very allocation sparse. Yeah... I'm the author of Cap'n Proto, which has the same property, and predates Flatbuffers.
- vvanders 6y ago> The common programming languages used on phones are very memory-allocation-friendly. I would hardly classify Dalvik or ART as "allocation friendly", they don't perform escape analysis and if you do it constantly you'll be in a world of constant hard GC pauses. Multiple times over the years I've had to build free-lists in Java to avoid this specific problem. Same for C++ if you use one of the built-in generic memory allocators. The fastest new/delete are the ones that you don't call. For what it's worth I tend to agree with the grand-parent thread. The lack of awareness of allocations, cache-invalidation via indirection are a significant contributor to why we see software clawing back hardware wins across the years on these platforms.
- cbsmith 6y ago> I would hardly classify Dalvik or ART as "allocation friendly", they don't perform escape analysis and if you do it constantly you'll be in a world of constant hard GC pauses. Multiple times over the years I've had to build free-lists in Java to avoid this specific problem. I don't know what you're doing, but I've generally found free-lists to be a net performance negative in Java libraries. Time and again, I've been called in to "optimize" Java code that uses them, and usually by simply removing them I can get rid of the performance problems entirely.
- vvanders 6y agoDalvik and ART have different behaviors than traditional Java VMs. Dalvik is aimed at low memory devices, ART trades some memory for speed but still is significantly different(i.e. it doesn't do escape analysis). As always benchmarking on hardware is the way to confirm this but I've had numerous cases where free lists or pre-allocated arrays has gained 10-30x performance improvements. Flatbuffers in particular excels here since once you have the ByteBuffer in memory you can immediately start accessing data without needing to do any extensive parsing. Even the official Android docs are very explicit[1] on the allocation point. Allocations are not cheap and even with generational collection you still will blow the 16.6ms frame window for 60 FPS if any of your operations allocate excessively. [1] https://developer.android.com/training/custom-views/optimizing-view#less https://developer.android.com/training/custom-views/optimizi...
- cbsmith 6y agoSure, with flatbuffers, I grok the case (though where possible I just use an off-heap mmap buffer for them anyway), but that's a very specific case.
- musicale 6y agoYou may be right - presumably iPhones have fewer GC pauses (though you can still have VM, loading, compression/decompression, network, and other pauses.) Be that as it may, lots of people still manage to use Android devices.