3 ms·
I'm going to add jemalloc, lockless and scalloc in the benchmark comparisons soon(ish). Regarding the files, only rpmalloc.[h|c] is needed. The malloc.c file i
by maniccoder 10y ago
I'm going to add jemalloc, lockless and scalloc in the benchmark comparisons soon(ish).
Regarding the files, only rpmalloc.[h|c] is needed. The malloc.c file is just for overriding the malloc family of entry points, and provide automatic init/fini glue.
The 16 byte alignment is just a note that all returned blocks are naturally aligned to 16 byte boundaries due to the fact that all size buckets are a multiple of 16 bytes internally. It is needed for SSE instructions, which is why I mention it.
- olliej 10y agoAny plan to add the allocators used in modern browsers? bmalloc (webkit), which ever tcmalloc variant chrome currently uses, the current firefox allocator, I have no idea if edge has an allocator that is not the default windows one?
- maniccoder 10y agoI will look into it
- rurban 10y agoFair would be a comparison single-threaded to ptmalloc2 (glibc). I assume the slowdown will be ~5%. And the size granularity is split into 3 (small, medium, large) which is fine, but other's have finer granularities. So benchmarks with object systems would be nice, not just random data.
- maniccoder 10y agoThere already is a benchmark against glibc single threaded, compile the rpmalloc-benchmark repository linked in the readme and run the runall script. Compare rpmalloc 1 thread case against crt 1 thread (with crt being the C runtime, i.e glibc in this case). Graphs and analysis for a Linux system will be added soon, but a spoiler: rpmalloc 22.7m memory ops/CPU second, 26MiB bytes peak, 15% overhead jemalloc 21.3m memory ops/CPU second, 30MiB bytes peak, 33% overhead tcmalloc 20.4m memory ops/CPU second, 30MiB bytes peak, 32% overhead crt 13.5m memory ops/CPU second, 24MiB bytes peak, 7% overhead rpmalloc is faster and has less overhead than both jemalloc and tcmalloc, and a lot faster than glibc at the cost of extra memory overhead. The reasoning behind the granularities is that 16 bytes is a reasonable granularity for smaller objects (below 2KiB in size). Medium sized requests (between 2KiB and 32KiB) have a 512 byte granularity, while large requests (up to 2MiB) have a 62KiB granularity. If you have a suggestion for additional benchmark models I would be happy to add an implementation for it.
- rurban 10y agoOuch, so this looks like the final coffin into ptmalloc2, and esp. a shout to emacs, which still hasn't finished their portable dumper. But not much missing there, should be ready soon. (I got a not working branch on my GitHub) Hope glibc will adopt this soon.