4 ms·
My point was that a malloc vs alloca in a full blown IO path, where you have sockets and posibly disk fsyncs is not all that significant, if at all. I understan
by int0x80 8y ago
My point was that a malloc vs alloca in a full blown IO path, where you have sockets and posibly disk fsyncs is not all that significant, if at all. I understand that journald cares about performance. Micro-optimizations like malloc vs alloca are the least of your problems. Batching/buffering messages, an asyncronous model, plain O_APPEND, different files for logs/metadata etc like others and you have said ITT is what will matter. You can even use a sane dynamic string library with preallocation if you care that much. I haven't looked at the src, to be honest, but am I wrong?
- newnewpdro 8y agojournald does all its journal writing via mmap. When the window to write into is already mapped (it caches the mappings), there's relatively little per-message syscall overhead required. Prior to process metadata caching, the per-message CPU cost was dominated by acquiring the sender process metadata from /proc, which was dominated by syscalls but the malloc()+free() overhead of all that junk was not insignificant. I haven't done any journald profiling or hacking since those process metadata caching changes landed, but I imagine it's brought the per-message syscall overhead down to little more than pulling the message off the socket. When I was experimenting with adding metadata caching, once caching could amortize the /proc-accessing syscall overhead across many messages the major pain was all the allocator activity in moving that metadata around and composing the KEY=VALUE clauses from that metadata. It was a death from a thousand cuts kind of situation, happening for every message. The stack is the convenient place to efficiently do much of this granular, ephemeral stuff. But it's obviously inappropriate for allocations of user-controlled/unbounded size.
- int0x80 8y agoThanks for the interesting and detailed info. I mean the cost of hitting the disk/socket. The actual IO. There is also the syscall overhead but that doesn't matter /that much/ with IO usually, in comparation with the actual IO (unless, again, you are doing something very wrong - very small buffers etc ...). This of course depends on the deatils. And yeah, if it uses mmap then that syscall problem is no more. Problem I was refering was actual IO cost per message (fsync/msync per message?) to mantain consistency instead of caching for example or batching and grouping IO at the expense of consistency guarantees (maybe) or coherent data re. messages in realtime. I agree that the stack is usefull and I see your point. But I would almost never resort to alloca. That is not a solution IMO. If you have that kind of costs with the joining of keys there are other solutions, I think. What about a dynamic string library, that over allocates (maybe a big chung on the first time in the 'new message' path) and then just uses the extra memory until you run out. Classic (and silly) strategy of allocating like n << 4 bytes every time you run out of space. This way you really just have a few allocations instead of hundreds maybe making it a non-issue (at the cost of prob. unused extra memory). Not as efficent as alloca but very close. I you want to get fancy you can have efficent malloc pools to. But I think a semi-clever dynamic string lib would be able to handle the problem best.