4 ms·
> In C++, this can be done in zero-copy fashion with string_view. In C, every string has to be null-terminated. Thus, you need to either manipulate the original
by thethirdone 6y ago
> In C++, this can be done in zero-copy fashion with string_view. In C, every string has to be null-terminated. Thus, you need to either manipulate the original buffer, or copy it over. I elected the latter.
Is the null-termination requirement due to returning the value of cells as C strings? I couldn't find an easy answer looking through the provided code. Otherwise if its just an internal detail, I would think a (pointer,length) string type would be ideal to allow you to use the buffer without modification or copying.
> It is about 2x slower than csv2. This is expected because we need to null-terminate strings and copy them to a new buffer.
I think it is strange to settle for a 2x slowdown in order to not modify the original buffer. Substituting a null byte for the separator / newline should be simpler than copying and null-terminating. So is keeping the original buffer free from modifications important? In most cases preserving the original buffer seems unimportant.
- willvarfar 6y agoIf you don’t modify the original buffer then it can me an mmapped file..?
- thethirdone 6y agoI was thinking about using MAP_PRIVATE to prevent write-back. I not sure that doesn't come with a performance penalty on write, but it should at least be better copying in userspace.
- saagarjha 6y agoThat would likely trigger the “copy” part of copy-on-write and not be particularly fast.
- thethirdone 6y agoDo you have any benchmarks that show that? It seems plausible, but a lot of COW systems optimize for the case where there is only 1 reference and modify in-place. I don't MMAP much, but it would be nice to know if Linux doesn't do the optimization I think is possible. In this particular case, I would hope it only copies to memory from disk. If only one program is accessing a file, I would expect the "copy" of copy-on-write to be the copy to memory.
- saagarjha 6y agoSadly I do not, and I’m on a device running Darwin that I can’t readily do benchmarks on :( But yes, if you are the exclusive owner of a page you won’t get an extra copy. However, you’re still not really going any faster, as this is just the same as the kernel giving you the file in a buffer that you can freely modify…
- sgtnoodle 6y agoWhen reading mostly sequentially, is a memory mapped file any faster than reading into a buffer? The data is presumably still copied into memory, it's just done on a page fault rather than an explicit function call.
- deleted 6y ago[deleted]
- sgtnoodle 6y agoApparently there is one less memory copy when using mmap, since read() copies from the intermediate buffer that mmap provides access to. Neat! https://medium.com/@sasha_f/why-mmap-is-faster-than-system-calls-24718e75ab37 https://medium.com/@sasha_f/why-mmap-is-faster-than-system-c...
- bleepblorp 6y agoMy assumption is that support for C strings (null terminated) was a design requirement. Note that just replacing separators and newlines with nulls isn't enough to parse CSVs because the file format supports escape sequences and optional quotes around values. If you're modifying a buffer in place, you'll still be stuck copying data to replace (longer) escape sequences with their (shorter) final values. I don't understand the line you've quoted from the article regarding why this code is slower than csv2. The csv2 code seemingly does copy from the original buffer rather than manipulating in place. Indeed, the csv2 parser copies data from a buffer into a C++ STL container object, which I would expect to be quite an expensive operation. Something does not add up here, either in my understanding or in the mechanics/description of the benchmark.
- thethirdone 6y ago> Note that just replacing separators and newlines with nulls isn't enough to parse CSVs because the file format supports escape sequences and optional quotes around values. If you're modifying a buffer in place, you'll still be stuck copying data to replace (longer) escape sequences with their (shorter) final values. I had forgotten that escape sequences needed to also be changed. Unescaping in place does make it non-trivial. > I don't understand the line you've quoted from the article regarding why this code is slower than csv2. The csv2 code seemingly does copy from the original buffer rather than manipulating in place. Indeed, the csv2 parser copies data from a buffer into a C++ STL container object, which I would expect to be quite an expensive operation. Something does not add up here, either in my understanding or in the mechanics/description of the benchmark. I'm not familiar with the csv2 parser, but if what you say is true, I guess the slowdown would come from doing a csv2 doing a larger copy(s) and the given code doing many smaller copies.
- formerly_proven 6y ago> I had forgotten that escape sequences needed to also be changed. Unescaping in place does make it non-trivial. Assuming escapes are rare, backshift unescaping is totally an option, though it scales poorly (interpreting escapes is linear, backshifting is quadratic) and is thus susceptible to slowdowns with malicious inputs. It does need very little code and no extra memory at all, though.