3 ms·
Simutf is designed for performance. The function is designed for you to provide a buffer of sufficient length. It will return the number of bytes it actually
by clausecker 3y ago
Simutf is designed for performance. The function is designed for you to provide a buffer of sufficient length. It will return the number of bytes it actually needed. You can call utf8_length_from_latin1 to find the required buffer length, or provide a buffer that is twice as large as the input buffer, which will always suffice. We intend that users build high-level APIs around these functions if so
desired.
> Also, for all that simdutf is really quite fast, it's really not obvious how to transcode latin-1 to utf-8 in a single pass, because you apparently need to compute the length first.
You do this by having the caller pre-allocate a buffer that will always be long enough (twice the length of the input). Some bytes of the buffer may end up being unused, but that's usually acceptable.
- amluto 3y ago> Simutf is designed for performance. I know. When I was less experienced, I believed more in this kind of isolated performance. But... > You can call utf8_length_from_latin1 to find the required buffer length For smallish input (that fits in L1 cache, for example), performance might be dominated by the cost of filling the cache, in which case two passes might not be that bad. For large inputs, especially at the kinds of speeds that simdutf8 achieves, memory bandwidth is likely to dominate or come close, and two passes is not so great. In either case, I bet that a high quality length-checked implementation of a function like convert_latin1_to_utf8 could be considerably faster than calling utf8_length_from_latin1 and then convert_latin1_to_utf8. > or provide a buffer that is twice as large as the input buffer, which will always suffice. and will always be fast as long as wasting up to half the memory is not a big deal. But in many cases, that factor of 2 of memory usage is a big deal. So I hate to say it, but I tend to think that simdutf's transcoding design is based on benchmarking the wrong thing. (The validation benchmark design is fine.)
- RaisingSpear 3y ago> wasting up to half the memory OSes generally implement overcommit - allocating a bunch of memory doesn't mean it's actually being used.
- amluto 3y agoSorry, but allocating approximately 2x the needed size for each string is not going to get covered up well by overcommit unless you are directly mmapping and munmapping (or DONTNEEDing or similar) each allocation and each allocation is more than 8kB. (Probably well over 8kB, and for actual good efficiency in the presence of huge pages, over 4MB so there’s a chance that half of it is unused.) I doubt this workload exists. Overcommit is kind of nifty but not at all magical. I bet a really good AVX-512 transcoder with output size checking could be almost as fast as the OP.
- RaisingSpear 3y ago> unless you are directly mmapping and munmapping (or DONTNEEDing or similar) each allocation You don't free up the memory you're not using? Sounds like you're doing it wrong then. > I bet a really good AVX-512 transcoder with output size checking could be almost as fast as the OP. Submit your pull request and we can evaluate your bet.