4 ms·
I thought you could use memoryview over a string to get rid of that allocation even in 2.7 https://docs.python.org/3/library/stdtypes.html#memoryview https://d
by daniel_rh 9y ago
I thought you could use memoryview over a string to get rid of that allocation even in 2.7
https://docs.python.org/3/library/stdtypes.html#memoryview https://docs.python.org/3/library/stdtypes.html#memoryview
- nostrademons 9y agoAllocations are challenging for HTML parsers, even in C, because of the presence of entity references and case-normalization of attribute & tag names. That means that a lot of the time when you think you ought to be able to just use a slice or memoryview into the original source text, you can't; for example, if any of your text nodes contains < ('<') or &ldquo (smart double quote), you can't use the original source buffer, because you're supposed to have decoded the entity to a unicode character, which will leave the string a different length. This happens stupidly often in real HTML. I initially had the API for Gumbo use string slices a lot more than the final released API, and then found that I couldn't do it and needed to allocate in order to maintain correctness. I'd done a patch that arena-allocated all memory used in the parse, which gave a fairly significant CPU speedup, but it also bloated max memory usage in ways that some clients found unacceptable, so I never merged it. Small C strings at least are quite lightweight; Python strings have a lot of additional overhead, and much of the PyObject structure itself requires chasing pointers.
- gsnedders 9y agoOnly over bytes objects, not over unicode objects.