3 ms·
The parser finds the start/end of sub-structures and doesn't necessarily process them all. So for large structures of which you only need a subset, you're only
by Tobani 11y ago
The parser finds the start/end of sub-structures and doesn't necessarily process them all. So for large structures of which you only need a subset, you're only doing the work you need to do. For a large structure in which you need all of the data, there is probably less of gain.
- valarauca1 11y agoThat plus the parsing is done in SIMD so even in UTF-8 you're processing 4 unicode points per CPU cycle [1], instead of 1 in most traditional parsers. [1] This is a board generalization not necessarily true for all SIMD opcodes.
- thechao 11y agoHm. I guess when thinking "SSE parsing" I didn't go with 4/8-wide parsing. I was thinking that they'd be grabbing 16/32-bytes, doing a compare against a fixed constant literal of 16/32-copies of, say, '{', or '}', then extracting the index of the match, exactly. Something like this: d := _mm_loadu_ps <json> b := _mm_loadu_ps <token: '{'> n := _mm_cmpistrc(d, b) You'd have to be clever to skip around strings ("... \"sonuva ... "), but once that's handled, you'd have significant speed ups to scan for ',', '{', '}', etc. I think the double-quote escape might look something like this: d := _mm_loadu_ps <json> q := _mm_loadu_ps <token: '"'> e := _mm_loadu_ps <token: '\\'> n := _mm_cmpistrc(d, q) m := _mm_cmpistrc(d, e) if m+1 == n: branch-to-top ... process ... Looks like cmpistrc has 1/2 reciprocal throughput. If you unrolled the loop 8 deep, you're probably looking at 10c per 16bytes scanned.