13 ms·
So someone used lower() from an unspecified version of Unicode when the standard was very specific about which to use. And they say "There's also a database of
by ghusbands 1mo ago
So someone used lower() from an unspecified version of Unicode when the standard was very specific about which to use. And they say "There's also a database of Unicode 3.2.0 data available on every version of Python (unicodedata.ucd_3_2_0) specifically for the StringPrep and IDNA algorithms", so the right version is available.
And then the fix is to hardcode a bunch of special cases which again depend on exactly which version of Unicode is in use, and so will break again in the same way in future, rather than just using the right version?
- akoboldfrying 1mo agoI'd assumed that the updating of the B3 dict takes place at runtime, either during module initialisation or lazily at the first call (that is, roughly as late as possible -- certainly, well after the wheel was built), just as a perf optimisation. But then I realised the code does (and must do) a lookup in the B3 table for each character anyway, so there doesn't seem to be any point. I suppose it means they can load the full 3.2.0 table once, use it to discover the exceptions and then immediately evict it from memory, keeping only the presumably smaller and faster-to-query B3 table of exceptions, but this seems pretty marginal...