3 ms·
The first sentence sounds as if they modified the implementemention of str.lower(). That would be bonkers, but that's not what they did. https://github.com/pyt
by jwilk 2mo ago
The first sentence sounds as if they modified the implementemention of str.lower(). That would be bonkers, but that's not what they did.
https://github.com/python/cpython/commit/7e109d084d55e7eb https://github.com/python/cpython/commit/7e109d084d55e7eb
The important part is:
# B.3 is mostly Python's .lower, except for a number
# of special cases, e.g. considering canonical forms.
+# To enforce Unicode 3.2.0 behavior of .lower instead of
+# whatever Unicode version is included with Python we
+# add unassigned or newly case-folding codepoints to
+# the exception map, too.
b3_exceptions = {}
for k,v in table_b2.items():
if list(map(ord, chr(k).lower())) != v:
b3_exceptions[k] = "".join(map(chr,v))
+for cp in range(0x110000):
+ ch = chr(cp)
+ # Assigned in current Unicode version
+ # and supports case folding, but not
+ # explicitly in B.2 or B.3 tables.
+ if (unicodedata_current.category(ch) != "Cn"
+ and ch.lower() != ch
+ and cp not in table_b2
+ and cp not in table_b3):
+ b3_exceptions[cp] = ch # Identity.
- quietbritishjim 2mo agoThat fragment doesn't mean much in isolation. You've just said that they didn't modify str.lower (because "that would be bonkers") but you've posted a fragment which, for all we know, is part of the str.lower implementation.