3 ms·
Perl seems to pass nearly all the tests (including uppercasing baffle): $ perl -E 'use utf8; binmode STDOUT, ":utf8"; say uc("baffle");' BAFFLE The only fail
by jbert 13y ago
Perl seems to pass nearly all the tests (including uppercasing baffle):
$ perl -E 'use utf8; binmode STDOUT, ":utf8"; say uc("baffle");'
BAFFLE
The only failure I can see is that it treats "no<combining diaresis>el" as 5 characters (so reports length as 5 and reversing places the accent on the wrong character). That's documented here: http://perldoc.perl.org/perluniintro.html#Handling-Unicode http://perldoc.perl.org/perluniintro.html#Handling-Unicode "Note that Perl considers grapheme clusters to be separate characters"
All else seems to work though (including precomposed/decomoposed string equiality etc). The docco also says that perl's regex engine with Do The Right Thing with matching the entire grapheme cluster as a single char.
- cursork 13y agoPerl is actually very good with Unicode. Note that a character is "The smallest component of written language that has semantic value" according to the Unicode glossary - I'd say Perl respects that meaning. As noted in the docs, graphemes can be handled with \X in regular expressions (although admittedly that's not pretty): my $length = 0; $length++ while $dec =~ /\X/g; Note that a grapheme is defined as "A minimally distinctive unit of writing in the context of a particular writing system" - i.e. context is required to determine what a grapheme actually is. A few others have pointed that out... Given the definitions from Unicode, Perl does a pretty good job (esp. when using Unicode::Normalize to normalize input). http://www.unicode.org/glossary/ http://www.unicode.org/glossary/
- llimllib 13y agoPython 3 gets that one, but python 2.7 doesn't: $ python3 -c 'print("baffle".upper())' BAFFLE $ python -c 'print "baffle".upper()' BAfflE $ python -c 'print u"baffle".upper()' BAϬE
- deathanatos 13y agoIt's interesting that you get BAFFLE for the first one. I get the same result in both 3 and 2. Note first that the reason you get "BAϬE" is a bit of garbarge-in garbage-out. Strangely, the interpreter isn't rejecting that with the typical "SyntaxError: Non-ASCII character <char> in file" error; instead, it appears to be assuming ISO-8859-1, and then performing .upper(). You can fix that: python2 -c '# coding: utf-8 print u"baffle".upper()' (Note, of course, that the #coding needs to match your terminals encoding, which is likely UTF-8, but it isn't guaranteed.) That, for me, prints "BAfflE" in both Python 2 and 3 (adjusting for 3 by adding parens around print, and removing the u prefix on the literal.) I'm on Python 3.2, so perhaps 3.3 does better. (I'm behind on updates, but last I did update, Gentoo stable was still on 3.2.)
- llimllib 13y agoInteresting. I am on python 3.3, but I don't know if it's the updated interpreter that fixes the bug.
- jmah 13y agoAlso Cocoa's NSString: [@"baffle" uppercaseString]; // @"BAFFLE"