3 ms·
Short answer: \d includes all the Unicode characters from http://www.fileformat.info/info/unicode/category/Nd/list.htm http://www.fileformat.info/info/unicode/
by Aurel1us 13y ago
Short answer: \d includes all the Unicode characters from http://www.fileformat.info/info/unicode/category/Nd/list.htm http://www.fileformat.info/info/unicode/category/Nd/list.htm
- wging 13y ago...at least in C# regexes.
- ars 13y agoAnyone know if this happens in other languages?
- yahelc 13y agoDoesn't appear to in JavaScript: "੧".match(/\d/); //null (Incidentally, this may explain the finding from http://stackoverflow.com/a/16622773/172322 http://stackoverflow.com/a/16622773/172322, as to why adding the RegexOptions.ECMAScript flag in the C# code eliminates the performance gap)
- deskglass 13y agoNor in python: print re.match(r'\d','੧') None
- wulczer 13y agoit does when using the re.U flag re.match(r'\d', u'੧', re.U) <_sre.SRE_Match at 0x3070ac0> sys.version 2.7.3 (default, Mar 4 2013, 14:57:34) \n[GCC 4.7.2]
- Falling3 13y agoYes, but not by default.
- tcas 13y agoAlso, when using Python 3.2 it seems to be the default behavior Python 3.2.3 (default, Oct 19 2012, 20:10:41) [GCC 4.6.3] on linux2 Type "help", "copyright", "credits" or "license" for more information. >>> import re >>> re.match(r'\d', '੧') <_sre.SRE_Match object at 0x7f188f6d4850>
- xudongz 13y agoNot true for Go http://play.golang.org/p/ls96RxJxpz http://play.golang.org/p/ls96RxJxpz
- nknighthb 13y agoI would be reluctant to rely on this until the Go documentation is clearer on the intended behavior. Right now it's very poorly specified. The regex doc[1] talks about "same general syntax" as Perl, but points to [2], which doesn't seem to understand what it's saying, describing '\d' in terms of its "Perl" meaning, but then saying that it's [0-9]. [1] http://golang.org/pkg/regexp/ http://golang.org/pkg/regexp/ [2] https://code.google.com/p/re2/wiki/Syntax https://code.google.com/p/re2/wiki/Syntax
- laumars 13y agoAs a Perl developer that's been making the switch to Go, I've been caught out a few times with Go's no-so-Perl-like regular expression syntax. In fact I wish I knew about your 2nd link before now, because that could have saved me a few hours over recent months.
- jamesmiller5 13y agoConsidering go's vocal support for UTF8 I'm surprised at this behavior and curious to the reason for excluding it.
- masklinn 13y agoSupporting UTF8 and correctly handling unicode are very, very different beasts. The former is absolutely trivial, the latter is extremely difficult. Go is vocal about the former, but seems to not give a shit about the latter.
- snogglethorpe 13y agoI'm not particularly fond of go, but "correctly handling unicode" can be subjective and case-dependent... I think making only minimal guarantees and punting to the application is often the only sane course.
- jeltz 13y agoHappens in Perl but not ruby or PostgreSQL.
- pfedor 13y agoDoesn't happen in Perl for me: pfedor@Pawels-iMac:~$ perl -ne 'print "Digit!\n" if /\d/' af 3 Digit! 23fa3 Digit! asdf ١٢٣٤٥٦٧٨٩۰۱۲۳۴۵۶۷۸۹ ৩৪৫৬৭৮৯੦੧੨੩੪੫੬੭੮੯૦૧૨૩૪૫ ୧୨୩୪୫୬୭୮ ౨౩౪౫౬౭౮౯೦೧೨೩೪೫೬೭೮೯൦൧൨൩൪൫൬൭൮൯๐๑๒๓๔๕๖๗๘๙໐໑໒໓ 234 Digit! (perl from Macports and perl from /usr/bin/perl behave the same in this respect.)
- xonea 13y agoYou have to tell to interpret stdin as UTF-8 (flag -C) - then it works: https://news.ycombinator.com/item?id=5734641 https://news.ycombinator.com/item?id=5734641
- pfedor 13y agoGood to know, thanks. I'd argue that perl gets it right--as the default behavior, this behavior would gravely violate the principle of least surprise, but for the 0.01% of people who want \d to match ੧, there's no harm to making it available as an option you need to specifically request.
- LawnGnome 13y agoHappens in PHP only if you enable Unicode regex handling via the /u modifier and are running libpcre 8.10 or later (which corresponds to PHP 5.3.4 and later, assuming you're using the bundled libpcre): http://3v4l.org/QD3k0 http://3v4l.org/QD3k0
- bodyfour 13y agoIf you're using pcre directly from C code, this is controlled by specifying the PCRE_UCP flag to pcre_compile(). By default, \d and friends only match ASCII characters even if the PCRE_UTF8 flag is set.
- Falling3 13y agoExactly what I was thinking. Doesn't in Ruby: /\d/.match "੧" #=> nil
- Argorak 13y agoJust for reference: /\p{Digit}/.match "੧" => #<MatchData "੧">
- cwmma 13y agoAll the same speed in JavaScrip http://jsperf.com/regexcwm/2 http://jsperf.com/regexcwm/2
- jrabone 13y agoNot true for Java. Docs even say: \d A digit: [0-9] \p{Digit} A decimal digit: [0-9] which is actually somewhat depressing. I'd expect the named class to include the full Unicode digit set. It's surprising to see: ab1234567890cd matched 1234567890 ab𝟣𝟤𝟥𝟦𝟧𝟨𝟩𝟪𝟫𝟢cd no match from code using Pattern.compile("(\\p{Digit}+)"); EDIT: and perhaps more surprising to see in the logs: Exception in thread "main" java.lang.NumberFormatException: For input string: "𝟤𝟥𝟦𝟧" at java.lang.NumberFormatException.forInputString(NumberFormatException.java:48) at java.lang.Integer.parseInt(Integer.java:449) That'll keep someone guessing for a while...
- nspragmatic 13y agoIt happens in Objective-C: NSString *pattern = @"\\d", *string = @"੧"; NSRegularExpression *regex = [NSRegularExpression regularExpressionWithPattern:pattern options:NSRegularExpressionCaseInsensitive error:nil]; NSUInteger numMatches = [regex numberOfMatchesInString:string options:0 range:NSMakeRange(0, [string length])]; numMatches ? NSLog(@"%@ found by %@", string, pattern) : NSLog(@"%@ not found", string); // 2013-05-20 09:38:42.650 Regexperiment[17848:c07] ੧ found by \d
- ars 13y agoIs that actually a good thing? If I'm using \d to validate numbers (for example to check before string to int conversion, or IP address, phone number, or any other use), other unicode digits are not helpful to me. It's great to support unicode, but I don't think the \d should have been extended this way. Add a \ud or something.
- Tuna-Fish 13y agoGiven that the category is specifically "decimal digit", I think it's good, so long as the number parsing code accepts them all too.
- dllthomas 13y agoYes. Assuming that, it's good. I think that assumption is likely to be invalid in many cases, though.
- bellbind 13y agoIf you use a preg engine you can add the /a modifier which excludes unicode chars from matches.
- chebucto 13y agoMaybe specify the subset of unicode you're expecting in the headers, and have the compiler do the nitty gritty?
- rmc 13y agoYes it's a good thing. There are other places in the world that don't just use ascii. If you want European style numbers just use [0-9]
- hkmurakami 13y agooh wow I had no idea that "full width digits" can actually be handled properly. (U+FF10 ~ U+FF19)
- coldtea 13y agoOr improperly. If you expect \d to be a shorthand for 0-9, your string can also contain junk.