5 ms·
>>> int("۲۶۷۹") 2679 Thats pretty cool.
by ianbertolacci 9y ago
>>> int("۲۶۷۹")
2679
Thats pretty cool.
- Tistron 9y agoIndeed. Seems to be a thing of python 3. $ ruby -v ruby 2.2.6p396 (2016-11-15 revision 56800) [x86_64-linux-gnu] $ ruby -e "puts \"۲۶۷۹\".to_i" 0 $ python2.7 -c "print(int(\"۲۶۷۹\"))" Traceback (most recent call last): File "<string>", line 1, in <module> ValueError: invalid literal for int() with base 10: '\xdb\xb2\xdb\xb6\xdb\xb7\xdb\xb9' $ python3.5 -c "print(int(\"۲۶۷۹\"))" 2679 $ node -v v6.9.1 $ node -p "parseInt(\"۲۶۷۹\")" NaN
- plus 9y agoIt works in Python 2, you just have to ensure that your string is unicode. I can't get it to work directly from the command line for some reason, but if you open an interactive Python session it works: Python 2.7.12 (default, Dec 14 2016, 13:32:53) [GCC 4.9.3] on linux2 Type "help", "copyright", "credits" or "license" for more information. >>> print(int(u'۲۶۷۹')) 2679
- Tistron 9y agoHmm, interesting. What's up with the terminal here? The shell is doing something to the text. $ python2.7 -c "print(int(u'۲۶۷۹'))" Traceback (most recent call last): File "<string>", line 1, in <module> ValueError: invalid literal for int() with base 10: '\xdb\xb2\xdb\xb6\xdb\xb7\xdb\xb9' $ python2.7 Python 2.7.13 (default, Jan 03 2017, 17:41:54) [GCC] on linux2 Type "help", "copyright", "credits" or "license" for more information. >>> print(int(u'۲۶۷۹')) 2679 Looking at what the string is: $ python2.7 -c "print(list(u'۲۶۷۹'))" [u'\xdb', u'\xb2', u'\xdb', u'\xb6', u'\xdb', u'\xb7', u'\xdb', u'\xb9'] $ python2.7 Python 2.7.13 (default, Jan 03 2017, 17:41:54) [GCC] on linux2 Type "help", "copyright", "credits" or "license" for more information. >>> print(list(u'۲۶۷۹')) [u'\u06f2', u'\u06f6', u'\u06f7', u'\u06f9']
- evincarofautumn 9y agoDB B2, etc. are the UTF-8 encodings of U+06F2, etc. So Python is seeing mojibake: U+00DB (Û), U+00B2 (²), etc. which are not digits. Well, one of them kinda is, but it’s No (“Number, other”), not Nd (“Number, decimal digit”).
- Tistron 9y agoYeah I get that, but why is that happening?
- evincarofautumn 9y agoI’d guess because CPython is assuming the input to -c is ISO-8859-1 (Latin-1) when it decodes it using Py_DecodeLocale(): main() … setlocale(LC_ALL, "") … argv_copy[i] = Py_DecodeLocale(argv[i], NULL) … mbstowcs() or mbrtowc() … setlocale(LC_ALL, oldloc) … Py_Main(argc, argv_copy) While the REPL’s encoding (sys.stdin.encoding) is set to UTF-8 due to LANG/LC_CTYPE settings. You can get the same error when invoking the REPL as: LANG="en_US.iso8859-1" python2.7 So the shell isn’t doing anything to the text—it’s providing UTF-8 bytes in both cases, it’s just that Python is interpreting them differently.
- acdha 9y agoThat's what I thought as well when I first found it – once the initial surprise passed, it seems nice to have something just work for anyone using another script. The Unicode consortium has put an enormous amount of work into building that database and it's nice to reuse that to make something friendlier for humans.
- deleted 9y ago[deleted]