3 ms·
It works in Python 2, you just have to ensure that your string is unicode. I can't get it to work directly from the command line for some reason, but if you ope
by plus 9y ago
It works in Python 2, you just have to ensure that your string is unicode. I can't get it to work directly from the command line for some reason, but if you open an interactive Python session it works:
Python 2.7.12 (default, Dec 14 2016, 13:32:53)
[GCC 4.9.3] on linux2
Type "help", "copyright", "credits" or "license" for more information.
>>> print(int(u'۲۶۷۹'))
2679
- Tistron 9y agoHmm, interesting. What's up with the terminal here? The shell is doing something to the text. $ python2.7 -c "print(int(u'۲۶۷۹'))" Traceback (most recent call last): File "<string>", line 1, in <module> ValueError: invalid literal for int() with base 10: '\xdb\xb2\xdb\xb6\xdb\xb7\xdb\xb9' $ python2.7 Python 2.7.13 (default, Jan 03 2017, 17:41:54) [GCC] on linux2 Type "help", "copyright", "credits" or "license" for more information. >>> print(int(u'۲۶۷۹')) 2679 Looking at what the string is: $ python2.7 -c "print(list(u'۲۶۷۹'))" [u'\xdb', u'\xb2', u'\xdb', u'\xb6', u'\xdb', u'\xb7', u'\xdb', u'\xb9'] $ python2.7 Python 2.7.13 (default, Jan 03 2017, 17:41:54) [GCC] on linux2 Type "help", "copyright", "credits" or "license" for more information. >>> print(list(u'۲۶۷۹')) [u'\u06f2', u'\u06f6', u'\u06f7', u'\u06f9']
- evincarofautumn 9y agoDB B2, etc. are the UTF-8 encodings of U+06F2, etc. So Python is seeing mojibake: U+00DB (Û), U+00B2 (²), etc. which are not digits. Well, one of them kinda is, but it’s No (“Number, other”), not Nd (“Number, decimal digit”).
- Tistron 9y agoYeah I get that, but why is that happening?
- evincarofautumn 9y agoI’d guess because CPython is assuming the input to -c is ISO-8859-1 (Latin-1) when it decodes it using Py_DecodeLocale(): main() … setlocale(LC_ALL, "") … argv_copy[i] = Py_DecodeLocale(argv[i], NULL) … mbstowcs() or mbrtowc() … setlocale(LC_ALL, oldloc) … Py_Main(argc, argv_copy) While the REPL’s encoding (sys.stdin.encoding) is set to UTF-8 due to LANG/LC_CTYPE settings. You can get the same error when invoking the REPL as: LANG="en_US.iso8859-1" python2.7 So the shell isn’t doing anything to the text—it’s providing UTF-8 bytes in both cases, it’s just that Python is interpreting them differently.