5 ms·
> Whitespace is needed between two tokens only if their concatenation could otherwise be interpreted as a different token (e.g., ab is one token, but a b is two
by css 5y ago
> Whitespace is needed between two tokens only if their concatenation could otherwise be interpreted as a different token (e.g., ab is one token, but a b is two tokens).
https://docs.python.org/3/reference/lexical_analysis.html#whitespace-between-tokens https://docs.python.org/3/reference/lexical_analysis.html#wh...
- tannhaeuser 5y agoI find that to be a really odd choice when OTOH tabs/indents are required for block structure.
- RHSeeger 5y agoPython: Semantic whitespace is the one true way Also Python: Meh, you can leave out the whitespace, I'll figure it out
- notacoward 5y agoSimilarly, Python is supposed to be a language that's easy to learn and easy to read code in. Pythonistas take pride in that, as IMO they should. With that in mind, preserving a counterintuitive behavior that exists solely because of how the lexer was implemented seems inconsistent at best. A language with a more relaxed attitude toward inscrutable brevity - e.g. Perl - can use that excuse. I don't think Python should.
- mywittyname 5y agoI first learned Python in the 2.6 days, and have been developing in it fairly regularly over the 15 or so years since. I've never encountered this issue before. The reason I've never encountered this issue, or even needed to know it was an issue is because I use good development tools. 1) Pycharm highlighted the or in orange, making it clear that it was being interpreted as a keyword, 2) I use a linter (as we all should) which explicitly highlights the lack of whitespace around the token as an issue (PEP 8: E225), and 3) I use a code formatter (which we all should) which, again, highlights this statement as requests that I fix it. Python has a lot of flaws. This one is down near the bottom of ones that are even really worth talking about.
- notacoward 5y agoStarted with 1.5 myself. I've never encountered this issue, nor would I expect to. It might be down near the bottom of the list, but it's still a wart and it showed up here for us to talk about. You seemed to think it was worth talking about at even greater length than I did. So do you just believe it's not worth other people expressing opinions?
- mywittyname 5y agoThis showed up on the front page because it's a novelty that even people who have been using the language for over a decade would never encounter. You're taking my statement way too personally and out of context. You and everyone else can express your opinions all you want. As a topic in the list of "flaws with Python," I don't think it's a very important or insightful topic because it's not something people really run into. This just doesn't fit there. This is a novelty, pure and simple. Filed under "weird programming hacks you won't expect" This would be fine. It's like, `([]+![])[+!+[]+!+[]+!+[]][([]+{})]` in javascript. Weird, kind of interesting, but it's not really a "flaw," at least not the sense of something that will burn you unexpectedly.
- CameronNemo 5y agoAny way to lint this pattern away? Say what you will about the language and lexer complexity, backwards compatibility. This is a code smell if I have ever seen one.
- tsbinz 5y agoEnforce consistent formatting, e.g. with black (https://github.com/psf/black https://github.com/psf/black) this gets formatted to [0xF or x in (1, 2, 3)].
- mywittyname 5y agoThis issue is caught by a linter and reports, "PEP 8: E225 missing whitespace around operator".
- sco1 5y agoSome cases are, but there are still plenty of patterns that (currently) are not. e.g. `1if 1else 2` See also: https://github.com/PyCQA/pycodestyle/issues/371 https://github.com/PyCQA/pycodestyle/issues/371
- baruchel 5y agoBut why "0xf or" rather than "0 xfor" (where xfor is a valid variable name) as long as we only care about splitting between valid tokens?
- css 5y agoBecause the hexadecimal parsing rule [0] happens before [1] the rule that parses names [2]: % cat 0xfor.py 0xfor 1 % python -m tokenize -e '0xfor.py' 0,0-0,0: ENCODING 'utf-8' 1,0-1,3: NUMBER '0xf' 1,3-1,5: NAME 'or' 1,6-1,7: NUMBER '1' 1,7-1,8: NEWLINE '\n' 2,0-2,0: ENDMARKER '' Also, literals have to be evaluated before names, otherwise you could overwrite them: >>> 0xzzz = 1 File "<stdin>", line 1 0xzzz = 1 ^ SyntaxError: invalid hexadecimal literal >>> 0xf = 1 File "<stdin>", line 1 0xf = 1 ^ SyntaxError: cannot assign to literal [0]: https://docs.python.org/3/library/stdtypes.html#float.fromhex https://docs.python.org/3/library/stdtypes.html#float.fromhe... [1]: https://github.com/python/cpython/blob/5ce227f3a767e6e44e7c41e0c845a83cf7777ca4/Lib/tokenize.py#L535 https://github.com/python/cpython/blob/5ce227f3a767e6e44e7c4... [2]: https://github.com/python/cpython/blob/5ce227f3a767e6e44e7c41e0c845a83cf7777ca4/Lib/tokenize.py#L589 https://github.com/python/cpython/blob/5ce227f3a767e6e44e7c4...
- tom_mellior 5y agoYou seem to suggest some sort of differences in rule priorities. I don't think they are prioritized. It's just that Python reads left to right, and the first thing it sees looks like the start of a number, so it starts parsing a number. It doesn't reason like "I have to parse like this, otherwise you could overwrite literals".
- css 5y agoI’m suggesting that it has to happen in a specific order, which it does. Regardless of the direction the tokenization process reads, the rules can’t be evaluated concurrently. > the first thing it sees looks like the start of a number Yes, because it checks if the token is a number before checking if it is a name.
- emmelaich 5y agoThis is amazing to me and I would never have suspected such a thing. Reminds me of Fortran. https://arstechnica.com/civis/viewtopic.php?t=862715 https://arstechnica.com/civis/viewtopic.php?t=862715
- PhantomGremlin 5y agoYeah, FORTRAN made some bad choices. In its defense, it first appeared in 1957. There wasn't much language design or compiler writing knowledge back in those days. When we were in college learning FORTRAN, students would ask for help from the teaching assistants at the computer center. One big problem was FORTRAN allowed horrible spaghetti code because GOTO statements could be used anywhere. It was easy to jump into and out from loops. The TAs had a tough job. It didn't help when people would deliberately mess with the TAs by asking about code with language features similar to those in the article you linked. Something like IIRC after 47 years: DO 15 I = (1, 100) purported loop stuff goes here 15 final statement of loop That is also not a loop control statement. Instead it is an assignment to a complex number.