4 ms·
Let's say you are parsing a language that contains both negative numbers, and the "minus" symbol. Most languages use the same ASCII glyph for both of those, but
by insulanus 7y ago
Let's say you are parsing a language that contains both negative numbers, and the "minus" symbol. Most languages use the same ASCII glyph for both of those, but they mean very different things, and are parsed as parts of different tokens.
So, I believe the author's point is that in cases like these, you have to remember surrounding context in order to correctly classify which type of (and which) token a character belongs to.
- johnday 7y agoIn many cases this won't be true as you can infer from the left hand side only (for example) what - means and mode switch appropriately.
- Sharlin 7y agoMany mainstream languages don't actually have negative number literals. They lex `-42` as two tokens: the operator `-` and the integer literal `42`.
- Doxin 7y agoThe unary minus operator is still different from the binary minus operator depending on context in languages that tackle it that way.
- Sharlin 7y agoYes, but that's a parsing-time distinction. The lexer does not need to care. And even though the token's syntactic meaning "depends on the context", the grammar itself can be (and usually is) perfectly context-free!
- Doxin 7y agoThat depends a lot on your parser. Personally if I write a parser I don't generally bother with having a separate tokenizing/lexing stage. Tokenizing happens as-needed and inline.
- JoeAltmaier 7y agoE.g. in an RPN calculator, one is '-' and the other is a different key 'CHS' for 'CHange Sign'
- canjobear 7y agoThat doesn’t imply formal context sensitivity. Context-free grammars can easily be locally ambiguous, as in the case of a - character.