5 ms·
the lexer can always tell with fixed lookahead whether `a-b` is three tokens or one How does that work? It doesn't seem obvious to me? I also wonder what appr
by nextInt 8y ago
the lexer can always tell with fixed lookahead whether `a-b` is three tokens or one
How does that work? It doesn't seem obvious to me?
I also wonder what approaches other languages have taken since I haven't seen dashed keywords in any of the other big languages.
- gizmo686 8y ago>How does that work? It doesn't seem obvious to me? Assume that "-" is the MINUS token. This means that, starting with the first character in "a", we have the token string "A MINUS B", which we can determine with a 3 token lookahead. Now, if either A or B are reserved keywords, the lexer knows that "A MINUS B" is not correct (unless the grammar happens to allow for a keyword to occur next to a minus sign, which I don't think Java does. If it does, then you just avoid using that keyword when deriving a new one). At this point, the lexer and lex it as "A-B", which is (hopefully) a keyword. I can't think of any examples, but I feel like I have seen languages reserve take a simmiliar approach where they reserve a prefix, which allows them to create as many new reserved words as they like.
- ridiculous_fish 8y agoOoh let's take a tour of big languages. - C has given up on new identifiers save the "reserved namespace", consisting of an underscore followed by an Uppercase, which is how you get _Bool and _Generic. Oof, nuff said. - C++ has stretched poor `static` to its limits, and now incorporates context-sensitive "identifiers with special meaning": `final` and `override`. No new hard keywords in C++17 AFAIK. - C# takes this even further with its LINQ-driven contextual keywords. - JavaScript stumbles about by adding new identifiers, BUT only in strict mode, and even then only sometimes; `yield` is especially ambiguous. Fill in the blanks please!
- ygra 8y agoI think C# has only ever added contextual keywords after its initial release and if I remember correctly there was only one breaking change for semantics of source code, which I find impressive. I also didn't get the impression (from working with Roslyn and looking at its source here and there, as well as reading posts from people like Eric Lippert or Mads Torgersen about the language design) that contextual keywords are as much of a hassle as Brian makes them out to be. About the only really annoying one I can think of in C# would be the nameof operator, which has to be parsed as a method invocation and is only the nameof operator when there's no symbol with that unqualified name accessible at that point that can be invoked as a method (e.g. an actual method named nameof (rare), or a local variable of a delegate type). Pretty much all other contextual keywords are only valid in a few places where parsing is not ambiguous. You just happen to have a token there that's still valid elsewhere as an identifier. It's interesting to read about the musings and decision making processes for different languages, though. I'm sure Java and C# are both designed very carefully, yet with radically different goals and outcomes. And I'm sure both language design teams must ponder pretty much the same issues.
- haldean 8y agoC is a little better than presented here because they add a header to the standard library with a #define in it that gives the new identifier a reasonable keyword; for example, you can access `_Bool` as `bool` if you `#include <stdbool.h>`. I actually think this is a reasonable compromise; you can opt in to the new "keyword" at the compilation-unit-level.