4 ms·
It's not actually implemented as a scannerless parser, but the grammar API mostly abstracts away the parser/lexer distinction. We have to handle JSX and all la
by maxbrunsfeld 8y ago
It's not actually implemented as a scannerless parser, but the grammar API mostly abstracts away the parser/lexer distinction.
We have to handle JSX and all language versions combined because for our use case, users need to be able to open up `.js` files and have them Just Work.
It's a similar story with Python and Ruby - we have to handle the union of all language versions.
- mncharity 8y ago> open up `.js` files and have them Just Work Nod. A single big grammar just wasn't what I'd thoughtlessly expected. When I did something similar, there was a gaggle of small grammars slicing up javascript aspects and history, then variously composed into assorted language grammar variants. But part of the motivation for that was sharing grammar fragments across languages, and "parse exactly language version X" in support of parser validation and compiler use. So a bit different use cases.
- mncharity 8y ago> It's not actually implemented as a scannerless parser, but the grammar API mostly abstracts away the parser/lexer distinction. Any aspects of the 'not' part of 'mostly' which one should bear in mind?
- maxbrunsfeld 8y agoYeah, the grammar author does fully control the division between the parser and lexer: every literal (string or regex) in the grammar corresponds to a token. There's also a `token()` function that you can use to specify that an arbitrary rule should be handled by the lexer as a single token. In most cases, you don't have to think about it; the obvious way to write something is the right way. There are cases where its helpful to have a mental model of how lexing works.
- mncharity 8y ago> grammar author does fully control the division between the parser and lexer Nifty. So explicit GLR forks at rule level (`conflicts:`), non-forking token "conflict" resolution[1], no token regex backtracking pressure across tokens, and token resolution at a single code position can(?) differ across conflicting rules? I'm uncertain on that last bit. > In most cases, you don't have to think about it Well yes, but, some of us crave much more syntactically flexible languages. :) [1] https://tree-sitter.github.io/tree-sitter/creating-parsers#conflicting-tokens https://tree-sitter.github.io/tree-sitter/creating-parsers#c...