3 ms·
There's a fine line to walk between "Not Invented Here" and suffering through an inadequate tool. We at Facebook have a long history of trying stock tools and s
by andralex 13y ago
There's a fine line to walk between "Not Invented Here" and suffering through an inadequate tool. We at Facebook have a long history of trying stock tools and see them failing at our scale.
Walter Bright puts out a good argument against lexer generators here: http://www.drdobbs.com/architecture-and-design/so-you-want-to-write-your-own-language/240165488 http://www.drdobbs.com/architecture-and-design/so-you-want-t.... Besides, tokenization is hardly the interesting/bulky part of the task. The verifications are.
- quotemstr 13y ago> Walter Bright puts out a good argument against lexer generators here: http://www.drdobbs.com/architecture-and-design/so-you-want-t... http://www.drdobbs.com/architecture-and-design/so-you-want-t.... I'm going to have to disagree with you and Walter. While I agree that for a fixed language, a well-written hand-built recursive descent parser will beat the pants off the output of a parser generator, languages aren't necessarily fixed and writing a good hand-build parser is harder than it seems. I still recommend that people start with parser generators. (Edit: yes, we're talking about lexers above, but the article mentions both lexers and parsers. I think the case for using a hand-built lexer is much weaker than the case for using a hand-built parser, not that I'm a fan of either.) If you use a parser generator, you can be sure that your language has a clear, (hopefully) LALR(1) syntax that anyone can understand and port to another parsing environment. If you hand-parse, then the temptation exists not only to introduce subtle irregularities and context dependencies into the language, but also to mix semantic interpretation and parse tree generation, and both of these anti-patterns make it difficult to build generic tools that work with your new language. With a hand parser, you also run the risk of making silly coding errors that end up being baked into your language specification. (Ahem, PHP.)
- haberman 13y ago> While I agree that for a fixed language, a well-written hand-built recursive descent parser will beat the pants off the output of a parser generator I'm not sure this is true, do you have numbers that demonstrate this? I know that Clang knows some tricks that lexer generators don't currently know like using SSE to skip long blocks of comments. But even that, to me, represents an opportunity for better parsing tools, not an inherent performance advantage of hand-written parsers/lexers.
- quotemstr 13y ago> I'm not sure this is true, do you have numbers that demonstrate this? No numbers offhand, but GCC's experience of moving toward a hand-written C++ parser was very positive. I don't believe hand-written parsers have an inherent speed advantage, but it does seems that it'd always be easier for a hand-written parser to produce more informative error messages --- if anything, because the hand-written parser has more embedded syntactic information.
- haberman 13y ago> GCC's experience of moving toward a hand-written C++ parser was very positive. Fom what I see here, the measured improvement was 1.5% -- http://gcc.gnu.org/wiki/New_C_Parser http://gcc.gnu.org/wiki/New_C_Parser > I don't believe hand-written parsers have an inherent speed advantage, but it does seems that it'd always be easier for a hand-written parser to produce more informative error messages From what I can see, this is the major shortcoming of parser tools at the moment. I haven't looked into this problem too deeply yet, but at the moment I suspect that it isn't possible to get good enough error messages automatically. But to me this doesn't mean "parsing tools will never have good error messages", it means "parsing tools need to evolve to give users the right hooks to produce good error messages themselves." This would let the algorithms do the part that they do well (analyzing the grammar and creating state machines) while letting the human do the part that can't be done automatically (creating good error messages).
- quotemstr 13y agoSure, things might change. I'd be happy with better error messages from automatically-built parsers. There's a lot of potential, I think, in methods that take advantage of having a formal description of the language grammar --- you could, for example, automatically generate a sentence that "fits in" at the point at which an error occurs. There's still a lot of progress to be made in the area, though, and for now, hand-generated error messages do seem to be more useful to people.
- sanxiyn 13y agoGetting good error messages from the parser generator is a solved problem, but it seems that solution is not widely known. As you said, it isn't possible to get good error messages automatically. The solution is to include error cases in the grammar. No hooking is necessary. If people often miss semicolons, include missing semicolon case in the grammar! Mono C# compiler uses this technique (including error cases in the grammar) to the great effect. I wonder why people don't do that and resort to hand-written parsers instead.
- andralex 13y agoI think this is starting to meander. We're talking about lexers, not parsers, and for a stable language, not a language being designed.
- cbsmith 13y agoC++14 just got finalized... not sure how stable the language really is...
- jejones3141 13y agoIf memory serves, the original preprocessor that turned "C with Classes" into C wasn't written with a parser generator. I think your argument applies.
- haberman 13y ago> We at Facebook have a long history of trying stock tools and see them failing at our scale. I can understand that, having experienced the same before. I just re-read my message and I regret the way it comes off -- my intention in writing it wasn't to be critical of you, but to explore the issue of how inconvenient parser tools are perceived as being, and try to analyze why that is. "Insanity" was meant to refer to the overall situation, and not you personally! I might re-word this a bit. I still think that we ought to be in a position where the tool is so easy and so valuable that it is a no-brainer. And I think this largely indicates that there is a long way for the tools to improve.
- tlrobinson 13y agoThe Walter Bright who's comment appears next to yours? https://news.ycombinator.com/item?id=7294024 https://news.ycombinator.com/item?id=7294024
- WalterBright 13y agoThere can be Only One.