4 ms·
> The author will have a hard time fitting the necessary semantic actions into this code, like constructing an abstract syntax tree, It's a declared non-goal.
by wott 5y ago
> The author will have a hard time fitting the necessary semantic actions into this code, like constructing an abstract syntax tree,
It's a declared non-goal.
> The code inconsistently mixes the use of concrete character constants like '.' and '{' and #define symbols that expand to character constants like #define TOK_PLUS '+'.
> Using the character constants in the code as in expect('+') or whatever is clearer.
That's not how it works, TOK_* represent whole elements, the element may happen to be a single character, but it can be a word. It would be highly confusing to pass single characters to expect() as you wouldn't know at first sight whether it expects a simple character or an element. The TOK_* could as well be assigned a random number as you say in your following sentence. And they were until yesterday :-) I guess it is easier to debug with mnemonic characters.
https://github.com/ibara/pl0c/commit/aba92627deac3f8cfe086793b28529935378b171#diff-adb748471157c609384d66b3e3c9dfa3368d313556696a53ec2702778614f96d https://github.com/ibara/pl0c/commit/aba92627deac3f8cfe08679...
- deleted 5y ago[deleted]
- kazinator 5y agoUsing character constants for denoting one-character tokens is a time-honored tradition that everyone understands. E.g. GNU Awk: http://git.savannah.gnu.org/cgit/gawk.git/tree/awkgram.y?h=gawk-5.1.0#n1774 http://git.savannah.gnu.org/cgit/gawk.git/tree/awkgram.y?h=g... This is readable; it says "I am just what I look like: a terminal symbol corresponding to the actual +".
- p4bl0 5y ago> [building an AST is] a declared non-goal. Then what will be going on after parsing? What will be compiled?
- kazinator 5y agoCompilers can go straight to output without building an AST, which is just one form of output. I should have more generally written about output. Output can be produced procedurally in a single pass. However, it will still require refactoring to work that in. For instance, say we go to a register machine (one that has an unlimited number of registers). The expression() function needs to know the destination register of the calculation. If it returns void as before, then it has to take it as an argument. Or else, it has to come up with the register and return it. I think that if the target is a stack-based machine, that would be the easiest to work into the parsing scheme. This is because without any context at all. Let's use term() as an example: Original recognizer skeleton: static void term(void) { factor(); while (type == TOK_MULTIPLY || type == TOK_DIVIDE) { next(); factor(); } } Stack-based output: static void term(void) { factor(); // this now outputs the code which puts the factor on the stack while (type == TOK_MULTIPLY || type == TOK_DIVIDE) { next(); factor(); // of course, likewise outputs code. output_stack_operation(type); // we just add this } } What is output_stack_operation_type: static void output_stack_operation_type(int type) { switch (type) { case TOK_MULTIPLY: output_line("mul"); break; case TOK_DIVIDE: output_line("div"); break; // ... } } For instance if we see 3 * 4 / 6 we have the grammar symbols factor TOK_MULTIPLY factor TOK_DIVIDE factor the term function calls factor() first, and that function produces output, which might be the line: push 3 the term function calls factor() again, so this time the output: push 4 is produced. Then it calls output_stack_operation_type(type) where type is TOK_MULTIPLY. This outputs: mul The loop continues and further produces push 6 div Because stack code implicitly accesses and returns operands using the stack, the syntax-directed translation for it doesn't have to pass around environmental contexts like register allocation info and whatnot. The stack-based operations do not take any argument, and therefore a parser whose functions don't return anything and don't take any arguments can accommodate stack-based code generation. Stack-based code can be an intermediate representation that is parsed again, and turned into something else, like register-based. If the plan is to go to stack-based output, the code can work.