3 ms·
> Or anything but just hand-hacking your own tokenizer. I've only just got into parsing and I know for sure that there are more challenging problems, but for a
by DavidMcLaughlin 17y ago
> Or anything but just hand-hacking your own tokenizer.
I've only just got into parsing and I know for sure that there are more challenging problems, but for a dynamic language like Javascript there is a pretty major example of where this worked. JS Lint: http://www.jslint.com/fulljslint.js http://www.jslint.com/fulljslint.js
1) Split source into lines
2) Run each line through a regex:
/^\s([(){}\[.,:;'"~\?\]#@]|==?=?|\/(\(jslint|members?|global)?|=|\/)?|\[\/=]?|\+[+=]?|-[\-=]?|%=?|&[&=]?|\|[|=]?|>>?>?=?|<([\/=!]|\!(\[|--)?|<=?)?|\^=?|\!=?=?|[a-zA-Z_$][a-zA-Z0-9_$]|[0-9]+([xX][0-9a-fA-F]+|\.[0-9]*)?([eE][+\-]?[0-9]+)?)/
3) Look at the resulting token character by character and define what kind of token it is.
Once it has these parsed tokens JS Lint then uses a Pratt Parser to check syntax and style, but you could easily use the same tokenizer and build a parse tree as part of a translator or compiler instead.