3 ms·
I wouldn't assume. The awk implementation likely stands up well for a 1-billion-row challenge, with its thoughtful bytecode-based design. Redditors ran some q
by jgarzik 2y ago
I wouldn't assume. The awk implementation likely stands up well for a 1-billion-row challenge, with its thoughtful bytecode-based design.
Redditors ran some quick performance tests on parsing, also: https://www.reddit.com/r/rust/comments/1fd7qgl/comment/lmelo6z/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button https://www.reddit.com/r/rust/comments/1fd7qgl/comment/lmelo...
- dundarious 2y agoParsing the awk input? While not an irrelevant concern, that is obviously on the low priority end when considering awk performance in general. My most often used awk program is just '{print $1}', but I use it on enormous files. The performance when operating on the enormous file is the concern wrt performance, not the initial parse of '{print $1}' or of command line arguments. I know you're just directly responding to the concerns of a parent comment though.
- jgarzik 2y agoIt is hoped that posixutil's awk's bytecode-based modern design should keep performance high, theoretically higher than ancient C-based awks. Inspired by Ray Gardner's "wak" awk implementation https://github.com/raygard/wak https://github.com/raygard/wak A volunteer benchmarking our awk on a 1-billion line text file would be welcome.
- dundarious 2y agoYes, that's the most important part to focus on. I have nothing to remark about it, I haven't seen/gathered any numbers on it.