4 ms·
> the mawk flavor is extremely fast Fast, but partly because it's not Unicode-aware: it treats strings as 8-bit character sequences rather than UTF-8. Often th
by mjn 8y ago
> the mawk flavor is extremely fast
Fast, but partly because it's not Unicode-aware: it treats strings as 8-bit character sequences rather than UTF-8. Often that's fine, if non-ASCII characters are only passed through unmodified, but requires some care to avoid problems.
$ echo $LANG
en_GB.UTF-8
$ echo "ÜNICÖDE" | gawk '{print tolower($0)}'
ünicöde
$ echo "ÜNICÖDE" | mawk '{print tolower($0)}'
�nic�de
I ran into this in practice because I was using awk to convert paper titles from "Title Case" to APA-style "Only first word of title capitalized" case. The garbled output led me down a rabbit hole where I discovered that only some awks support Unicode locales, and the default awk on Debian (mawk) isn't one of them.
- davidgould 8y agoSee my post way down the bottom, I do mention this. GNU awk also has a lot of useful extensions and builtins so that sometimes it's painful to use a plain Posix awk. But, when you need the speed, it's nice to know mawk is out there. I'd love to see the mawk compilation technology merged to GNU awk. Or mawk updated with Unicode support and a few of the GNU extensions. The other item on my awk-like wishlist is CSV support, ie split $1 .. $N by CSV rules instead of just a field separator. I usually end up copying CSVs into postgresql because it is fast and then I can process it very flexibly, but it's a bit heavy for things that could be one line-ers. Also, postgresql won't load malformed CSV, but I suspect awk could be less picky. There is miller which hits some of these points, but I haven't really grokked it yet and the syntax seems ... awkward.