6 ms·
What does linux utilities like sed, awk use for text manipulation because they were very slow when I was changing a few table names in a sql file.
by akarambir 8y ago
What does linux utilities like sed, awk use for text manipulation because they were very slow when I was changing a few table names in a sql file.
- zorked 8y agoI don't think they use anything in common. Try to set your locale to "C" as otherwise string comparisons will do extra work handling your locale's notions of equivalent characters.
- masklinn 8y agoNote that this and that are not necessarily related: you're talking about performing unicode-aware text matching and manipulation, TFA is solely about validating a buffer's content as UTF-8.
- akx 8y agoHow slow? On my 2013 MBP, `gsed` (sed from coreutils) can do a replacement like that at about 350 MiB/s (of which most seems to be spent writing to disk, since writing to /dev/null hikes it up to 800 MiB/s).
- akarambir 8y agoIt was sed substitute command on a ~800Mb file on Thinkpad T470 with SSD. It was taking around 40-50 sec for each substitution. Though as others have pointed, it may not be directly related to article in discussion.
- coldtea 8y ago>It was taking around 40-50 sec for each substitution. Substitution should not be really a relevant metric as it wouldn't influence the result much. Sed/Awk will still have to go through the whole file to find all occurrences they should substitute (and when they do find an occurrence, the substitution would take nanoseconds). The size of the file is a better metric (e.g. how many seconds for that 800mb in total). Also, whether you used regex in your awk/sed, and what kind. A badly written regex can slow down search very much.
- etatoby 8y agoDid you use any quadratic or worse regex algorithm? Such as having more than one .* in a single regex. Did you set LANG=C before running sed, to bypass the UTF-8 logic? Also, if you had a list of substitutions to perform, did you try writing them as a single sed script?
- coldtea 8y agoWhat was the size of the SQL file? A "few table names" doesn't mean much if the SQL file is 20GB. In any case, sed and awk are plenty fast, but not the fastest methods of text manipulation. You could write a custom C program for that.
- Thiez 8y agoWhile it sure is possible to do text manipulation in C, I don't think it should ever be the first choice, even if 'fastest' is a goal. A 0 byte is perfectly acceptable in a utf8 string (or any unicode string, really). But C has those annoying zero-terminated strings, so if you want to manipulate arbitrary unicode strings the first thing you can do is kiss the string functions in the C standard library goodbye. Which you probably want to do anyway because pascal-strings are simply better. I would use Rust or C++ for this task.
- masklinn 8y ago> Which you probably want to do anyway because pascal-strings are simply better. They're not though. While having an explicit length is great, p-strings means the length is the first item of the data buffer, which is just awful, and why Pascal was originally limited to 255 byte strings. Rust or C++ use record-strings, where the string type is a "rich" stack-allocated structure of (*buffer, length[, capacity], …) rather than just a buffer/pointer.
- Thiez 8y agoThat is a fair point, I misunderstood the term to refer to any type of string where the length is stored explicitly. I'll try and refer to them by their correct name ('record strings') from now on :-)
- Dylan16807 8y ago> p-strings means the length is the first item of the data buffer, which is just awful You can represent it as a struct of (length, char[]) which isn't awful.
- rurban 8y agoThey are still mostly not multi-byte string (i.e. unicode) aware after decades of work. I.e. you cannot really search for strings, with case-folding or normalized variants. See http://crashcourse.housegordon.org/coreutils-multibyte-support.html http://crashcourse.housegordon.org/coreutils-multibyte-suppo... and http://perl11.org/blog/foldcase.html http://perl11.org/blog/foldcase.html for an overview of the performance problems. This tool only does the minor task of validation of the UTF-8 encoding, nothing else. There are still the major tasks of decoding, folding and normalization to do.