6 ms·
What are the reasons for grep not being replaced/improved? This topic seems a bit old by now.
by mariopt 3y ago
What are the reasons for grep not being replaced/improved?
This topic seems a bit old by now.
- capableweb 3y agoThere are multiple alternatives you can already use as an alternative, like ripgrep. What are you proposing, switching out the command `grep` for another utility? Sounds like that could introduce a ton of breakage, for little value. People who want a faster grep will use a different thing, while people who use grep can continue to use it. Sounds like an ideal situation already.
- anonymous_sorry 3y agoThese benchmarking results are seven years old, so perhaps it has been. My entirely anecdotal and unscientific impression is that rg and grep perform similarly on Linux (though rg has nicer defaults for searching through source code). The old version of grep that Apple preinstalls on the Mac was slower last time I checked though.
- burntsushi 3y agoThere's like a whole host of things you could use to explain it. Inertia. Compatibility. Resistance to change. Innovator's dilemma. And so on. (I do not say any of these things pejoratively! All of those things apply to me too.) With respect to compatibility, see my FAQ on the topic: https://github.com/BurntSushi/ripgrep/blob/master/FAQ.md#posix4ever https://github.com/BurntSushi/ripgrep/blob/master/FAQ.md#pos...
- alkonaut 3y agoSomeone designed unix based on the idea that some system functions are both core OS functions AND tools for human use. That leads to some bizarre outcomes decades later like "there must be a program called xyz that accepts these arguments and works exactly like this".
- wruza 3y agoFor the same reason the 40yo chair I currently sit in is not being replaced with Razer UltraSeat XR3000-A. It's comfortable, fits the workplace around it, and there's no reason for getting a replacement and rebuilding everything. (Partially because a Razer-like chair already stands nearby taking care of my clothes, but that's where the analogy ends.)
- gosub100 3y agoComplete guess: it works just fine for 99.9999% of users, but greater than (1-0.999999)% chance that it would break compatibility or have a bug. Anyone who would need the performance gain would know about specialized alternatives.
- Tarucho 3y agoI guess it is because after decades of use, grep has probably been fixed to handle lots of user cases that the new tools don´t handle because they haven´t found them yet.
- burntsushi 3y agoAuthor of ripgrep here. Like automatic encoding detection and transparently searching UTF-16? Or simple ways for composing character classes, e.g., `[\pL&&\p{Greek}]` for all codepoints in the Greek script that are letters. Another favorite of mine is `\P{ascii}`, which will search for any codepoint that isn't in the ASCII subset. Or more sophisticated filtering features that let you automatically respect things like gitignore rules. Those are all things that ripgrep does that grep does not. So I do not favor this explanation personally. ripgrep has just about all of the functionality that GNU grep does. I would say the two biggest missing pieces at this point are: * POSIX locale support. (But this might be a feature[1].) * Support for "basic" regexes or some equivalent that flips the escaping rules around. i.e., You need to write `\+` to match 1 or more things, where as `+` will just match `+ literally. Otherwise, ripgrep has unfortunately grown just about as many flags as GNU grep. [1]: https://github.com/mpv-player/mpv/commit/1e70e82baa9193f6f027338b0fab0f5078971fbe https://github.com/mpv-player/mpv/commit/1e70e82baa9193f6f02...
- deleted 3y ago[deleted]
- amiga386 3y agogrep is a general purpose tool for searching for text in all types of files, baked into the standards for UNIX. Some programmers use it to search source code. Other people use it for other types of text searches that have nothing to do with source code, they rely on it in scripts, they don't use it as part of a text-based programmer UI, they rely on it to never crash, etc. ripgrep is a specialist, opinionated tool, designed primarily to search through source code repositories. There's not much you can add to general purpose text search to make it faster; you can make it use mmap() at the risk of it crashing on truncated files, you can reduce the expressiveness of regular expressions so they can be computed faster. You could throw out general support for all locales and charsets and hardcode support for only UTF-8 / UTF-16, but you shouldn't.
- bawolff 3y ago> you can make it use mmap() at the risk of it crashing on truncated files I was under the impression that grep removed mmap() support because it was slower than normal file i/o
- burntsushi 3y agoI talked about this in the OP. Memory maps are sometimes a little faster. See: $ ls -l full.txt -rw-rw-r-- 1 andrew users 13113340782 Sep 29 12:30 full.txt $ time rg -c --no-mmap Clipton full.txt 294 real 1.337 user 0.470 sys 0.866 maxmem 15 MB faults 0 $ time rg -c --mmap Clipton full.txt 294 real 1.045 user 0.722 sys 0.323 maxmem 12511 MB faults 0 But in recursive search, especially when used for lots of little files, they end up provoking substantial overhead that slows everything down. And this might change depending on the platform.
- burntsushi 3y ago> There's not much you can add to general purpose text search to make it faster Oh I beg to differ! The blog post goes into this. Here's a simple demonstration using ripgrep 14: $ ls -l full.txt -rw-rw-r-- 1 andrew users 13113340782 Sep 29 12:30 full.txt $ time rg -c --no-mmap 'Clipton' full.txt 294 real 1.419 user 0.539 sys 0.879 maxmem 15 MB faults 0 $ time LC_ALL=C grep -c 'Clipton' full.txt 294 real 6.911 user 6.078 sys 0.829 maxmem 15 MB faults 0 $ time rg -c --no-mmap 'DMZ|Clipton' full.txt 1070 real 1.643 user 0.747 sys 0.894 maxmem 15 MB faults 0 $ time LC_ALL=C grep -E -c 'DMZ|Clipton' full.txt 1070 real 8.317 user 7.384 sys 0.930 maxmem 15 MB faults 0 No memory maps. No multi-threading. No filtering. No fancy regex engine features or reducing expressiveness. No locales. No UTF-8. No UTF-16. Just a simple literal and a simple alternation of literals. It's just better algorithms. Also, you can disable ripgrep's opinions with `-uuu`. It's not designed to just be for code searching. You can use it for normal grepping too. It will even automatically revert to the standard grep line format in shell pipelines.
- p0nce 3y agoLack of interest maybe? I don't use grep, I use an UI that lets me click and jump to a file. Or the builtin search in my IDE.
- ilc 3y agoHonestly: I rely on my shell scripts working and any "grep replacement" has to work with all the old crusty shell scripts out there, likely including ones that use odd "quirks" and GNU options. If you want to innovate in this space, why sign up for all that? Invent a better wheel, and if people like it, they'll migrate over time. I remember using ag in the old days, and I use rg now. But there's things rg does by default that I don't like at times... so I go back to old fashioned grep. rg is at the point where many programmers use it. I think it is on its way to becoming one of those "standard tools". It needs... another 5 years? When POSIX has a rg standard... we'll know ripgrep "succeeded" and teargrep will soon come into existence ;)
- burntsushi 3y agoOh my heavens, I would never let ripgrep into POSIX. You can pry it out of my cold dead hands. :-) > so I go back to old fashioned grep If you do `rg -uuu` then it should search the same stuff grep will. Not sure if that's what you meant though.