5 ms·
Developer of the tool here :) Glad to see it posted here, I still actively use it myself. Also check out the fzf integration in the README: https://github.com/p
by phiresky 6y ago
Developer of the tool here :) Glad to see it posted here, I still actively use it myself. Also check out the fzf integration in the README: https://github.com/phiresky/ripgrep-all/blob/master/doc/rga-fzf.gif https://github.com/phiresky/ripgrep-all/blob/master/doc/rga-...
Currently the main branch is undergoing a refactor to add support for having custom extractors (calling out to other tools), and more flexible chains of extractors.
Ripgrep itself has functionality integrated to call custom extractors with the `--pre` flag, but by adding it here we can retain the benefits of the rga wrapper (more accurate file type matchers, caching, recursion into archives, adapter chaining, no slow shell scripts in between, etc).
Sadly, during rewriting it to allow this, I kind of got hung up and couldn't manage to figure out how to cleanly design that in Rust. I'd be really glad if a Rust expert could help me out here:
In the currently stable version, the main interface of each "adapter" is `fn(Read, Write) -> ()`. To allow custom adapter chaining I have to change it to be `fn(Read) -> Read` where each chained adapter wraps the read stream and converts it while reading. But then I get issues with how to handle threading etc, as well as a random deadlock that I haven't figured out how to solve so far :/
- cb321 6y agoOne possibility is the almost dirt-simple solution wherein you just have a "make"/"Makefile" (or your favorite other build system) maintain a shadow tree of parallel pre-translated files. You get parallelism via `make -j$(nproc)` or its equivalent. Every name in the shadow is built from the name in the origin but maybe with ".txt" added (or .txt.gz if you want to keep the compressed with whatever is the fastest decompressor builtin to ripgrep as a library not called as a program). Untranslated names can be just symbolic/hard links back to the origin. Build rules become as flexible as your build system. This also scales to deployments that have more disk space than memory. Admittedly, in that case, the whole procedure probably becomes disk-IO bound, but maybe not. Maybe some translations cannot even keep up with disk IO - NVMe storage is pretty fast, for example. Or available memory may vary dynamically a lot, sometimes allowing the shadow to be fully in the buffer cache, other times not. It strikes me as less presumptuous to assume you can find disk space vs. having that much memory available. (EDIT2: though I may be confused about how `rga` operates - your doc says "memory cache", though.) On the pro-side, but for updating the shadows based on origins, the user could even just `rg` from within the shadow and translate filenames "in their head", although stripping an always present string is obviously trivial. Indeed, you won't need `rg --pre` at all and the grep itself could become pluggable. I doubt any of your other `fzf`/etc. integrations would be made more complicated by this design, either. This all strikes me as simple/nice enough that someone has probably already done it...EDIT1: Oh, I see from thumbs ups and other comments over at [1] and [2] that @phiresky is probably already aware of this design idea, but maybe some HN person knows of an existing solution along these lines. [1] https://github.com/BurntSushi/ripgrep/issues/978 https://github.com/BurntSushi/ripgrep/issues/978 [2] https://github.com/BurntSushi/ripgrep/pull/981 https://github.com/BurntSushi/ripgrep/pull/981
- burntsushi 6y ago> In the currently stable version, the main interface of each "adapter" is `fn(Read, Write) -> ()`. To allow custom adapter chaining I have to change it to be `fn(Read) -> Read` where each chained adapter wraps the read stream and converts it while reading. But then I get issues with how to handle threading etc, as well as a random deadlock that I haven't figured out how to solve so far :/ I don't quite grok the problem here. If you file an issue against ripgrep proper with code links and some more details, I can try to assist. Taken literally, ripgrep uses that exact same approach. There are potentially multiple adapters being used. Each adapter is just defined to wrap a `std::io::Read` implementation, and the adapter in turn implements `std::io::Read` so that it can be composed with others. The part that I'm missing is why this has anything to do with threading or deadlocks. I/O adapters shouldn't be having anything to do with synchronization. So I'm probably misunderstanding your problem.
- phiresky 6y ago> If you file an issue against ripgrep proper with code links and some more details Sorry, I don't think I explained my issue very well. In general it has nothing to do with the interaction with ripgrep, that works fine. It's that each adapter (e.g. zip -> list of file streams) needs to have an interface of fn(Read) -> Iter<ReadWithMeta> But then if there's a PDF within the zip, I have to give the returned ReadWithMeta to the PDF adapter - but it can't take ownership, because the Archive file iterators only give borrowed reads. I maybe worked around this by creating a wrapper type [3] and adding an unsafe here [2], but something deadlocks when adapting zip files currently. Also, for external programs, I have to copy the data from the Read into a Write (stdin of the program) - which needs to happen in a separate thread, otherwise the stdout is never read [1], but some Reads I have aren't Send since they come from e.g. zip-rs, so they can't be passed to a thread. [1] https://github.com/phiresky/ripgrep-all/blob/baca166fdab3d2466ac9c482c3c9b8e160014b9a/src/adapters/spawning.rs#L118 https://github.com/phiresky/ripgrep-all/blob/baca166fdab3d24... [2] https://github.com/phiresky/ripgrep-all/blob/baca166fdab3d2466ac9c482c3c9b8e160014b9a/src/recurse.rs#L23 https://github.com/phiresky/ripgrep-all/blob/baca166fdab3d24... [3] https://github.com/phiresky/ripgrep-all/blob/baca166fdab3d2466ac9c482c3c9b8e160014b9a/src/adapters/zip.rs#L52 https://github.com/phiresky/ripgrep-all/blob/baca166fdab3d24...
- maximz 6y agoLove this. I appreciate your building on ripgrep versus my own bulky lucene-based approach a while back (https://github.com/maximz/sift https://github.com/maximz/sift), and that you don’t require pre-indexing but build up a cache as you go.
- one-punch 6y agoThe integration with fzf seems nice. Any plans to integrate with skim, a Rust implementation of fzf? https://github.com/lotabout/skim https://github.com/lotabout/skim
- cutemonster 6y agoI also hope for that :-)
- pdimitar 6y agoSeconding this, I love fzf but I love skim a little more.
- aembleton 6y agoAUR has both a ripgrep-all [1] and ripgrep-all-bin [2] package. Both were addded by you. The bin package has a newer version. What is the difference between the two? 1. https://aur.archlinux.org/packages/ripgrep-all/ https://aur.archlinux.org/packages/ripgrep-all/ 2. https://aur.archlinux.org/packages/ripgrep-all-bin https://aur.archlinux.org/packages/ripgrep-all-bin
- jhardy54 6y agoYou can (and should!) read the PKGBUILDs, they're very small and should always be manually inspected before install. The -bin suffix is an AUR convention to let you know that it's downloading a precompiled binary rather than building from source.
- scaladev 6y agoIf we're talking about conventions, it's also an AUR convention to use -vcs suffix for source builds (like -git), for example: https://aur.archlinux.org/packages/?O=0&SeB=nd&K=-git&outdated=&SB=v&SO=d&PP=50&do_Search=Go https://aur.archlinux.org/packages/?O=0&SeB=nd&K=-git&outdat... https://aur.archlinux.org/packages/?O=0&SeB=nd&K=-hg&outdated=&SB=v&SO=d&PP=50&do_Search=Go https://aur.archlinux.org/packages/?O=0&SeB=nd&K=-hg&outdate...
- trynewideas 6y agoThanks for this tool, I'm already getting a ton of use from it. For fun, I pointed a 12-core/32GB RAM 2018 MBP at a 9GB network share full of PDFs, while still using the laptop for other things (so not a benchmark, just an anecdote). Initial cold/uncached run: rga -j 12 testword share 1140.70s user 77.58s system 31% cpu 1:03:55.85 total Cached: rga -j 12 testword share 8.09s user 4.88s system 92% cpu 14.048 total Cache after the run is 77M.