4 ms·
No. Neither is gz. In fact, how could any compression algorithm possibly be greppable?
by vectorjohn 11y ago
No. Neither is gz. In fact, how could any compression algorithm possibly be greppable?
- labster 11y agoApparently it is possible, as zgrep can search gzipped files.
- pdkl95 11y agoThe 'xzgrep' script (and related xzdiff, xzless, xzmore scripts) are part of the standard xz package, though they are an optional feature so YMMV between distros. ~ $ xx /usr/portage/distfiles/xz-5.0.8.tar.gz ~ $ cd xz-5.0.8/ ~/xz-5.0.8 $ ./configure --help | grep -A1 scripts --disable-scripts do not install the scripts xzdiff, xzgrep, xzless, xzmore, and their symlinks
- andreasvc 11y agozgrep simply decompresses all the data and feeds it into regular grep. If the data is indexed in some way, it is possible to do better by not having to look at all the data exhaustively.
- JustSomeNobody 11y agoZgrep
- tbrownaw 11y agoProbably something like this: $ cat zgrep #!/bin/sh # zgrep FILE ARGS... FILE="$1" shift gzip -d <"$FILE" | grep "$@"
- andreasvc 11y agoIt's possible to compress and index a file at the same time, gaining both a size and speed advantage over the original. For example: https://en.wikipedia.org/wiki/FM-index https://en.wikipedia.org/wiki/FM-index
- dalke 11y agoOne example is "Fast and Flexible Word Searching on Compressed Text". A copy is at http://www.cs.uml.edu/~haim/teaching/iws/tirsaa/sources/ACM_Transactions_on_Information_Systems/word_searching_compressed_text.pdf http://www.cs.uml.edu/~haim/teaching/iws/tirsaa/sources/ACM_... . > We present a fast compression and decompression technique for natural language texts. The novelties are that (1) decompression of arbitrary portions of the text can be done very efficiently, (2) exact search for words and phrases can be done on the compressed text directly, using any known sequential pattern-matching algorithm, and (3) word-based approximate and extended search can also be done efficiently without any decoding. The compression scheme uses a semistatic word-based model and a Huffman code where the coding alphabet is byte-oriented rather than bit-oriented.