3 ms·
I have been wanting this for a long time, although my primary use case would be indexing a particular field. It is nice that you were able to make it work with
by xaa 11y ago
I have been wanting this for a long time, although my primary use case would be indexing a particular field.
It is nice that you were able to make it work with ordinary gzip files. In bioinformatics, an approach to this problem has been to create a gzip-backwards-compatible (i.e., can be read by zcat but not written to by gzip) format called BGZF [1] which acts like ordinary gzip except each compressed block is the same size, which eases indexing and partial decompression.
I would be interested to know how you solved the problem of finding the right block and offset when each compressed block can be an arbitrary size as it is in ordinary gzip format.
[1] http://blastedbio.blogspot.com/2011/11/bgzf-blocked-bigger-better-gzip.html http://blastedbio.blogspot.com/2011/11/bgzf-blocked-bigger-b...
- mattgodbolt 11y agoFantastic! I had hoped others would have the same problem as me. I didn't have the luxury of being able to change my many input gzip files (unlike BGZF). The trick is to store the gzip internal state as a checkpoint every few hundred MB. Then to decode data at a particular offset, you find the further internal state that's before your offset; reinitialize the compression internal state with that, and then scan on until you get to the offset you need. The zindex file that's generated contains both these checkpoints and the index that lets one scan there. There's an open issue I'd like to resolve to add support for indexing a particular field: would support for arguments like those the "cut" UNIX utility allow for your use case?
- xaa 11y agoSmart. I will see for myself in a minute how much space such an approach takes; I would imagine the compressor size gets nontrivially large on big files. Well, my exact use case, which I imagine I'm not alone in, is indexing a tab-delimited file with a single-line header containing the column names. So indexing by column index would be adequate, but it would present two very minor problems: 1) I'd have to figure out a way to exclude the header row from output (or alternately, to always output it), and 2) it would be slightly more convenient to refer to the columns to be indexed by name rather than position, as some of these have thousands of columns. Can you have multiple indexes on the same GZIP file? And if so, does it create multiple .zindex files or re-use the same one? I don't know if SQLite supports custom data retrieval backends like Postgres, but what would be REALLY interesting is to write one such that the indexing information is in SQLite, and you can use SQL to query row data directly from the GZIP index (I know this is beyond the scope of your project, I'm just thinking out loud now).
- mattgodbolt 11y agoShort reply as I'm on a phone, but thanks! I can add an option to pick the nth field, and to skip the first X lines of the input; I think this'd cover your use case! The code and file format support multiple indices in the same zindex file but I haven't got a decent way to pass command line arguments to define them yet! Watch this space: also feel free to ping me on email for more comments!
- mattgodbolt 11y agoAs of just now, zindex and zq now support field-based indexing. Hope this helps your use case!