4 ms·
Let's say I have a web visits analytics log database, with a column "user-agent" and a column "url" that have most of the time the same values (99.9% of rows ha
by josephernest 4y ago
Let's say I have a web visits analytics log database, with a column "user-agent" and a column "url" that have most of the time the same values (99.9% of rows have one of 100 typical values).
Would this work well with compression here?
ie. from 50 bytes on average for these columns... to a few bytes only?
- hnarn 4y agoProbably yes, but presumably you could achieve this by compressing the data on the file system level as well (for example with ZFS).
- Scaevolus 4y agoYes, if you use the "dictionary" option it will compute build a buffer of common strings and use that during compression, so your user-agent columns will largely become references to the dictionary. Make sure you set compact to true! See this for more information: https://facebook.github.io/zstd/#small-data https://facebook.github.io/zstd/#small-data
- josephernest 4y agoWill this mode require a separate file along the data.db Sqlite file, to store the compression dictionary? From your last link: > The result of this training is stored in a file called "dictionary", which must be loaded before compression and decompression. Also, what happens if a column has been compressed based on 100 typical values, and then later we insert thousands of new rows with 500 new frequent values. Does the compression dictionary get automatically updated? Then do old rows need to be rewritten? PS what is compact mode?
- Scaevolus 4y agoIt stores these compression dictionaries in the database: > If dictionary is an int i, it is functionally equivalent to zstd_decompress(data, (select dict from _zstd_dict where id = i)). Compact saves you 8 bytes per compressed column by removing some metadata zstd uses : > if compact is true, the output will be without magic header, without checksums, and without dictids. This will save 4 bytes when not using dictionaries and 8 bytes when using dictionaries. this also means the data will not be decodeable as a normal zstd archive with the standard tools. The same compact argument must also be passed to the decompress function. You will need to update the dicts occasionally, but there's a helper function to do this for you, read the docs about zstd_incremental_maintenance.