6 ms·
Shoco: a fast compressor for short strings
- Khao 11y agoI get negative compression percentage when I put words with "é" in the test box.
- jozan 11y agoIn default it doesn't work well with non-ASCII characters. https://ed-von-schleck.github.io/shoco/#how-it-works https://ed-von-schleck.github.io/shoco/#how-it-works
- Semiapies 11y agoBetween this (an ASCII-only compressor in 2015?) and the other aspects brought up here, it seems downright toylike.
- techwizrd 11y agoI wonder what'd happen if you used this on base64 strings.
- bmh100 11y agoI would love to see a blog post about that test, if you're willing.
- thrownaway2424 11y agoI can't tell you how many times I've said to myself "if only these very short ASCII strings were even shorter!"
- BrandonSmith 11y agoAt scale, and if you are paying for transmission costs, it can have a massive impact.
- rurban 11y agoWill test against smaz for our internal JSON compressed protocol. smaz compressed fine but was too slow. The ability to train the model sounds convincing.
- knodi123 11y agoLook how well it can compress "fofofofofofofofofofofo". 50% Look how well it can compress "ababababababababababab". 0%
- dalke 11y agoI have a background project of exploring how to compress SMILES strings, which is a notation for storing chemical information. For example, "C" is methane, "CC" is ethane, "C=C" is ethene, "CCO" is ethyl alcohol, "C1CCCCC1" is cyclohexane, and "c1ccccc1", which contains aromatic carbons, is benzene. The average length of a SMILES string for real-world molecules is about 50 characters. I previously evaluated a special purpose tool which identifies the best n-grams and uses dynamic programming during encoding. That gets about 70% compression on SMILES string. I also tried the off-the-shelf femtozip which got about 60% compression but had more decompression overhead than I like. Shoco, trained on 1,455,763 SMILES strings (average of 56 letters each), and tested with 100,000 strings from the training set, reports "average compression ratio: 47%".
- bmh100 11y agoCould you provide more information about your SMILES test? How many unique symbols were there? How does gzip do? This is an interesting use case.
- dalke 11y agoSure. I'm switching this conversation to email though, using the gmail account in your profile. Short version is, I trained it on the RDKit-generated SMILES strings from ChEMBL-20. Three of the strings look like this: CC(C)=CCC/C(C)=C/C=C/C(=O)N1CCCC1 CC(=O)NC(C(=O)N1CCSCC1)[C@H]1CC(C(=O)O)C[C@@H]1N=C(N)N O=C(CC(c1ccc(F)cc1)(c1ccc(F)cc1)c1ccc(F)cc1)N1C[C@H](O)C[C@H]1C(=O)N1CCC[C@@H]1C(=O)NC[C@@H]1CCCNC1 On the raw data set (on record per line), wc reports: 1455763 1455763 82882385 while | gzip -c | wc -c reports 18773892.
- TheLoneWolfling 11y ago> I'm switching this conversation to email though I wish you wouldn't do that. That defeats the entire point of a website such as this. Just because you don't think that this is interesting to random people doesn't mean that random people don't think this is interesting.