4 ms·
Ask HN: Looking for compression algorithm that took 6GB to 800MB. Anyone know?
Looking for a good media compression & archiving algorithm for an app I'm building. More specifically, I'm looking an algorithm I had used 10 years ago but had totally forgotten the name of. Description: It was a self-extracting archive (about 800MB) that spat out a folder of over 6GB (mainly media: movies & audio). Now, the media files themselves were already in compressed format (MP3, MP4, etc...) so I was really impressed. Only drawback (as expected) was that it took over 1 hour to decompress on an Intel dual core. Does this kind of performance ring any bells for anyone? I just need a name.
- drvc33 11y agoI wanted to add: The self-extractor had no interface -- it was all in the windows command prompt. A message in the prompt said something about "Media compressor" and below it was the percentage extracted.
- toreriklinnerud 11y agoPied Piper?
- deleted 11y ago[deleted]
- samfisher83 11y agoinside out compression algorithm
- ChuckMcM 11y agoWe just to joke at NetApp that the Oil and Gas industry had the best compression algorithm, it could compress 100TB of seismic imaging data into a single bit {oil / no-oil } :-) I created a theoretical compressor which I haven't yet been able to implement which uses the fact that every sequence of bits appears in pi somewhere, so my compressor would just return the digit offset and the length of data. I keep looking for a source for all the digits of pi though, have yet to find it.
- digital_ins 11y agowhy not just create an online API for this? Sequences that are not yet "found" would go uncompressed, and as time goes by and internet speeds get faster and computing power gets cheaper, more sequences will be "found" for reference in that API and the algorithm will keep getting better. Alternatively, you could release a program that comes with an already long sequence of the decimal section of PI and connects regularly to the internet to download more digits (you could even run the computation & API on a google compute / amazon ec2 instance to keep costs low)
- komon 11y agoAfter having experimented a bit in high school with ideas for compression algorithms like this, I can tell you that one of the problems you'd run into is that you may have to look so deep into pi that the offset may actually take the same number of bits (or even more bits) than the sequence you're trying to find.
- ChuckMcM 11y agoIf you're up for it there is some great number theoretic work to be had in that statement :-). Given a sequence of digits X, P{X} E {Offset n .. Offset p} ? len{Offset n} > len{X}
- simon_acca 11y agoYou might find PiFs interesting: https://github.com/philipl/pifs https://github.com/philipl/pifs P.s. As the author states in issue #2, "the release date of pifs is an important part of understanding what's going on" ;)
- ChuckMcM 11y agoI really like that one too :-) Apparently there was a scam to sell a data set of all the US social security numbers, and the "decompressor" was just computing progressive digits of pi and printing them in xxx-yy-zzzz formatting.
- rahimnathwani 11y agoI don't believe there is any known lossless algorithm which can achieve 7.5:1 compression on already-lossy-compressed media files. If you want further lossy compression for existing media files, look at the newest algorithms supported by ffmpeg. If you want generic lossless compression, you can do a bit better than the usual suspects (gzip et al), but only if you're willing to put up with very slow compression times. If you have some other type of specific data (e.g. sparse files) then you could do something custom, but I guess this is unlikely to be your situation.
- insoluble 11y agoCall me a skeptic, but there is no way there was an algorithm available to humans 10 years ago that reversibly compressed 6GB of unique, already-compressed media files down to 800MB. The only way this could have happened is if there were shared files or shared segments between the files. For example, if a DVD had a bunch of audio tracks but some of the tracks were basically just direct copies of the others, then the compressor could recognise the similarity and capitalise thereon. For lossless compression of general data, 7-zip set on Ultra is probably the best available right now. On the other hand, algorithms such as FLAC or PNG work well for losslessly compressing uncompressed media.
- DanBC 11y agoYou can look through the software here: http://www.maximumcompression.com/benchmarks/benchmarks.php http://www.maximumcompression.com/benchmarks/benchmarks.php Or the software linked via here: http://prize.hutter1.net/ http://prize.hutter1.net/ As other people say, what you're asking for probably isn't possible.
- ta808945 11y agoI did quick internet search and only app mentioned by people that is able to achieve such compression rate is KGB Archiver[1]. And according to its wikipedia page it uses PAQ[2]. [1] https://en.wikipedia.org/wiki/KGB_Archiver https://en.wikipedia.org/wiki/KGB_Archiver [2] https://en.wikipedia.org/wiki/PAQ6 https://en.wikipedia.org/wiki/PAQ6