5 ms·
Fifteen years after the initial release of FLAC – have there been any significant developments in the lossless compression of audio since then? I know there’s
by mdf 10y ago
Fifteen years after the initial release of FLAC – have there been any significant developments in the lossless compression of audio since then?
I know there’s FLIF[1] for lossless image compression and Zstandard[2] for general purpose lossless compression that have recently hit the Hacker News front page. Are their adopted techniques not suitable for audio?
[1] http://flif.info/ http://flif.info/
[2] https://code.facebook.com/posts/1658392934479273/smaller-and-faster-data-compression-with-zstandard/ https://code.facebook.com/posts/1658392934479273/smaller-and...
- niftich 10y agoLet's see: - Wavpack [1], which is a rough contemporary but offers three tiers of presets (normal scale, high scale, extra high scale) and an innovative (and optional) lossy/hybrid mode - TAK [2] which compressed better and decoded faster than either, but was initially closed-source until the dev was persuaded to open it up - LossyWAV [3] which isn't lossless but chops off least-significant-bits while using noise shaping to pre-process audio and make it compress better when fed to a lossless compressor Most of these developments were first publicized on Hydrogenaudio. But as for innovations in the last two years, not that I'm aware. [1] http://wiki.hydrogenaud.io/index.php?title=WavPack http://wiki.hydrogenaud.io/index.php?title=WavPack [2] http://wiki.hydrogenaud.io/index.php?title=TAK http://wiki.hydrogenaud.io/index.php?title=TAK [3] http://wiki.hydrogenaud.io/index.php?title=LossyWAV http://wiki.hydrogenaud.io/index.php?title=LossyWAV EDIT (for some more background): generally in lossless audio compression you want to use linear prediction to predict an approximate signal for the next few samples, then encode the difference between your predicted guess and the actual signal in some entropy coder, like Golomb-Rice codes or Huffman or Arithmetic coding. Although most of Zstandard's improvements are algorithmic or implementation-related and not related to data theory, the part that could show promise is the tANS entropy coder [4] used in Zstandard; but Golomb-Rice codes perform well for data that comes from linear predictors; so I'm not sure what to expect [5]. [4] https://github.com/Cyan4973/FiniteStateEntropy https://github.com/Cyan4973/FiniteStateEntropy [5] 'Benchmarks' section under [4]
- cpeterso 10y agoThere is also ALAC (Apple Lossless Audio Codec), which has been open source and royalty-free since 2011: https://en.wikipedia.org/wiki/Apple_Lossless https://en.wikipedia.org/wiki/Apple_Lossless
- niftich 10y agoIn all fairness ALAC is very similar to FLAC in its inner workings, but differs in arbitrary, minor ways that result in more complexity over FLAC (and of course, incompatibility) for little gain. I'm paraphrasing from a post from the FLAC developer himself [1], after someone released an ALAC decoder created through reverse-engineering the format, back in 2005. [1] https://hydrogenaud.io/index.php/topic,32111.msg279843.html#msg279843 https://hydrogenaud.io/index.php/topic,32111.msg279843.html#...
- sjwright 10y agoJudging a codec by a reverse-engineered implementation may be very misleading. For example, the original codec may have been written to satisfy a different set of goals, e.g. optimised for energy consumption on a particular ARM CPU.
- arthurfm 10y agoWould lossless and/or lossy compression algorithms perform better if each element of an audio track was compressed and stored separately (similar to the MP4-based STEM audio format [1], but without any limitations on the maximum number of elements and the stereo master which increases the file size)? The constituent parts of the track would then be merged together during playback by the audio player. [1] http://www.stems-music.com/stems-faq/ http://www.stems-music.com/stems-faq/
- niftich 10y agoWe had these a long time ago; 'module files' made with music trackers [1]. And yes, they compress very well. Unfortunately, the transformations that make up 'mastering' can be pretty elaborate, and for professionally-produced music there is little incentive to let the general public see their project files -- although they are occasionally made available for remixers. [1] https://en.wikipedia.org/wiki/Module_file https://en.wikipedia.org/wiki/Module_file
- ars 10y ago> have there been any significant developments in the lossless compression of audio since then? One thing that is lacking in audio compression is the same thing as the difference between JPEG and MPEG - using past data to predict future data and only storing the difference. Music has lots of repeating notes and passages. Isolating them (from other notes played at the same time), and only storing the slight change of how the note was played this time, vs last time should greatly increase compression ratios. But I have not seen any audio compression that does this. (Note this applies equally to lossless and lossy compression.)
- niftich 10y ago> using past data to predict future data and only storing the difference Actually what you've described is essentially how all of lossless audio compression really works. You pre-process the signal to make it so that the 'important' parts don't take a bunch of space to store, then you feed it to a predictor, and then encode your difference in a way that it doesn't take a bunch of space to store. You can try to tune each step to try to make your next step perform better. Lossy compression can be made to work similar, because you can just store a less accurate difference between what you predict and the original.
- ars 10y agoMy understanding is that the predictor only looks at the wave directly before the current one, and not the entire file. Delta encoding basically. They also (as far as I know) don't attempt to isolate notes or voice phonemes. I.e arithmetic coding for sound.
- niftich 10y agoThe predictor looks at a 'window', which is usually defined in terms of n samples. If you make your window too short, you are not capturing any meaningful periodicity; if you make it too long you have sacrificed some error resilience and decoding convenience in exchange for hoping to get a better match for your predictor. Speech codecs (lossy) basically operate just like you describe, see [1] [1] https://en.wikipedia.org/wiki/Vocoder#Modern_implementations https://en.wikipedia.org/wiki/Vocoder#Modern_implementations
- ris 10y agoAs much as I like FLAC, I do start to wonder about the necessity of native support for the more minority file formats in browsers when decoders can be written in js/asm.js/webassembly. Every natively supported file format adds to the attack surface of the browser - another piece of decoding code running outside the sandbox/vm that js decoders would be forced to.