17 ms·
FLAC – Format overview
- jsbg 4y agotl;dr: the algorithm splits the audio into blocks (default is 4096 samples per block). Then a polynomial is used to approximate the wave for the block, and residuals are calculated for each sample by subtracting the approximation. The residuals, because their magnitude is smaller than the original signal, require fewer bits to encode. Hence the smaller file without loss of information.
- aidos 4y agoNice. Not dissimilar from what jpeg2000 does (from memory - correct me if I’m wrong!). They use the previous pixel values in order to “guess” the next value and then store the delta, which tends to be a smaller number on average, giving better compression.
- mistrial9 4y agoproximity in a 2D image has neighbors in north-south-east-west (NSEW) on a 2D grid, not just pixel neighbors per-scanline (row of bytes); proximity in a sound file has harmonics and transitions of some plural set of signal, but the sound file neighbors are roughly at a point in time. corrections welcome
- robotresearcher 4y agoTrue. The fact of locality in real-world 2D images means you can predict a pixel's value by picking any neighboring pixel, or indeed any nearby pixel, with worsening performance with L2 distance. There's no reason I can think of to prefer a direction, except for convenience in the data structure.
- cryptonector 4y agoPlus the use of Rice codes for encoding the residuals in as few bits as possible.
- derefr 4y agoCurious to me why only such a simple representation (polynomial) is used for the approximation. Couldn't you get a much better approximation (⇒ more compressible residual) in about the same amount of storage space + CPU decoding effort, using a formula based on e.g. wavelet LUTs as your primitive (plus maybe a larger block size)? Are there other, more advanced lossless encodings that do this? And if so, why didn't they catch on compared to FLAC? For that matter, is there a lossless format that just embeds a regular lossy encoding of the audio as the approximation, and then computes+stores the residual relative to that? (I'm guessing that this wouldn't work well for some reason, but I'm not sure what that reason would be.) (ETA: the other later lossless audio formats that I'm personally aware of — ALAC, Monkey's Audio, and WavPack — all seem to use linear prediction. Seemingly they were all designed under the presumption of the constraint that the encode step must be able to be done in hardware / with fixed-sized memory buffers; rather than allowing that you can load the whole PCM audio file into memory and do things like FFT to it. Possibly made sense in the late 90s, when a PC's RAM wasn't much larger than five minutes of uncompressed audio. Doesn't really make sense today. Maybe we're due for a new lossless audio encoding?)
- hcs 4y agoI just ran across MPEG-4 SLS [1] (née "AAZ" for "Advanced Audio Zip", aka HD-AAC), which is a lossless format on top of AAC's MDCT, and which is patented, unfortunately. That led me to other coders: DTS-HD Master Audio [2] (née DTS++) and OptimFROG DualStream [3]. OptimFROG's DualStream mode is similar to WavPack's hybrid mode, and DTS-HD MA uses DTS Coherent Acoustics (based on ADPCM), so none of these are based on a perceptual lossy codec besides MPEG-4 SLS. (Sorry to keep replying here, I keep stumbling on interesting things after the edit window closes) [1] https://en.wikipedia.org/wiki/MPEG-4_SLS https://en.wikipedia.org/wiki/MPEG-4_SLS [2] https://en.wikipedia.org/wiki/DTS-HD_Master_Audio https://en.wikipedia.org/wiki/DTS-HD_Master_Audio [3] http://losslessaudio.org/DualStream.php http://losslessaudio.org/DualStream.php
- hcs 4y ago> is there a lossless format that just embeds a regular lossy encoding of the audio as the approximation, and then computes+stores the residual relative to that? It's not "regular lossy", but WavPack does allow separating lossy from the residual in hybrid mode. I think this is rarely done with DCT-based stuff because there's so much potential imprecision in the decoders.
- black_knight 4y agoAs always with compression there is no magic. There are wav files that are smaller than the equivalent flac file (perhaps even the majority of possible WAV files are shorter than their equivalent FLAC). But it so happens that the sounds we actually want to store are very well approximated by polynomials.
- drawingthesun 4y agoWould applying a lossless compression algorithm to the data such as zip or 7z after the algorithm you describe reduce the file even size even more? Wondering if FLAC already does this or if such a feature could be added?
- tgv 4y agoThese methods perform poorly on audio, so even worse on compressed audio. Just try it on a wav and a flac file.
- liftm 4y agoSounds quite simple already. I still wonder if for FLAC there's an equivalent of what QOI is to PNG.
- pier25 4y agoAnyone knows if FF has solid FLAC support now? I know a couple of years ago it used to be flaky. Eg: https://bugzilla.mozilla.org/show_bug.cgi?id=1528265 https://bugzilla.mozilla.org/show_bug.cgi?id=1528265
- shoghicp 4y agoStill flaky with double ID3 tags (front/back) or issues parsing metadata (so it fails decoding), but the issues with variable block size were solved. Most FLAC I handle on a large system play fine. There are other issues related to streaming the FLAC via Range requests depending if it is WebAudio, <audio> or directly in a tab, however this applies to all audio/media in general.
- brnt 4y agoFlac has Flac tags, how did Id3 get up in there?
- lifthrasiir 4y agoID3 is format-agnostic so it shouldn't be no surprise to see it in flac. Even the linked article mentions that flac does recognize and skip ID3 tags.
- brnt 4y agoYou have to go out of your way to add them though. It's not standard.
- shoghicp 4y agoA lot of tagging programs add them as they just get appended or prepended to the file regardless their content. Sometimes one program reads a format but writes in the other, I have personally found a FLAC in the wild with ID3v2 at the front, ID3 at the back (corrupt) and FLAC metadata as well. The FLAC metadata inside was wrong, the valid one was outside at the front (sadly it was a broken JIS encoding). There is no sane approach to media. Between legacy formats, legacy tagging, and all kinds of implementation specific bugs, you are in for a ride if you want to cater to more than one decoder out there :)
- deleted 4y ago[deleted]
- bretbernhoft 4y agoI've only ever used .WAV and .MP3 formats. What would a .FLAC file format be used for?
- andai 4y agowav but smaller
- redcalx 4y agoWAV is uncompressed. MP3 is highly compressed, but lossy. FLAC is compressed and lossless. If you wanted to store lots of master copies of audio, you could use WAV or you could reduce storage space by using FLAC instead.
- samstave 4y agoBut you cant (up)convert WAV to FLAC? right? so if your source is WAV or MP3, FLAC is irrelevant?
- alexanderh 4y agoWhat? Most all FLAC is created (converted) from a WAV source. If your source is MP3, then yes, FLAC is irrelevant... But FLAC is basically a lossless compression format for WAV. I'm really confused by what you're talking about lol... especially "up"convert...??? WAV is the ultimate lossless audio on PC. It really doesn't get any better than WAV. There is no "up" from WAV. FLAC is a compression format for WAV, that does not lose any data. The output of FLAC will be identical to the WAV file, even though its compressed. MP3 is a compression format for WAV that loses data, and will not be identical to the original WAV file.
- Maursault 4y agoWAV this and WAV that. In 1988, Apple developed the Audio Interchange File Format (AIFF), which is uncompressed pulse code modulation (PCM). PCM is what is stored on CDs, so any Mac with a CD-ROM drive attached will recognize the PCM information on Red Book audio CD's as AIFF files. Inexplicably, 3 years later, Microsoft and IBM developed the Resource Interchange File Format (RIFF) in 1991, of which the WAV format is one implementation. RIFF doesn't store PCM. Instead it stores various formats of data in 4 byte "chunks." Depending on the audio file format specified, one can always distinguish a Windows user from an audio professional (or a Mac user), because since about 1990, the vast majority of professional audio recording (tracking, mixing and mastering) studios have been exclusively Mac shops, including such greats as Skywalker Sound and Abbey Road Studios.
- jansan 4y agoI do not have very sensitive ears, but maybe an audio enthusiast can explain this to me: In the 80s and 90s some people were going crazy over HiFi, only the absolute high end products were just good enough. I remmeber seeing stereo systems for 50,000$ and more. CDs were already seen as inferior to records quality-wise, and speakers had to be huge if possible. Today Wifi speakers are all the rage. The music is downloaded (precompressed) and then sent over Wifi or Bluetooth with (sometimes very) limited bandwidth to a single speaker which has the size of a laptop. How does the audio quality compare? Is it like day and night? Or do the new multi room systems play in the same league as the old system that were used by enthusiasts? I often have the feeling that overall sound quality does not matter anymore as long as the bass is strong enough, but as I said at the beginning, my ears are not very sensitive.
- jsmith99 4y agoThe signal is digital until it gets to the speaker so there is no loss of quality after the initial encoding, which might itself be lossless. Some systems (eg Sonos) can stream lossless audio or even 24 bit 192 khz, which is certainly overkill.
- malermeister 4y agoWith some wireless systems (Bluetooth comes to mind), there's a lossy reencoding between playback device and speaker too!
- AshamedCaptain 4y agoLossy, but at >300kbps.... this is not going to be the part that introduces most artifacts.
- chrisseaton 4y ago> The signal is digital until it gets to the speaker so there is no loss of quality after the initial encoding, which might itself be lossless. This ain’t true - for example Bluetooth is lossy compression. ‘Digital’ doesn’t mean lossless from the source.
- MrStonedOne 4y ago
- mustache_kimono 4y agoFLAC makes lots of smart decisions. One constraint on lossless audio should, of course, be that it's actually lossless (which the test suite confirms), but also that it makes it easy for the user to confirm the audio streams are the same as the decoded WAV file (`flac -t`) by storing a checksum of the stream. I have often wondered why other media formats don't do a similar thing, especially since changing a media file's tags (which can change the checksum of a file) or name (which makes external verification from txt file difficult) is quite common. I even wrote a utility[0] that uses ffmpeg to hash all ffmpeg compatible bitstreams, and store their hashes in a xattr (yes, with lots of other options to test and compare, etc.), but all media formats should just be as clever (and care as much) to do this natively, like FLAC. I mean -- why not? [0]: https://github.com/kimono-koans/dano https://github.com/kimono-koans/dano
- lake_vincent 4y agoWhat are the reasons why not? What are the obstacles/drawbacks to the FLAC approach? I have no domain knowledge here, just curious
- lazide 4y agoFor the most part, no one cares to do so is why. Hell, 99% of all computers are still running on non-checksummed filesystems, if not 99.99%. For the most part, only a tiny minority of people are even aware of bit rot, let alone have anything they simultaneously care enough about to protect against, and have enough ownership of it to attempt to try.
- hnuser123456 4y agoAre there any good ways to tell if audio has been lossy compressed but then re-encoded in a lossless format, or deceptively high lossy bitrate?
- dmitri_ignat 4y agoYou can view it with a spectrogram, audio which has been lossily encoded will have a telltale cutoff in the high frequencies. The lossier the encoding, the lower the frequency ceiling will be.
- Cupertino95014 4y agoMy pandemic project was to flac-ify all my music, including the vinyl records. It still fits on a microSD card, so the phone now has everything. It's funny how often people still assume "on the computer" means "MP3." I don't know why you'd put up with any loss of quality anymore, even if you personally can't hear the difference.
- wheels 4y agoI don't have a crazy large music collection, and mine is 85 GB of 192 kbps MP3s. With FLAC that'd be 250-ish GB of music. That's a significant chunk of most modern SSDs. Why would you use 3x the storage space if you can't hear the difference for a non-trivial percentage of your available storage? Literally, by definition, according to your own terms, it serves no purpose. I'm a musician and audio developer, and it's only really in my own music, that I've listened to over and over again while creating it, that I notice the degradation in a 192 kbps MP3 – and occasionally in the high-hats of CDs that I listened to hundreds of times in high school. FLAC's great, and definitely serves a purpose, but I use it mainly for archives, not for casual listening.
- blibble 4y ago> With FLAC that'd be 250-ish GB of music. at current prices that's about $25 of SSD
- doubleunplussed 4y agoOne cannot simply add more SSDs to one's laptop, regardless of price.
- Cupertino95014 4y agoBut one can buy a multi-TB drive that plugs into the laptop's USB 3.0 port. The "multi" part of it keeps getting larger, for the same price. And you can keep an extra one at another site in case your house burns down.
- 4y ago
- armchairhacker 4y agoI wish apple supported FLAC. There’s an amazing free tool to convert to/from ALAC (https://tmkk.undo.jp/xld/index_e.html https://tmkk.undo.jp/xld/index_e.html), but still: some players only support FLAC and some ALAC and sometimes you can’t transfer songs in either format because of the no overlap between source/target.
- esperent 4y agoI'm always mystified when someone says they wish their general purpose computing device running a Unix based OS supported X feature. How can an Apple device not support FLAC? Are they really that locked down that you can't do something basic like install a codec?
- CharlesW 4y ago> How can an Apple device not support FLAC? Are they really that locked down that you can't do something basic like install a codec? You can't install a codec globally on iOS, but there are many, many player apps that support FLAC: https://www.igeeksblog.com/best-music-player-apps-for-iphone-ipad/ https://www.igeeksblog.com/best-music-player-apps-for-iphone... For folks who think Apple Music is the best music app for iOS, it's not difficult to convert songs, albums, or whole music libraries to ALAC.
- brewdad 4y agoI've been happy with Flacbox. Obviously it supports FLAC but it also supports OPUS which is great for the portability/sound quality tradeoff.
- savoytruffle 4y agoALAC (Apple Lossless …) is not much newer than FLAC — it started in mid 2004 as the audio transport to AirPlay (previously called AirTunes) to AirPort Express. I think they prefer it because it's in a MPEG4 container, it can be DRM encumbered with their "FairPlay" technology that they hadn't used for a while, but now use again for subscription Apple Music.
- doublepg23 4y agoThe videos Xiph did a while back are excellent for understanding some digital audio concepts (though it’s somewhat dense with jargon). https://www.xiph.org/video/ https://www.xiph.org/video/
- jancsika 4y agoFor the modelling step, is there an aw-fuck-it special case branch for white noise?
- dTal 4y agoI think I recall reading that there is, and another one for silence, making a total of four frame types.
- pornel 4y agoZeroing out data in the residual coding step would make FLAC a lossy compression format! (and most likely not a very good one). I'm tempted to implement it for giggles, but OTOH I'm worried it could spark some kind of uprising of audiophiles.
- suzzer99 4y agoVery disappointed to search this comment stream and not find one instance of "middle out".
- eapriv 4y agoI can’t understand the very first step in the compression process: “The left and right channels are converted to center and side channels through the following transformation: mid = (left + right) / 2, side = left - right. This is a lossless process, unlike joint stereo.” How is this a lossless process?
- ex3ndr 4y agoBecause you can recover left and right from mid and side, you can also multiply both signals by two to make division result to be always integer value.
- eapriv 4y agoSurely multiplying both signals by two is changing them? What if overflow happens? How do you store the information that it needs to be divided back when decoding?
- ruuda 4y agoThe result is one bit wider, so it does not overflow.
- eapriv 4y agoThen why divide by two at all? Also, adding one bit for every value (or pair of values) means that data size gets a significant increase from the start.
- edflsafoiewq 4y agoThe point is to decorrelate the channels. The left and right channels are usually mostly the same, so coding them both would basically send the same signal twice. With perfect decorrelation, the side channel would become zero, instantly saving half the bits. (It's the same idea for images, where the RGB channels mostly all look like grayscale copies of the image, so the YCbCr/YCoCg transform is done to decorrelate them.) There are actually three methods in FLAC: mid/side, left/side, and right/side. Each frame can use a different method and stores what method it used (if any) in the frame header. The difference has a range 1 bit larger than the original, but this doesn't matter that much since everything is getting compressed anyway. Anyway the bit is only used if the side channel is very large, ie. the correlation was poor, in which case it would be better not to use a decorrelation for this frame. Interchannel Decorrelation: https://www.ietf.org/archive/id/draft-ietf-cellar-flac-04.html#section-5.2 https://www.ietf.org/archive/id/draft-ietf-cellar-flac-04.ht... Channel Bits in Frame Header: https://www.ietf.org/archive/id/draft-ietf-cellar-flac-04.html#name-channels-bits https://www.ietf.org/archive/id/draft-ietf-cellar-flac-04.ht...