8 ms·
The Quite OK Audio Format for Fast, Lossy Compression
- rockstarflo 3y agoWhat is the tradeoff there?
- DamonHD 3y ago> QOA is slower than ADPCM, doesn't compress as much as MP3 and sounds worse than FLAC (duh). But I believe it fills a gap that was worth filling.
- jandrese 3y agoMP3 compression is very fast on modern hardware. This may have a niche for low power devices, especially if they are battery constrained.
- dale_glass 3y agoIt's probably something we (https://overte.org/ https://overte.org/) can use. We have a 3D environment with spatial audio. Audio is encoded server-side, and since it's spatial everyone needs their own mix. We're using Opus, and audio encoding turns out to be the usual limiting factor on small servers. So this kind of thing is exactly up our alley: an alternate option that uses less CPU than Opus, but consumes less bandwidth than raw audio. But adding supporting for FLAC is also on our list. It seems nicely performant when compared to Opus.
- brnt 3y agoDoesnt Opus (speex?) have some low CPU settings?
- dale_glass 3y agoIt does, and I've tried tweaking that, but the performance difference isn't very significant. I appear to be able to get maybe 30% better performance -- pretty nice, but not nearly big enough especially on low end servers.
- doublepg23 3y agoI'm not sure much gets better latency than Opus but LyraV2 seemed interesting https://opensource.googleblog.com/2022/09/lyra-v2-a-better-faster-and-more-versatile-speech-codec.html https://opensource.googleblog.com/2022/09/lyra-v2-a-better-f...
- dale_glass 3y agoCould be an option, but we take high audio quality as a point of pride and encode in Opus 128k by default. Audio doesn't only include speech but also any sound effects, media present in-world, etc. But that might be an interesting experiment. Right now the low cpu usage/high quality/faily high bandwidth usage category is something we're looking to have an option for.
- nmfisher 3y agoLyra is a speech-only codec, so it's apples and oranges to compare with Opus for general-purpose audio compression.
- a2128 3y agoI'm curious, why encode audio server-side? Other games in this genre I've seen seem to have clients do encoding/decoding, and do the spatial audio clientside, with the server just passing each user's audio and position data along from client to client. Especially in VR where ideally there should be no latency between turning your head and the audio shifting. Are there any reasons to do this on the server, or am I misunderstanding something?
- kragen 3y ago"very fast" could mean many different things that vary by orders of magnitude
- kragen 3y ago"very fast" could mean many different things that vary by orders of magnitude in https://phoboslab.org/log/2023/04/qoa-specification https://phoboslab.org/log/2023/04/qoa-specification he got ffmpeg on one core of an i7-6700k (which is arguably 'modern hardware') to encode a 9807-second file in mp3 in 146.2 seconds, 67× faster than real time. but qoa was 25.75 seconds, 5.7 times faster than that. qoa decoding was 2.5× as fast as dr_mp3 you can imagine situations where reducing the number of audio encoding servers in your audio encoding cluster by a factor of 6 would be a big win, or where you want to encode 100+ audio streams in real time on your laptop (maybe an sdr tuned to every am radio station at once), but i agree with you that battery-constrained devices are a more likely application area: making your audio recorder battery last twice as long is a much bigger win
- daneel_w 3y agoThat in terms of quality per any bitrate it comes nowhere near ubiquitous formats like AAC or MP3 when produced with good encoders. But it's good to have (possibly) patent-free solutions available.
- Pet_Ant 3y agoWhat is the LFE channel? It should be spelled out explicitly, but I figured out the rest L-Left,R-Right,C-Center,FL-Front Left,FR-FrontRight,SL-SideLeft,SR-SideRight,BL-BackLeft,BR-BackRight --- Edit: LFE-LowFrequencyEffects... so subwoofer? https://www.dolby.com/uploadedFiles/Assets/US/Doc/Professional/38_LFE.pdf https://www.dolby.com/uploadedFiles/Assets/US/Doc/Profession...
- bravura 3y agoLow frequency energy, I assume. Ie bass. Your subwoofer or „bottoms“ if you have several.
- ok_dad 3y agoLFE is usually a bass shaker which is a subwoofer but it moves a weight instead of a cone, so you get vibrations in your seat. It stimulates movement to your body somewhat, I use two for my sim racing rig, one under my seat to inform me of the car dynamics and immersive feeling, one under my pedals to inform me when ABS is active and when my tires are spinning.
- samplatt 3y agoLFE can mean "bass shaker", but it's an industry-standard term invented by Dolby that effectively means "between 3 and 120hz", which usually means "subwoofer". These days crossover points are very configurable. Most bass shakers are rated for use between 20hz and 200hz.
- entropicdrifter 3y agoLFE is an industry standard term for the subwoofer channel. It's the ".1" in "5.1","6.1","7.1" etc
- ogurechny 3y agoLFE audio channel is different from subwoofer output. Subwoofers come with multichannel audio systems in which directional speakers usually can't cover the lower range of audio frequencies. They are responsible for bass content from all channels, and get it from software or hardware crossover filter which is independent from specific input formats. Placement of low frequency speaker does not matter much because of human perception. LFE track is an additional effects channel for movie theaters and similar amusement rides in which audio system plays low frequencies from other channels just fine. Dedicated LFE emitter then adds rattling and other wub-wub effects without overloading audio speakers with all that extra energy. Movies that lack car chases and explosions routinely have completely silent LFE tracks.
- ape4 3y agoWhat's going to be the next Quite OK thing?
- bartwe 3y agoHopefully a movie format
- WithinReason 3y agoMPEG1 is actually quite OK
- p1mrx 3y agoQuite OK Food. It tastes like sand but the shelf life is above average.
- dvh 3y agoQuite OK browser. It doesn't have webgl, webgpu or other fancy and easy to exploit stuff, but it renders 95% of websites and source code is easy enough to be maintained with very few people.
- kragen 3y agomaybe links2 or dillo?
- extua 3y agoTinyVG follows the similar goals: an alternative to SVG with a specification which trades off features for simplicity. https://tinyvg.tech/ https://tinyvg.tech/
- Turing_Machine 3y agoI looked around, but didn't see any mention of potential patent issues. I assume that this has been considered? The Ogg Vorbis people spent a lot of time on that back when they were developing their format. Other than that, looks great!
- speedgoose 3y agoThe website says it’s made in Hesse. No software patents to care about there. https://en.m.wikipedia.org/wiki/Software_patents_under_the_European_Patent_Convention https://en.m.wikipedia.org/wiki/Software_patents_under_the_E...
- Turing_Machine 3y agoMaybe not, but that doesn't help people who aren't using it in the EU.
- speedgoose 3y agoTrue, it hasn’t stopped hobbyists from using x264, ffmpeg or VLC in the past but that would probably prevent companies in some markets to use this audio format.
- morelisp 3y agoProbably the most infamous audio format patent ever was owned by a German research institute.
- speedgoose 3y agoI guess you mean this one: https://patents.google.com/patent/US5812672 https://patents.google.com/patent/US5812672 That was an USA patent from Fraunhofer, who made quite some cash from mp3 license fees (100 000 000€ according to Wikipedia).
- nullc 3y ago
- g0xA52A2A 3y agoSome previous discussions. 3 months ago - https://news.ycombinator.com/item?id=35738817 https://news.ycombinator.com/item?id=35738817 6 months ago - https://news.ycombinator.com/item?id=34625573 https://news.ycombinator.com/item?id=34625573
- kragen 3y agothank you very much these had crucial information for me
- codeflo 3y agoIt's interesting that this works in the time domain (instead of frequency domain), and I wonder what the resulting quality limitations are, if any. The sound samples on the demo page, at the least the dozen I clicked on, didn't seem all that challenging. Few, mostly synthesized instruments, low dynamic range. My ears aren't good enough to evaluate audio codecs anyway, however.
- marcoc 3y agoHow can one create a professional looking pdf like the QOAF specification one?
- deleted 3y ago[deleted]
- GraemeMeyer 3y agoTwo-column layout in Microsoft Word, large header, smaller footer, with appropriate font choices would get you basically all the way there.
- crumpled 3y agoI looked at the PDF, and can confidently say I could typeset that in a word processor, using a stylesheet to sustain it. That's not what they did, apparently. The document properties call out https://cairographics.org https://cairographics.org
- jfk13 3y agoHTML+CSS, converted to PDF via the Save As PDF feature in Firefox. (Or the same could be done with other browsers, but this one apparently comes from FF.)
- kragen 3y agohas anyone benchmarked qoa to see roughly how many instructions per sample it needs? all i see here is that it's more than adpcm and less than mp3, but those differ by orders of magnitude like, can you reasonably qoa-compress real-time 16ksps audio on a 16 megahertz atmega328? hmm, https://phoboslab.org/log/2023/04/qoa-specification https://phoboslab.org/log/2023/04/qoa-specification has some benchmark results, let's see... seems like he encoded 9807 seconds of 44.1ksps stereo in 25.8 seconds and decoded it in 3.00 seconds on an i7-6700k running singlethreaded. what does that imply for other machines? it seems to be integer code (because reproducibility between the predictor in encoding and decoding is important, and a significant part of it is 16-bit. https://ark.intel.com/content/www/xl/es/ark/products/88195/intel-core-i76700k-processor-8m-cache-up-to-4-20-ghz.html?ui=BIG https://ark.intel.com/content/www/xl/es/ark/products/88195/i... says it's a 4.2 gigahertz skylake. agner says skylake can do 4–6 ipc (well, μops/cycle) https://www.agner.org/optimize/blog/read.php?i=628 https://www.agner.org/optimize/blog/read.php?i=628, coincidentally testing on an i7-6700k himself, but let's assume it's 3 ipc, because it's usually hard to reach even that level of ilp in useful code so that's about 380 μops per sample if i'm doing my math right; that might be on the order of 400 32-bit integer instructions per sample on an in-order processor. if (handwaving wildly now!) that's 600 8-bit instructions, the atmega328 should be able to encode somewhere in the range of 16–32 kilosamples per second so, quite plausibly for decoding the same math gives 43 μops per sample rather than 380 i'm very interested to hear anyone else's benchmarks or calculations
- MobiusHorizons 3y agoSeems to have similar design criteria as opus but I don’t see any comparison.
- Aldipower 3y agoAn _audio_ format which is _quite_ ok? Not sure, if I need that.
- ericls 3y agoThe smaple page preloads all the files before playing... Which wastes lots of bandwidth.
- mips_r4300i 3y agoComparing against 4bit ADPCM, which is already able to give quite good performance as long as your sample rates are relatively modern, this only improves it to 3.2 bits. It is fast, but ADPCM is also fast. Would be nice to see joint stereo support. If you were to take ADPCM or this OK format and try to encode any stereo music with it, you will need 2 channels. However, there is an extremely advantageous optimization that can be made here - most music is largely center panned, so both channels are almost the same. With joint stereo you record one channel (either by picking one or mixing to an average) and then you can store the difference for the other channel which will occupy a lot fewer bits, assuming you are able to quantize away the increased entropy. For example, instead of using two 4bit ADPCM channels for stereo, which would only be a 50% savings over uncompressed, you could probably use an average of 5 bits per sample.
- anotherhue 3y ago> Would be nice to see joint stereo support This was/is available in MP3 since forever, so seems a reasonable request. https://wiki.hydrogenaud.io/index.php?title=Intensity_stereo https://wiki.hydrogenaud.io/index.php?title=Intensity_stereo
- gaazoh 3y agoI like the philosophy of QOA (and other similar projects, including QOI and TinyVG), but unlike others, it seems like it's not ready to use yet, see https://github.com/phoboslab/qoa/issues/25 https://github.com/phoboslab/qoa/issues/25 > I have just pushed a workaround to master. [...] > This still introduces audible artifacts when the weights reset. It prevents the LMS from exploding, but is far from perfect :/ This, combined with the fact that that issue is still open mean that a breaking change is still to be expected.