13 ms·
Opus 1.5 released: Opus gets a machine learning upgrade
- travisporter 3y agoVery cool. seems like they addressed the problem of hallucination. would be interesting to see an example of it hallucinating without redundancy and corrected with redundancy
- CharlesW 3y agoIsn't packet loss concealment (PLC) a form of hallucination? Not saying it's bad, just that it's still Making Shit Up™ in a statistically-credible way.
- derf_ 3y agoThe PLC intentionally fades off after around 100 ms so as not to cause misleading hallucinations. It is really just about filling small gaps.
- skybrian 3y agoIn a broader context, though, this happens all the time. You’d be surprised what people mishear in noisy conditions. (Or if they’re hard of hearing.) The only thing for it is to ask them to repeat back what they heard, when it matters. It might be an interesting test to compare what people mishear with and without this kind of compensation.
- jmvalin 3y agoAs part of the packet loss challenge, there was an ASR word accuracy evaluation to see how PLC impacted intelligibility. See https://www.microsoft.com/en-us/research/academic-program/audio-deep-packet-loss-concealment-challenge-interspeech-2022/results/ https://www.microsoft.com/en-us/research/academic-program/au... The good news is that we were able to improve intelligibility slightly compared with filling with zeros (it's also a lot less annoying to listen to). The bad news is that you can only do so much with PLC, which is why we then pursued the Deep Redundancy (DRED) idea.
- tialaramex 3y agoRight, this is why the Proper radio calls for a lot of systems have mandatory read back steps, so that we're sure two humans have achieved a shared understanding regardless of how sure they are of what they heard. It not only matters whether you heard correctly, it also matters whether you understood correctly. e.g. train driver asks for an "Up Fast" block. His train is sat on Down Fast, the Up Fast is adjacent, so then he can walk on the (now safe) railway track and inspect his train at track level, which is exactly what he, knowing the fault he's investigating, was taught to do. Signaller hears "Up Fast" but thinks duh, stupid train driver forgot he's on Down Fast. He doesn't need a block, the signalling system knows the train is in the way and won't let the signaller route trains on that section. So the Up Fast line isn't made safe. If they leave the call here, both think they've achieved understanding but actually there is no shared understanding and that's a safety critical mistake. If they follow a read-back procedure they discover the mistake. "So I have my Up Fast block?" "You're stopped on Down Fast, you don't need an Up Fast block". "I know that, I need Up Fast. I want to walk along the track!" "Oh! I see now, I am filling out the paperwork for you to take Up Fast". Both humans now understand what's going on correctly.
- a_wild_dandan 3y agoTo borrow from Joscha Bach: if you like the output, it's called creativity. If you don't, it's called a hallucination.
- CharlesW 3y agoI love that, what's it from? (My Google-fu failed.) Unexpected responses are often a joy when using AI in a creative context. https://www.cell.com/trends/neurosciences/abstract/S0166-2236(00)02081-6 https://www.cell.com/trends/neurosciences/abstract/S0166-223...
- a_wild_dandan 3y agoIt was from one of his podcast appearances. Which doesn't narrow it down much, unfortunately. Most likely options: https://www.youtube.com/watch?v=LgwjcqhkOA4 https://www.youtube.com/watch?v=LgwjcqhkOA4 https://www.youtube.com/watch?v=sIKbp3KcS8A https://www.youtube.com/watch?v=sIKbp3KcS8A https://www.youtube.com/watch?v=CcQMYNi9a2w https://www.youtube.com/watch?v=CcQMYNi9a2w
- Aachen 3y agoThat sounds funny, but is it true? Certainly there's a bias that goes towards what you're quoting, but would you otherwise genuinely call the computer creative? Is that a positive aspect of a speech codec or of an information source? Creative is when you ask a neural net to create a poem, or something else from "scratch" (meant to be unique). Hallucination is when you didn't ask it to make its answer up but to recite or rephrase things it has directly observed That's my layman's understanding anyway, let me know if you agree
- skybrian 3y agoThat's almost the same. You could say it's being creative by not following directions. Creativity isn't well-defined. If you generate things at random, they are all unique. If you then filter them to remove all the bad output, the result could be just as "creative" as anything someone could write. (In principle. In practice, it's not that easy.) And that's how evolution works. Many organisms have very "creative" designs. Filtering at scale, over a long enough period of time, is very powerful. Generative models are sort of like that in that they often use a random number generator as input, and they could generate thousands of possible outputs. So it's not clear why this couldn't be just as creative as anything else, in principle. The filtering step is often not that good, though. Sometimes it's done manually, and we call that cherry-picking.
- jmvalin 3y agoWell, there's different ways to make things up. We decided against using a pure generative model to avoid making up phoneme or words. Instead, we predict the expected acoustic features (using a regression loss), which means that model is able to continue a vowel. If unsure it'll just pick the "middle point", which won't be something recognizable as a new word. That's in line with how traditional PLCs work. It just sounds better. The only generative part is the vocoder that reconstructs the waveform, but it's constrained to match the predicted spectrum so it can't hallucinate either.
- Sonic656 3y agoThere something darkly funny that Opus could act psychotic because It glitched out or was fed something really complex. But you could argue transparent lossy compression at 80 ~ 320kbps is a controlled Deliriant like hallucination going by how only rare few can tell them apart from Lossless.
- p1esk 3y agoTwo inrelated “Opus” releases today, and both use ML. The other one is a new model from Anthropic.
- deleted 3y ago[deleted]
- out_of_protocol 3y agoWhy the hell opus still not in Bluetooth? Well i know - sweet sweet license fees (aKKtually, there IS opus codec, supported by pixel phones - google made it for VR/AR stuff. No one uses it, there are about ~1 headphone with opus support )
- lxgr 3y agoAs you already mention, it's already possible to use it. As for why hardware manufacturers don't actually use it, you can thank beautiful initiatives such as this: https://www.opuspool.com/ https://www.opuspool.com/ (previous HN discussion: https://news.ycombinator.com/item?id=33158475 https://news.ycombinator.com/item?id=33158475).
- giantrobot 3y agoThe BT SIG moves kind of slow and there's a really long tail of devices. Until there's a chip with native Opus support (that's as cheap as ones with AAC etc) you wouldn't get Opus support even if it was in the spec. Realistically for most headphones people actually buy AAC (LC and HE) is more than good enough encoding quality for the audio the headphones can produce. Even if Opus was in the spec and Opus-supporting chips were common there would still be a hojillion Bluetooth devices in the wild that wouldn't support it. It would be cool to have Opus in A2DP but it would take a BT SIG member that was really in love with it to get it in the profile.
- out_of_protocol 3y agoThey chose to make totally new inferior LC3 codec though. Also, on my system (Android phone + BTR5/BTR15 Bluetooth DAC + Sennheiser H600) all options sound realy crappy compared to plain old usb, everything else is the same. LDAC 990kbps is less crappy, by sheer brute force. I suspect it's not only codec but other co-factors as well (like mandatory DSP on phone side)
- giantrobot 3y agoI've got AirPods and a Beats headset so they both support AAC and to my ear sound great. Keep in mind I went to a lot of concerts in my 20s without earplugs so my hearing isn't necessarily the greatest anymore. AFAIK Android's AAC quality isn't that great so aptX and LDAC are the only real high quality options for Android and headphones. It's a shame as a lot of streaming is actually AAC bitstreams and can be passed directly through to headphones with no intermediate lossy re-encode. Like I said though, to get Opus support in A2DP a BT SIG member would really have to be in love with it. Qualcomm and Sony have put forward aptX and LDAC respectively in order to get licensing money on decoders. Since no one is going to get Opus royalties there's not much incentive for anyone to push for its inclusion in A2DP.
- mikae1 3y agoThey’ll have my upvote just for writing ML instead AI. Seriously, this is very exciting developments for audio compression.
- claudiojulio 3y agoMachine Learning is Artificial Intelligence. Just look at Wikipedia: https://en.wikipedia.org/wiki/Artificial_intelligence https://en.wikipedia.org/wiki/Artificial_intelligence
- declaredapple 3y agoMany people are annoyed by the recent influx of calling everything "AI". Machine learning, statistical models, procedural generation, literally an usage of heuristics are all being called "AI" nowadays which obfuscates the "boring" nature in favor of "exciting buzzword" Selecting the quality of a video based on your download speed? That's "AI" now.
- deleted 3y ago[deleted]
- sitzkrieg 3y agoim quite tired of this. every snake oil shop now calls any algorithm "a i" to sound hip and sophisticated
- mikae1 3y ago> Many people are annoyed by the recent influx of calling everything "AI". Yes, that was the reason for my comment. :)
- mook 3y agoOn the other hand, it means that you can assume anything mentioning AI is overhyped and probably isn't as great as they claim. That can be slightly useful at times.
- xcv123 3y ago
- behnamoh 3y agoIsn't it a strange coincidence that this shows up on HN while Claude Opus is also announced today and is on HN front page? I mean, what are the odds of seeing the word "Opus" twice in a day on one internet page?
- declaredapple 3y agoWell it was released today Very likely a coincidence. https://opus-codec.org/release/stable/2024/03/04/libopus-1_5_1.html https://opus-codec.org/release/stable/2024/03/04/libopus-1_5...
- mattnewton 3y agoNot that strange when you consider what “opus” means- product of work, with the connotation of being large and artistically important. It’s Latin, so it’s friendly phonemes to speakers of Romance languages and very scientific-and-important-sounding to English speaking ears. Basically the most generic name you can give your fine “work” in the western world.
- deleted 3y ago[deleted]
- spacechild1 3y agoI'm using Opus as one of the main codecs in my peer-to-peer audio streaming library (https://git.iem.at/cm/aoo/ https://git.iem.at/cm/aoo/ - still alpha), so this is very exciting news! I'll definitely play around with these new ML features!
- RossBencina 3y ago> peer-to-peer audio streaming library Interesting :)
- spacechild1 3y agoHa, now that is what I'd call a suprise! The "group" concept has obviously been influenced by oscgroups. And of course I'm using oscpack :) AOO has already been used successfully in a few art projects. It's also used under hood by SonoBus. The Pd objects are already stable, I hope to finish the C/C++ API and add some code examples soon. Any feedback from your side would of course be very appreciated.
- luplex 3y agoI wonder: did they address common ML ethics questions? Specifically: Are the ML algorithms better/worse on male than on female speech? How about different languages or dialects? Are they specifically tuned for speech at all, or do they also work well for music or birdsong? That said, the examples are impressive and I can't wait for this level of understandability to become standard in my calls.
- radarsat1 3y agoThis is an important question. However, I'd like to point out that similar biases can easily exist for non-ML, hand-tuned algorithms. Even in the latter case test sets and often even "training" and "validation" sets are used for finding good parameters. Any of these can be a source of bias, as can the ears of evaluators making these decisions. It's true that bias questions often come up in ML context because fundamentally these algorithms do not work without data, but _all_ algorithms are designed by people, and _many_ can involve data in setting their parameters. Both of which can be sources of bias. ML is more known for it, I believe, because the _inductive_ biases are less than in traditional algorithms, and therefore are more keen to adopt biases present in the dataset.
- thomastjeffery 3y agoUsually regular algorithms aren't generating data that pretends to be raw data. That's the significant difference here.
- shwaj 3y agoCan you precisely define what you mean by "generating" and "pretends", in such a way that this neural network does both these things, but a conventional modern audio codec doesn't? "Pretends" is a problematic choice of words, because it anthropomorphizes the algorithm. It would be more accurate and less misleading to replace "pretends to be" with "approximates". But then it wouldn't serve your goal of (seeming to) establish a categorical difference between this approach and "regular algorithms", because that's what a regular algorithm does too. I apologize, because the above might sound rude. It's not intended to be.
- yalok 3y agoThe main limitation for such codecs is CPU/battery life - and I like how they sparsely applied ML in it here and there, combining it with classic approach (non-ML algos) to achieve better tradeoff of CPU vs quality. E.g. for better low bitrate support/LACE - "we went for a different approach: start with the tried-and-true postfilter idea and sprinkle just enough DNN magic on top of it." The key was not to feed raw audio samples to the NN - "The audio itself never goes through the DNN. The result is a small and very-low-complexity model (by DNN standards) that can run even on older phones." Looks like the right direction for embedded algos and it seems to be a pretty unexplored one, as compared to the current fashion to do ML E2E.
- indolering 3y agoIt's a really smart application of ML: helping around the edges and not letting the ML algo invent pheonems or even whole words by accident. ML transcription has a similar trade-off of performing better on some benchmarks but also hallucinating results.
- kolinko 3y agoA nice story about Xerox discovering this issue in 2003, when their copiers began slightly changing random numbers in copied documents https://www.theverge.com/2013/8/6/4594482/xerox-copiers-randomly-replacing-numbers-in-documents https://www.theverge.com/2013/8/6/4594482/xerox-copiers-rand...
- h4x0rr 3y agohttps://youtu.be/zXXmhxbQ-hk https://youtu.be/zXXmhxbQ-hk Interesting yet funny CCC video about this
- cedilla 3y agoI don't think machine learning was involved there at all. As I understand it, it was an issue of a specifically implemented feature (reusing a single picture of a glyph for all other instances to save space) being turned on in archive-grade settings, despite the manual stating otherwise.
- rhdunn 3y agoI find the interplay between audio codecs, speech synthesis, and speech recognition fascinating. Advancements in one usually results in advancements in the others.
- brcmthrowaway 3y agoThis is game changing. When will H265 get a DL upgrade?
- h4x0rr 3y agoDoes this new Opus version close the gap to xHE-AAC, which is (was?) superior at lower bitrates?
- AzzyHN 3y agoDepends on whether you're encoding speech or music.
- brnt 3y agoWhat if there was a profiler or setting that helps to reencode existing lossy formats without introducing too many more artifacts? An sizeable collection runs into the issue, if the don't have (easily accessible) lossless masters. I'd be very interested if I could move a variety of mp3s, aacs and vorbis to Opus if I knew additional quality loss was minimal.
- Dwedit 3y agoI just want to mention that getting such good speech quality at 9kbps by using NoLACE is absolutely insane.
- qingcharles 3y agoI was the lead dev for a major music streaming startup in 1999. I was working from home as they didn't have offices yet. My cable connection got cut and my only remaining Internet was 9600bps through my Nokia 9000 serial port. I had to re-encode the whole music catalog at 8000kbps WMA so I could stream it and continue testing all the production code. The quality left a little to be desired...!
- kristopolous 3y agoI wanted to see what it would sound like in comparison to a really early streaming audio codec, realaudio 1.0 $ ffmpeg -i female_ref.wav - acodec real_144 female_ref.ra And if you can't support that I put it back to wav and posted it: http://9ol.es/female_ref-ra.wav http://9ol.es/female_ref-ra.wav This was seen as "14.4" audio, for 14.4kb/s dialup in the mid-90s. The quality increase over those nearly 30 years for what you can get out of what's actually a fewer number of bytes is really impressive.
- anthk 3y agoI used to listen opus avant agarde music from https://dir.xiph.org https://dir.xiph.org at 16kb/s under a 2G connection and it was usable once mplayer/mpv cached back the stream for nearly a minute.
- recursive 3y agoI don't know any of the details but maybe the CPUs of the time would have struggled to stream the decoding.
- deleted 3y ago[deleted]
- WithinReason 3y agoSomeone should add an ML decoder to JPEG
- viraptor 3y agoYou can't do that much on the decoding side (apart from the equivalent of passing the normally decoded result through a low percent img2img ML) But the encoders are already there: https://medium.com/@migel_95925/supercharging-jpeg-with-machine-learning-4471185d885d https://medium.com/@migel_95925/supercharging-jpeg-with-mach... https://compression.ai/ https://compression.ai/
- WithinReason 3y agoYou can more accurately invert the quantisation step
- viraptor 3y agoYou can't do it more accurately. You can make up expected details which aren't encoded in the file. But that's explicitly less accurate.
- adgjlsfhk1 3y agoIf the encoders know what model the decoders will be running, they can improve accuracy. You could pretty easily make a codec that doesn't encode high resolution detail if the decoder NN will interpolate it correctly.
- viraptor 3y agoThat's changing the encoder and sure, you could do that. But that's basically a new version of the format. It's not the JPEG we're using anymore + ML in decoder. It's JPEG-ML on both the encoder and decoder side. And with the speed that we adopt new image formats... That's going to take ages :(
- 3y ago
- aredox 3y ago>That's why most codecs have packet loss concealment (PLC) that can fill in for missing packets with plausible audio that just extrapolates what was being said and avoids leaving a hole in the audio ...How far can ML PLC "hallucinate" audio? A sound , a syllable, a whole word, half a sentence? Can I trust anymore what I hear?
- xyproto 3y agoIt can already fill in all gaps and create all sorts of audio, but it may sound muddy and metallic. Give it a year, and then you can't trust what you hear anymore. Checking sources is a good idea in either case.
- jmvalin 3y agoWhat the PLC does is (vaguely) equivalent to momentarily freezing the image rather than showing a blank screen when packets are lost. If you're in the middle of a vowel, it'll continue the vowel (trying to follow the right energy) for about 100 ms before fading out. It's explicitly designed not to make up anything you didn't say -- for obvious reasons.
- aredox 3y agoReassuring - thanks for clarifying that up.
- samus 3y agoYou never can when lossy compression is involved. It is commonly considered good practice to verify that the communication partner understood what was said, e.g., by restating, summarizing, asking for clarification, follow-up questions etc.
- m3kw9 3y agoSome people hyping it as AGI on social media
- nimish 3y agoThat 90% loss demo is bonkers. Completely comprehensible after maybe a second.
- cedilla 3y agoThe quality at 80% package loss is incredible. It's straining to listen to but still understandable.
- frumiousirc 3y agoHow about adding a text "subtitle" stream to the mix. The encoder may use ML to perform speech-to-text. The decoder may then use the text, along with the audio surrounding the audio drop outs, to feed a conditional text-to-speech DNN. This way the network does not have to learn the harder problem of blindly interpolating across the drop outs from just the audio. The text stream is low bitrate so it may have substantial redundancy in order to increase the likelihood that any given (text) message is received.
- jmvalin 3y agoActually, what we're doing from DRED isn't that far from what you're suggesting. The difference is that we keep more information about the voice/intonation and we don't need the latency that would otherwise be added by an ASR. In the end, the output is still synthesized from higher-level, efficiently compressed information.
- Sonic656 3y agoLove how Opus 1.5 is now actually transparent at 16kbps for voice and 96kbps is still beats 192kbps MP3. Meanwhile xHE-AAC still feels like It was farted out since It 96 ~ 256kbps area Is legit worse than AAC-LC(Apple, FDK) are at ~160kbps.