21 ms·
AI produces realistic sounds that fool humans
- macawfish 10y agoThey trained the algorithm to watch the stick and play sounds from the database where the stick moved similarly. But the title makes it seem like the algorithm is synthesizing the sounds from scratch!
- a1k0n 10y agoThey do both; there's a parametric synthesis module later on in the video. It doesn't work all that well for water.
- jobigoud 10y agoThey also train for the material the stick is hitting. But yeah, it's sound transfer, not synthesis from scratch. But that's also the approach of current speech synthesis algorithms and works better than trying to create the waveform from scratch.
- mwcampbell 10y ago> But that's also the approach of current speech synthesis algorithms and works better than trying to create the waveform from scratch. I don't think it's that simple. Speech synthesis by concatenation does produce more natural-sounding results, at least until you notice its quirks, so casual users tend to prefer it. But I know some heavy speech synthesis users, specifically blind programmers and power-users, and they tend to prefer parametric synthesis, because it's more intelligible at high speeds.
- 542458 10y agoThey can do pure parametric synthesis as well, but it's not nearly as convincing so most of the video is devoted to the more convincing match method. FWIW, constructing realistic sounds from first principles is much more difficult than you'd think. > where the stick moved similarly where the stick moved similarly and was hitting similar things, which is a non-trivial task.
- tuewocnc 10y agoyes, it would have to learn to simulate the physics of the system to match the video, which would be cool
- jordache 10y agoYes the title may give that impression, but the video explains clearly the source of the audio
- c3534l 10y agoAnd for what it is, it's really unimpressive. It's a cool idea and all, just with disappointing results.
- colordrops 10y agoExactly, most of the work here is video processing - analyzing a portion of one video, then finding another video clip that has similar characteristics. Then they copy the sound from the second clip into the first clip. But they could be copying any metadata, not just sound. This isn't really about sound _at all_.
- visarga 10y ago> But the title makes it seem like the algorithm is synthesizing the sounds from scratch! It does. If you read the paper, they say that first they went with matching sounds from a database, but later turned on to full synthesis. Reference: look for "parametric synthesis" in the paper https://arxiv.org/pdf/1512.08512v2.pdf https://arxiv.org/pdf/1512.08512v2.pdf
- slr555 10y agoAI that produces sound through analysis of a source video is impressive. Fooling humans is not. Since most of us have grown up on a steady diet of film and television many of the sounds we have in our memories are the work of foley artists that add sound effects to sequences in post. The sound of horse hoofs on cobblestones is likely created from a percussive technique that has no equine participation. The sounds of people being punched may be the sound of a large piece of meat being struck with a club. Similarly crunching snow likely is not the sound of a person walking through actual snow. Our perceptions of sound within a video/film source is already deeply skewed and therefore the notion that this AI is a Turing test of sorts is a weak analogy.
- kdeldycke 10y agoThe Turing test in this case might consist of feeding the algorithm with a mute video of a scene from Monthy Python and the Holy Grail, when coconut halves are used to simulate the sound of galloping horses.
- cleeus 10y agoYes, and often sounds from huge sound libraries are used. I allways smile when I hear a sound from Unreal Tournament 4.36 in a movie or on TV.
- Karunamon 10y agoOr Doom.. I remember very clearly the intro to Modern Marvels on the History Channel uses the Doom door opening noise. I guess they used the same SFX library or something?
- bitwize 10y agoThe library used by Doom is REALLY common. The sound DSBOSPIT (sound of boss demon spitting a telecube in Doom 2) in particular is so overused it's not even funny. You hear it everywhere: in budget movies when a house or plane explodes, sometimes in other video games.
- 10y ago
- mpitt 10y agoIt fools humans more often than a baseline algorithm. What that means in numbers, they carefully avoid saying.
- danielmorozoff 10y agoThis work looks very cool from the demo video. Have yet to read the paper, but the parametric inversion for generating sounds from features seems very intriguing.
- vernie 10y agoSample size: 3
- whatever_dude 10y agoImportant little side note.
- ThomPete 10y agoI think people here are underestimating how big a thing this actually is and I think the headline is kind of to blame for that. This is much less about fooling humans than about what this actually means. Humans spend a huge part of our early live learning to listen and to connect the dots between what we see and what we hear. The fact that Deep Learning algos now can simulate audio based on what they see thats the big thing here. Not the production of the sound that fools humans. You can almost sense how imagination and inspiration is inside the reach of machine learning (yes there are some way to go yet) We are now not only seeing individual senses being simulated but also the relationship between them. And as a bonus what one machine learns one place can be instantly added to the knowledge of the other. Thats IMO the big deal here.
- YeGoblynQueenne 10y ago>> And as a bonus what one machine learns one place can be instantly added to the knowledge of the other. That's actually one big problem with machine learning algorithms: it's not at all clear how to integrate their knowledge with that of other algorithms (and that includes different instances of the same algorithm). Such algorithms build a single model of one domain at a time, and we're talking about very strict domains. What we're seeing lately is many teams announcing that they trained an algorithm to do this or that pretty damn amazing thing, but watch closely: how many of those announcements describe a system that can integrate its learning into a wider cognitive architecture? There's teams that trained models to recognise images, to combine images, or to map images to strings, but all these things are simple tasks, that are only useful in a very limited range of circumstances. Machine learning algorithms unfortunately are one trick ponies. They do one thing well- and that's it. >> You can almost sense how imagination and inspiration is inside the reach of machine learning (yes there are some way to go yet). That's an understatement- the bit about having some way to go. We're not even close, really. To train a machine learning algorithm the first thing you need is a lot of examples of the thing you want it to learn. It's really hard to see how one would compile a set of examples of imagination, not least because it's inside peoples' heads. Not to mention we don't even know what human imagination is in the first place.
- 10y ago
- jderick 10y agoPaper is here: http://arxiv.org/abs/1512.08512 http://arxiv.org/abs/1512.08512
- EGreg 10y agoWe've been fooling humans for years because the humans in question were conditioned by tropes on TV: http://tvtropes.org/pmwiki/pmwiki.php/Main/TheCoconutEffect http://tvtropes.org/pmwiki/pmwiki.php/Main/TheCoconutEffect
- kelvin0 10y agoThis has the potential to do the same for Videogames as did Mo-Cap.
- anotheryou 10y agoFooling humans is much easier when about 50% of the sound is always accurate (in any case it's a wooden drumstick hitting something). The human mind is very forgiving, especially when vision accompanies sound (McGurk effect [1]) Furthermore this fixed variable makes the pool of samples to choose from much much smaller. Still impressive of course ;) just far from dubbing any video there is. [1] https://youtu.be/G-lN8vWm3m0?t=32 https://youtu.be/G-lN8vWm3m0?t=32
- bekimdisha 10y agoWhy is AI focused so much into fooling/imitating/trap-ing humans these days?
- tmalsburg2 10y agoI don't think these systems' purpose is to fool humans. Tasks that test whether a system can fool people are simply a good way to evaluate the performance a system. If a speech synthesis fools people into thinking a real person is speaking, that means the speech synthesis is really good. You might say it's not important that a speech synthesis sounds perfectly human but our speech perception evolved to be optimal for human speech, so it's likely that any deviation from that makes the signal harder to process.
- Scarblac 10y agoBecause with deep learning we quite recently got a new tool (someone figured out how to use GPUs for training) that lets of do a lot of those imitating things we couldn't do before. A lot of fruit that suddenly became low hanging.
- logicallee 10y agoGood research, but some parts of the videos are like badly dubbed sound effects - actually hilarious: https://www.youtube.com/watch?v=0FW99AQmMc8&t=1m1s https://www.youtube.com/watch?v=0FW99AQmMc8&t=1m1s (the drumstick noise in the middle). still impressive.
- andreyk 10y ago"The first step to training a sound-producing algorithm is to give it sounds to study. Over several months, the researchers recorded roughly 1,000 videos of an estimated 46,000 sounds that represent various objects being hit, scraped, and prodded with a drumstick. (They used a drumstick because it provided a consistent way to produce a sound.)" Whilst the production of natural seeming sound is cool, that quote right there perfectly shows just how limited AI/ML still is. Sure, Deep Learning systems can be taught to do perception tasks (such as understanding or creating sounds and images) very very well, but those perception tasks are incredibly specific and narrow. Not only that, but they are trained large datasets hand-labels through an intensely laborious process, and indeed this laborious process is necessary because we are still using simplistic supervised learning. At this point good recognition or generation with deep learning is entirely old news, and I think zero shot or unsupervised/semi-supervised learning is where the real challenges still are.
- YeGoblynQueenne 10y ago>> They used a drumstick because it provided a consistent way to produce a sound. More to the point, that bit. The article makes a big todo about how humans use sound to learn about their environment and so on, but imagine if we needed to get a drumstick to make sounds consistent enough to learn to recognise them. Supervised and unsupervised learning is not the real challenge. The real challenge is to get to the point where algorithms can build a model without the benefit of an insanely expensive data pre-processing pipeline. Deep learning's big promise is exactly that, but it's not always delivered (for instance, there's a paper by Hinton and I forget who else, where they report that training LSTM RNNs on raw characters does not give best performance) (so we're stuck with tokenisation and the implicit assumptions they impose on your data- my corollary). Also, there are ways to avoid the expenses of hand-labelling, for instance co-training: https://en.wikipedia.org/wiki/Co-training https://en.wikipedia.org/wiki/Co-training
- andreyk 10y agoTrue, that bit is especially bad. "The real challenge is to get to the point where algorithms can build a model without the benefit of an insanely expensive data pre-processing pipeline." Arguably, that is equivalent to saying the real problem is still unsupervised/semi-supervised learning. IE, being able to just throw a bunch of raw data and maybe just a bit of hand configuration at an algorithm and have it do complicated things for you. The success of Deep learning is to scale to tons of data and build really complicated models, but as it is used today that data is still hand labeled for supervised learning in an insanely expensive data pre-processing step. Good unsupervised or semi-supervised learning could hopefully let us get out of this, but I don't think anyone really knows how to get there yet. Co-training is an older example of semi-supervised learning, and more recently there were Ladder Networks, but I don't think any algorithm has been shown to work really well and become the norm in the way LSTM RNNs or CNNs have.
- _mhr_ 10y agoIt would be cool to create the drums for a song by taking a video of the performance and then using this software to create the recording to be put into the song.
- tmalsburg2 10y agoJust a couple of years ago no one would have called this AI. Interesting how this old term has become so fashionable again. Perhaps it's also being overused and AI winter is coming.
- Animats 10y agoThe demo seems to recognize three basic categories of things hit - shrubbery, dirt, and other solid objects. It doesn't distinguish much between hitting metal, wood, and asphalt. It's another step toward common sense. Predicting what will happen if a robot does something in the real world is essential to making robots less stupid.
- Lxr 10y agoNext we need to do it in reverse - given a sound, generate a video to match. I wonder how well that would work?
- dingle_thunk 10y agoMuch of this is beyond me, but is it producing (synthesizing) sounds? Or just sampling sounds?
- systoll 10y agoMost of the video has them using an algorithm to predict the sound of the hits, then finding the closest match in the database & using it as a sample. At 1:50, they switch to using the prediction as a synthesised replacement. It is not as good. They switch back to the first method for the tests
- tednoob 10y agoSince the discussion seem to have died down a bit I just have to say it, sorry. Quit beating around the bush.
- ourcat 10y agoPercussionists beware. Your obsolescence clock is ticking.
- grondilu 10y agoFunny coincidence that I just stumbled upon a nice video about sound crafting for cinema: http://sploid.gizmodo.com/wonderful-short-film-reveals-the-painstaking-magic-of-f-1782096219 http://sploid.gizmodo.com/wonderful-short-film-reveals-the-p...