17 ms·
Deep Learning for Guitar Effect Emulation
- 317070 6y agoThat is very cool. Though, part of the pedal are of course the knobs. You'd need to condition the wavenet on the knobs. Did that work well (I assume that you tried that already)? Also, what is the inference latency on your model? A nice thing about analog guitar effects is that they are blazingly fast.
- fab1an 6y agoPretty cool, though I wonder what the latency of this would be if used as a plugin? The author says it works in real-time, but to non music/audio folks this could mean '100 ms latency is real-time enough, right?' Generally, I think the audio VST business is a really fun space to be in for a lifestyle business, as it is way too small to be attractive for VCs. It seems like a space that provides many niches for lots of small players to thrive in. As an aside, it's really quite interesting that a lot of cutting edge tech is now used to emulate the hardware-based tech of yesteryear. Think film filters for photoshop, and about 90% of all audio plugins that emulate high end hardware, compressors, pedals, etc etc.
- eru 6y agoReal-time has a few slightly different meanings. So it's hard to say what the author means. One meaning is just that you can guarantee specific deadlines. So if your programme can react within an hour guaranteed, that would be real-time. (Though usually we are talking about tighter deadlines, like what's needed to make ABS brakes work.) For 'real time' music usage you wouldn't need strict guarantees, but something that's usually fast enough.
- aea12 6y agoImplementing a VST plugin is literally the exact definition of requiring strict latency guarantees. Your comment winds through a lot of unrelated comparisons to ultimately not make any sense. “Usually fast enough” are three words that guarantee failure in a live show/MIDI environment, which is a large use case of VST and its peers beyond production. By extension, “usually fast enough” further guarantees nobody will ever use your software. That’s noticeable right away. The question isn’t about compsci real-time theorycrafting, it’s “here’s a buffer of samples, if you don’t give it back in a dozen milliseconds the entire show collapses.” That’s pretty clearly meant by “real time“ contextually.
- microcolonel 6y agoNot to mention if the inference is done on the CPU, it shouldn't be that hard to control it. The matrices are of a set size by the time you're running a VST; this is the actual simple answer. The medium answer is "this is a wavenet model, so inference is probably really expensive unless the continuous output is a huge improvement to performance".
- mochomocha 6y agoIndeed. Having myself spent some time in the "VST lifestyle business" when I was in grad school (was selling a guitar emulation based on physical modelling synthesis), and now working in ML, I think there's no chance for such an approach to hit "mainstream" anytime soon. Even if you do your inference on CPU, most deep learning libraries are designed for throughput, not latency. In a VST plugin environment, you're also only one of the many components requiring computation, so your computational requirements better be low.
- jerf 6y agoYou might be able to combine it with the recent work on minimizing models to obtain something that is small enough to run reliably in real time. Although the unusual structure of the net here may mean you're doing original and possibly publication-level work to adapt that stuff to this net structure. If you were really interested in this, there could also be some profit in minimizing the model and then figuring out how to replicate it in a non-neural net way. Direct study of the resulting net may be profitable. (I'm not in the ML field. I haven't seen anyone report this but I may just not be seeing it. But I'd be intrigued to see the result of running the size reduction on the net, running training on that network, then seeing if maybe you can reduce the resulting network again, then training that, and iterating until you either stop getting reduced sizes or the quality degrades too far. I've also wondered if there is something you could do to a net to encourage it not to have redundancies in it... although in this case the structure itself may do that job.)
- microcolonel 6y agoI wonder if teddykoker has looked at applying FFTNet or similar methods as a replacement for Wavenet. I'm not sure but it seems to me like FFTNet is a lot more tractable than Wavenet, and not necessarily that much worse for equivalent training data.
- alexlarsson 6y agoI assume by real-time he meant "able to produce samples at a rate equal to or higher than the audio output sample rate".
- amelius 6y agoLatency is equally important.
- TheRealPomax 6y agoThat's not what real time, means though. Real time processing means taking signals as they come in, and outputting the transformed result such that there is as close to no signal lag as possible. The output can in fact be wildly lower or higher resolution, real-time does not particularly say anything about that. It's all about whether the output plays (for practical purposes) at the perceived "same time" as the input signal. There will always be some delay, but that delay can't get perceivable, and for obvious reasons there can't be any (significant) buffering.
- jeffbee 6y agoIs that your private definition of "real-time"? I think it is common to define real-time processing by a specified, finite time between input and output. Many real-time processes are concerned more with the consistency of the latency than with its absolute value.
- avisser 6y agoFor guitar pedals, there is an implied sub-perceptibility. The output needs to happen as I play - if the delay is too long, it's now a delay pedal. So realtime might match your definition, but it is consistent in audio production. For humans, you can start to notice the lag @ 50ms. (A selection of experimental results summarized here https://gamedev.stackexchange.com/a/74975 https://gamedev.stackexchange.com/a/74975)
- NobodyNada 6y ago
- whiddershins 6y agoDo solo or small shop vst plugin developers make any money? I’m curious if anyone has any direct knowledge about that. There are so many professional activities similar to that where no one makes any money and people really just do it for the love, and then there are seemingly similar things like that where people make surprisingly large amounts of money.
- ff7f00 6y agoThere are definitely big players making a lot of money from plugins they develop. Here are a few to check out: ($1200) https://www.native-instruments.com/en/products/komplete/bundles/komplete-12-ultimate/ https://www.native-instruments.com/en/products/komplete/bund... ($300) https://www.soundtoys.com/product/soundtoys-5/ https://www.soundtoys.com/product/soundtoys-5/ ($500) https://www.arturia.com/products/analog-classics/v-collection/overview https://www.arturia.com/products/analog-classics/v-collectio... However, piracy is also pretty big when it comes to plugins.
- moralestapia 6y agoSure, but, >Do solo or small shop vst plugin developers make any money?
- kleer001 6y agoThe implication seems to be "no".
- TheRealPomax 6y agoThat implication would be wildly incorrect.
- kleer001 6y agoWith the operative word being "developers" I disagree.
- deleted 6y ago[deleted]
- samplenoise 6y agoThere's latency and there's the somewhat separate question of how much time is needed to make a prediction. Wavenet is causal (no look-ahead) and operates on the sample level so there are no buffers and thus no latency in the strict sense, beyond encoding/decoding into the sample rate and format required by the ML model, which should take <1ms. Whether a model manages to make a prediction in that amount of time depends on things like the receptive field and number of layers. The linked paper says their custom implementation runs at 1.1x real-time. I guess this isn't impossible; their receptive field is ~40ms, vs. 300 for the original (notoriously slow) wavenet, and the model is likely to have less layers and channels.
- cjlars 6y ago"Round trip," or guitar to processing to speakers needs to be sub 10ms to be transparent to the musician. Source: spent years playing guitar through my guitar -> DAC -> PC -> DAC -> speaker signal chain
- samplenoise 6y agoThe receptive field size is how much 'history' the algorithm requires, it doesn't affect the round trip time, which can still be sub ms
- TheRealPomax 6y agowhile training? terrible. As finalised model running in an AU/VST3 wrapper? probably extremely low.
- qppo 6y agoI know of a few shops that took VC money. The big problem isn't the market size so much as how slow the market moves. The product lifetime of a plugin is around a decade. And users hate subscriptions. And it's really hard to determine the value you add to your customers. And no one wants to pay you. It's basically a terrible place to be a developer in it for the money. Really fun work otherwise. The cool gigs are the ones where you build custom plugins for someone's crazy idea. In consumer applications, plugins are used all the time for prototyping before you go to hardware. MATLAB is way too slow for anything useful.
- nil-sec 6y agoThe success of splice would disagree with your notion that “users hate subscriptions”. Given the horrendous price point of many of these plugins it seems to be perfect for a subscription based model. To me it always seemed there is more of a pushback from the industry producing vsts than from the consumers.
- qppo 6y agoSplice's numbers aren't public so I can't comment on their success. Avid's are, and they had a terrible quarter - and they're the poster child (alongside Adobe) for subscription licensing in creative software. But I'd be interested to see what the breakdown in revenue is for plugin licenses versus preset/sample packs (bit of a blade & razor model there). The price points really aren't horrendous if you consider how expensive the engineering is, how little demand there is, and how long you need to maintain a product. You aren't being ripped off by spending a couple hundred bucks on a plugin. I think we'll end up at a place where everything is a subscription, but I can tell you from experience that it creates friction for the users.
- nil-sec 6y agoAgreed. The business model seems to be to give access to the rent to own deals via the sample subscription fee. Don’t think they make any money of their plugin deals. I’m also not arguing it’s too expensive or a rip off. But it’s still a large amount of money for software, in the private space at least. The rent to own thing seems like a smart tool to get rid of the barrier of entry.
- exabrial 6y agoPretty cool! Is this how Kemper amplifiers work when they do a capture?
- ratww 6y agoAFAIK Kemper performs multiple passes of impulse-response capture, all at multiple signal levels in order to model non-linearities (like distortion). This is called dynamic convolution. [1] [2] There are other ways to do that, like Volterra Series, used by Nebula plugins [3] [1] https://www.uaudio.com/webzine/2004/july/text/content2.html https://www.uaudio.com/webzine/2004/july/text/content2.html [2] http://www.sintefex.com/docs/appnotes/dynaconv.PDF http://www.sintefex.com/docs/appnotes/dynaconv.PDF [3] https://en.wikipedia.org/wiki/Volterra_series https://en.wikipedia.org/wiki/Volterra_series
- deleted 6y ago[deleted]
- svantana 6y agoEnd-to-end modelling is very enticing for the lazy engineer, unfortunately parameter control (knobs) are an important feature of most audio effects, and sampling enough of the parameter space will become prohibitive for more complex effects. That's why the traditional approach is divide-and-conquer. Also, I don't think this approach won't work well with time-varying effects such as chorus, although I'm happy to be proven wrong.
- fxtentacle 6y agoSadly, I believe you will be proven correct. What that neural network learns is basically an approximation of a static impulse response. So while it can simulate linear time-invariant effects such as reverb quite nicely, it'll surely have issues with chorus.
- sk0g 6y agoReverb is time invariant? You can set custom decay time, rate etc, so the one not can be heard for say, 10 seconds if you want to go full Devin Townsend. I'd think Chorus would work better. I wanted to do a very similar project, but with an overdrive. Let's see if I get time anytime soon!
- vnorilo 6y agoReverb is indeed linear time invariant (sans some rarer internal modulation techniques) but it's quite a high order filter.
- sk0g 6y agoAh righto, the reverb pedal I'm most familiar with turns out to not be just reverb - EQD Afterneath does a whole bunch of funky stuff. Plain reverb though, yeah. I was approaching this more from the angle of training a neural network, where the input and output waves have to be correlated over a great span of time/ samples.
- InitialLastName 6y ago
- ericfrederich 6y agoSo this seems similar to an IR (impulse response) where you get a snapshot of an amp mic'd up in a room with knobs fixed at a particular position. In the end, you don't get knobs to fiddle with. Awesome, I'd love to hear Josh from JHS Pedal's opinion on this.
- ratww 6y agoThis is even more impressive since regular IRs can't duplicate the distortion effect itself, only the frequency response
- munificent 6y agoWhat is the difference between "distortion itself" and "only the frequency response"? Are you saying the phase response is important?
- ratww 6y agoImpulse responses can only represent linear time-invariant systems. Like delays, reverbs, equalization curves. Distortion is non-linear, it is something like a max(-1, min(1, input)) function (a waveshaper, like you said), and it produces harmonics when applied to audio signals. However guitar pedals also have some additional circuitry to "sweeten" the distortion, removing the extra harmonics added by the clipping diodes. Tubescreamers are notable for cutting bass and enhancing mids. An IR is able to capture this. This is important for guitar pedals, and the reason multiple of them exist. If you capture the impulse response of an overdrive pedal you'll be capturing only the frequency response of a distorted impulse. If you process clean guitar trough this you'll simulate the frequency response but not the distortion itself, so it will just be a clean guitar with a tinny, shrill, sound, not an overdriven guitar sound. One way around it (other than the idea in this article!) is doing multiple passes of Impulse Response capture with different amplitudes, this will capture this distortion non-linearity. This is supposedly how a Kemper Profiler works.
- jelling 6y ago> We find that the model is able to reproduce a sound nearly indistinguishable from the real analog pedal. Maybe for the average person or buried in the mix, but the audio samples were easy to distinguish for me as a guitarist. The NN samples unnatural decay were a dead give away.
- Tade0 6y agoAlso the pretty obvious quantization noise which sounds as if the effect had a wide bandwidth, which is impossible with their op-amps at these gains.
- sailfast 6y agoYeah - this was clearly audible on my phone speakers, especially during more muddy / multi-note sequences. While it may not be able to emulate a real pedal to create one’s own sound, it would be interesting / fun for amateurs when applied as a post-filter with an interface that says “make this sound like X famous incredible track” coming out of a stock guitar signal.
- hashkb 6y agoConfirmation bias overrides ear training. Always have an unbiased tone junkie do your blind test.
- magicalhippo 6y agoEven as a regular Joe it was easy for me to distinguish them, and though I was not very confident in my guess, I did guess correctly as well. It was close though, so maybe for say a beginner on a shoe-string budget it would be perfectly acceptable.
- finder83 6y agoAgreed, the NN had that "digital" sound you typically get from a simulated tube screamer, such as in a POD HD or something. Very impressive given it's from a NN, but I specifically moved to analog for that reason.
- zwieback 6y agoYeah, real pedal sounds much "better" but maybe we're just used to how they sound.
- baylessj 6y agoExcellent writeup, I love seeing real engineering applied to guitar pedals rather than black magic tone chasing. I'd be really curious to see if the model could be expressed as a transfer function and compared to the schematic for the pedal. The Tubescreamer is a fairly simple circuit but the mystery surrounding it indicates that there are some weird variables at play with the component properties that would lead to additional factors in the transfer function. Wonder if those variables could be identified somehow.
- hashkb 6y agoThe "weird variables" may have to do with the various changes in manufacturing over the years. "Tube screamer" refers to at least 10 different units. Maxon, Ibanez, TS9, TS808, and zillions of clones.
- Tade0 6y agoSounds great and I had to listen to both of the samples to guess correctly. That being said the Tube Screamer is a somewhat simple effect: it's just a distortion with the clipping diodes moved to the feedback loop. How possible would it be to get the famous A/B class amplifier voltage sag and associated changes in parameters of the whole amplifier, or in other words "will it chug"?
- cesaref 6y agoI think this would be very possible - there was quite a bit of discussion of using NN techniques for modelling fx discussed at DAFx2019 (http://dafx2019.bcu.ac.uk/ http://dafx2019.bcu.ac.uk/). There are a number of papers discussing different techniques in the paper archive. Many of the techniques discussed were variations on image processing - transforming the input to the frequency domain then converting this to an image, and applying standard techniques to transform the image, then back to the time domain. There are many compromises with this approach (loosing phase information for example) but with a suitable overlap/add the results were better than I expected, and certainly there's room for further investigation to see if there's useful stuff in there. Another time domain approach that was applicable to your amplifier model question was an attempt to determine hidden variables in a circuit. Basically, the circuit under test is examined, and rather that build a spice model (which can be laborious) the technique was to expose the interval voltages following components with memory (so capacitors for example). These outputs were included in the NN training model, and so in effect the normally hidden internal state was exposed and allowed for a very good approximation. Here's the paper: http://dafx2019.bcu.ac.uk/papers/DAFx2019_paper_42.pdf http://dafx2019.bcu.ac.uk/papers/DAFx2019_paper_42.pdf
- Tade0 6y agoThank you very much. Do you know if there will be a DAFx2020? That would make it the first conference in years that I would really want to attend.
- cesaref 6y agoUnfortunately not, it's been delayed. DAFx2020 was due to be in Vienna, and i'm assuming they are still planning on being there, but it's scheduled to be in 2021. It's a great conference, well worth attending. It's heavy on the maths, but that's DSP for you!
- hashkb 6y agoTrey Anastasio of Phish famously uses 2 stacked tube screamers. (And so do many of us phans). He deserves to be mentioned because more notes have hit audience ears through his screamers than anyone else's. Also, the modern TS9 isn't exactly right. I'd love to see this work applied to vintage vs current TS vs modded units.
- ZoomZoomZoom 6y agoFor anyone planning to try this, don't forget about impedance matching and use a transformer/active reamper. Some pedals may react very differently.
- willis936 6y agoA neat approach for sure. I am more interested in SPICE style modeled VSTs though. There's no need to throw ML at a simple math problem to get a bad approximation. I have not found many VSTs that seem like they're doing proper simulation of analog circuits. The VST space is filled with people claiming awesome results, but never revealing the sauce. If you're making a convincing sounding zener limiter, what are you actually doing? There are a dozen different levels of approximations you could make. Shouldn't a VST that is really simulating the analog circuit advertise that? On paper it should be easy, right? I've sat down with pen and paper to try to write out a simple input/output equation for a zener limiter circuit and I decided it was probably more worth my time to just plop a zener SPICE model into some language that could evaluate expressions and compile to VST (or use a systems of equations solver). And then there's the real holy grail of analog simulation: the tube amplifier. I'm not sure SPICE models really capture the limiting behavior of tubes very well. You might need to implement the spec sheet in code. All fun sounding problems, and I'm not sure anyone has even done them yet.
- ben7799 6y agoRight..the Spice modeled version has a much better chance of catching the oddball behavior of guitar effects across the wide span of possible inputs.
- dsharlet 6y agoFunny you mention SPICE to VST compilation... It was on my list for this (my) side project but I never got around to it: http://livespice.org/ http://livespice.org/ edit: And a Tubescreamer is one of the examples!
- TrackerFF 6y agoIsn't this essentially just learning the case of learning one function, with set parameters? I.e, if you want to build a complete model of the tubescreamer, you'd essentially have to train a model for each possible setting on the pedal - or in other words, every combination of the knobs. Sounds like a real chore, if you were to actually do that physically - and in the end, don't you just want to learn the impulse response of the circuit? I know some tools - like the Kemper modelling gear, are made for that exact purpose, and with extremely convincing results.
- Scene_Cast2 6y agoNot quite. As long as the knobs make consistent changes, just feed some large amount of tests and the model should generalize (smartly interpolate) the rest. What I do have a problem with is that if the pedal is already implemented digitally, then all the human interpretability, along with the classic DSP machinery, is thrown out the window. A better approach would be to build the pedal via a differentiable programming language and then try to gradient descent toward some analog "can't get this juicy tube sound digitally" variant.
- ben7799 6y agoThe knobs actually don't behave linearly on a tube screamer. Even the "tone" knob (EQ) doesn't behave at all linearly like you might expect out of consumer audio gear. Tube Screamers have an S-curve potentiometer in use for that knob. That would be part of the problem with this approach. Also with this approach you pretty much have to train the model with a near infinite collection of guitars in front of the model and a near infinite number of other effects turned on and off in front of the model.
- Scene_Cast2 6y agoThe knobs don't have to be linear at all, just differentiable - that's the beauty of ML. As for the collection of guitars and samples - not necessarily, it would depend on how you set up the training.
- saadalem 6y agoThis is actually impressive, I'm wondering if we could transfer the smartphone mic to a high quality one with AI
- wintermutestwin 6y ago"many purists argue that the sound of analog pedals can not be replaced by their digital counterparts." Truly effective modelling of analog pedals, tube amps and guitar cabs has been around for years and is way more cost effective from the bedroom to touring bands. The "purists" are hipsters who value the rarity of some pedals, massive pedalboards and their tube amps. I'm not knocking them - I understand why there is a nostalgia factor and tweaking dials is cool. As a computer guy though, I much prefer the ability to make things like this in my bedroom: https://i.imgur.com/OqMoBxz.png https://i.imgur.com/OqMoBxz.png And when I want to tweak a dial, I program an expression foot controller to tweak any parameter (or multiple). All that said, great to be looking at modelling techniques...
- maeln 6y agoA lot of band stopped using analogue hardware for sound also because they tend to be way less reliable than their digital counter part. A lot of analog amp, pedals and synth will tend to change their sound due to the analogue hardware aging. Digital stay virtually the same. And the same can be said about weather condition. Change in temperature and humidity affect analogue hardware, not so much digital. You will have the same sound from gig to gig and a lot of band really value this.
- selykg 6y agoI am by no means a musician or an experienced one at that. I tinker and enjoy playing and learning. But I have limited experience overall. My personal experience with electronic tools is the lack of feel. Can I make music with digital tools like AxeFX and similar? Absofreakinglutely. No doubt about it. But those digital tools feel VERY different to me than the real thing. I'm not just talking about a speaker moving air, though that's certainly part of it. My tube amp simply responds differently than any digital model of a similar amp. I find tools like the Kemper to be amazing, but they're just a snapshot of an amp in a particular configuration in a particular room. From a technical standpoint, all this modeling stuff is super cool. But it doesn't feel the same at the end of the day and this is a personal opinion and preference on my part. I look forward to the day that I can get an amp in a pedal (like the Strymon Iridium) and it behaves the same as the real amp. I think Fender's Deluxe Reverb (Tonemaster model) is as close as it has ever gotten, but it very specifically emulates a single amp and does so within a real amp cabinet rather than pushing it out to an audio interface. Anyway, anything that gets people playing guitar is, in my opinion, a great thing. We live in a golden age of guitar equipment. I don't think it can honestly get much better than it is right now. It's an amazing time to be a guitar player and incredible options are available at amazing prices.
- ben7799 6y agoI play guitar and own a tube amp & a tube screamer. All of this sounds horrible.. it doesn't even sound like his input is an actual guitar, it sounds like he's using a synth guitar sound or something. There's no dynamics, almost no sustain, no articulations. The outputs barely even sound distinguishable as a guitar through a tube screamer, even his actual tube screamer samples. (Possibly cause his interface is terrible?) The conclusion is ridiculous given how simplistic everything is. You can't use two tiny little clips to justify your model being high quality. The true test has to even allow a bunch of guitarists to move all the knobs, plug the model into different amp & guitar combinations, put other effects in front of and behind it, etc.. The Tube screamer is called a Tube screamer because it's intended use case is to make the tubes in a tube amp "scream". Using it with all the knobs at noon is not consistent with this, it usually gets used with a tube amp that is already on the verge of distortion, and then you use the TS with the volume turned up a lot (3/4-max) and the gain quite low, this might be part of why this sounds so bad to me. There are actually two different trains of thought on guitar effect modeling: - Model it based on input & output waveforms like he's doing - Actually model the circuit as an electrical simulation and then pass the signal through that. I have personally found the second approach to be way more realistic and satisfying. The Yamaha THR amps work this way and they're really amazing. One of the tricks here is a listener might not be able to tell a difference, but the guitar player picks up on a perceived change in how the guitar feels with these effects. A tube screamer has a lot of compression built into it for example. It causes everything to play to sound a little dirtier for the same amount of picking energy you put into the guitar. It will cause the player to play a little more lightly than they would without the effect. This is the kind of thing that makes a player reject the model and want to stick with the real thing, whereas the guy in the naive lab building the model thinks it's great cause they're not even playing an actual guitar through it. Once a skilled player tries it the "feel" is a dead giveaway which is which. It's easy for some of this stuff to get lost on the electronics crowd if the background is electronic music. An actual acoustic piano is the only keyboard based instrument that has anywhere near the nuance that a guitar has, and a guitar still has way more weird stuff going on with dynamics and articulation. The range of inputs you have to feed into any kind of computer model to simulate guitar well is huge.
- ateamtexas 6y agohttps://ateam-texas.com/things-to-take-care-while-outsourcing-mvp-development/ https://ateam-texas.com/things-to-take-care-while-outsourcin...
- mgamache 6y agoIt would be interesting to see how this responds to dynamics. For example, a favorite guitar sound is a fuzz cranked, but with the guitar volume turned down. This results in a compressed dirty sound that can overdrive into distortion if you hit the strings harder (attack).
- munificent 6y agoI'm not an expert on machine learning or DSP, but I do know just enough of each to suspect this isn't anywhere near as impressive as it seems. A distortion pedal is essentially just a waveshaper [1]. Think of audio in digital terms as just a series of numbers. A waveshaper is just a simple mathematical function. To apply it, you literally just apply the function to each value in the input stream and there's your output stream. There's no memory or interesting algorithms going on. It's the audio equivalent to calling map() on your list of samples with some lambda to produce a new list of samples. Of course distortion pedals do that in the analogue domain using circuitry, which has some additional complexity because transistors and diodes and friends don't behave exactly like mathematical functions. There's "sag" and some other physical effects that cause the output to also somewhat depend on previous input. Even so, that can generally be modelled using a simple convolution. Each output sample is calculated by taking some finite number of previous input samples, multiplying each of them by a weight factor, and then summing the results. Does that sound like a neural net? It is. That's what we call them convolutional neural networks. Convolution is bread and butter in DSP. You can easily generate one that produces the same effect as some piece of hardware or acoustic environment by running an impulse (a single 1.0 sample surrounded by silence) through the system and then recording the result. That "impulse response" essentially is your set of convolution weights. So using a deep neural network and then training sounds a lot to me like overkill to me. You could accomplish much the same by using a "depth-1 network" and running an impulse through it. Caveat, though: I am just a novice here, so there could very well be a lot of subtlety I'm missing out on. [1]: https://en.wikipedia.org/wiki/Waveshaper https://en.wikipedia.org/wiki/Waveshaper
- ndm000 6y agoI think the real innovation here is that this was done on just a few minutes of training data, opening up the possibility for all kinds of effects / amps to be modeled through this same method somewhat easily. I'm not sure how current DSPs are designed, but this is likely orders of magnitude more simple than designing the audio transformations (digital or analog) manually.
- deleted 6y ago[deleted]
- EamonnMR 6y agoAdd the ability to train on arbitrary effects as inputs and this will a best-selling VST for whoever can make it first.
- dharma1 6y agoHere is the original paper from 2019 by Eero-Pekka Damskägg- https://research.aalto.fi/en/publications/realtime-modeling-of-audio-distortion-circuits-with-deep-learning(5f5abe8a-9875-40a0-8e27-39109077f4e3).html https://research.aalto.fi/en/publications/realtime-modeling-... It was also published as a realtime JUCE project, which might be more useful for actual (realtime VST/AU) use: https://github.com/damskaggep/WaveNetVA https://github.com/damskaggep/WaveNetVA Alec Wright has done more work on this since then, using it for amplifiers: https://www.aalto.fi/en/news/deep-learning-can-fool-listeners-by-imitating-any-guitar-amplifier https://www.aalto.fi/en/news/deep-learning-can-fool-listener... And time variant effects: https://github.com/Alec-Wright/NeuralTimeVaryFx https://github.com/Alec-Wright/NeuralTimeVaryFx
- deleted 6y ago[deleted]
- SeanFerree 6y agoLove this! Great read!
- sdenton4 6y agoIt has been said that if we achieve the ability to fully simulate the universe from initial conditions, the first application will be creating a perfect recreation of Marvin Gaye's Roland 808 drum machine in a 1982 performance.
- mrob 6y agoThis isn't bad, but the note decays sound noticeably different. My guess is that the NN doesn't know that human ears have non-linear response that makes them more sensitive to errors in the decay than the attack, so it treats them equivalently. If this is the case then it might be fixable by using logarithmic scale audio samples instead of linear. The non-linearity of the ear is frequency dependent[0], but in practice I suspect it would be sufficient to pre-process the linear PCM data with x=sqrt(x) and undo before playback with x=x^2. [0] https://en.wikipedia.org/wiki/Equal-loudness_contour https://en.wikipedia.org/wiki/Equal-loudness_contour
- rubatuga 6y agoWhy square root and not log?
- mrob 6y agoCheap and dirty fast calculation. I don't actually know what the best mapping is, so I'd start with this.
- thesausageking 6y agoI came into the comments to say the same thing. To my ears, the NN versions roll off unnaturally at the end and that makes them really easy to identify as artificial.
- veenkar 6y agoB-but a simple convolution would do the same. Or for faster operation - a transfer function obtained using least squares method. NN is kinda overkill for this, but it's cool POC anyways ;)
- ssalazar 6y agoNope- Tubescreamer is non-linear so a simple transfer function won't do it.
- fallingfrog 6y agoI think that whereas most guitar effects are really very simple (gain and clipping, or delaying the signal and adding it back in), this approach will probably work just fine. But, it is sort of using a sledgehammer where a tap from a spoon will do- the original tube screamer is just an op amp and a couple diodes, plus a bit of eq! Not much to it. Plus, your real problems are going to be noise level (tube screamers in particular are noisy but a discrete transistor distortion can be made very very quiet). your a/d converter, your power requirements (comparable analog distortion effects use a few milliwatts) and cost. Edit: But that said, this is a super cool project! Good job! Sorry I just realized that what I wrote was kind of negative.