5 ms·
- De-reverberation is being intensely (in some niches at least) studied with these tools, and adding the reverb back (if you can learn and remove it from mixed
by kastnerkyle 9y ago
- De-reverberation is being intensely (in some niches at least) studied with these tools, and adding the reverb back (if you can learn and remove it from mixed records) means you have probably modeled it well enough to do whatever you want to. But de-reverberation is very hard especially with single source (1 mic) recordings, maybe looking at better reverb generation is easier with the right data. It would be nice to do something that didn't require the physical setup of impulse response recording.
- I haven't heard the term transposition for samples - is this the name for moving samples in pitch and compensating in time so that it sounds undistorted?
- On pure sound synthesis, see NSynth [0]. It's a beginning step toward this with abstract sounds.
- Sample morphing and changing the words with voices is pretty tough (overlaps neural TTS heavily, which is pretty recently starting to work well), but if you can synthesize a single voice decently the rest seems doable with manual intervention. Check Neural Parametric Singing Synthesizer [1], or try the online demo [2]
- Every few months there is a new convnet which improves the baselines for transcription, hopefully it is only a matter of time. The current best is quite a bit better than only a few years ago.
There is still a huge gap between "what is possible in research" and "what is easily usable in a practical editing/creation workflow" on a typical DAW, but hopefully that gap will diminish over time.
There are also a lot of tools from statistical signal processing that were doing these tasks to some degree, digging them back up and "neuralizing" the probability parts with neural nets is a promising way to get quality improvements, and would mirror what has been successful in neural TTS so far.
[0] https://experiments.withgoogle.com/ai/sound-maker https://experiments.withgoogle.com/ai/sound-maker
[1] http://www.dtic.upf.edu/~mblaauw/NPSS/ http://www.dtic.upf.edu/~mblaauw/NPSS/
[2] https://www.voiceful.io/demos.html https://www.voiceful.io/demos.html
- romaniv 9y ago>- I haven't heard the term transposition for samples - is this the name for moving samples in pitch and compensating in time so that it sounds undistorted? Yes, that's exactly it. There are traditional algorithms that do it, but most are variations on time stretching and granular synthesis. They aren't as good as they could be. A good algorithm would be able to take in a handful of examples and generate a model that seamlessly interpolates across different pitches and velocities, accounting for differences in timbre, attack, decay, noises and so on. Thank you for the links.