Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kastnerkyle
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
31.
▲
by
kastnerkyle
9y ago
This has changed recently, full seq2seq is now matching hybrid models [0]. [0] https://arxiv.org/abs/1712.01769
32.
▲
by
kastnerkyle
9y ago
For people who want simple, out of the box stuff (not necessarily in Python) for just getting phonemes I can also recommend [0]. Not amazing recognition quality, but dead simple setup, and it is possible to integrate a language model as wel
33.
▲
by
kastnerkyle
9y ago
You might find the work of Mason Bretan on this [0][1], and the video demo here https://www.youtube.com/watch?v=BbyvbO2F7ug relevant to your future work. Nice writeup! [0] https://arxiv.org/abs/1612.037
34.
▲
by
kastnerkyle
9y ago
On rule based solvers - this is true if you simply stop at "be within the rules" and do strict constraint solving ala the coloring problem. I think there is potential to do something by blending all 3 (rules and constraint checkin
35.
▲
by
kastnerkyle
9y ago
- De-reverberation is being intensely (in some niches at least) studied with these tools, and adding the reverb back (if you can learn and remove it from mixed records) means you have probably modeled it well enough to do whatever you wan
36.
▲
by
kastnerkyle
9y ago
It can be, though some covers are "straight up", while others (generally the memorable ones) are practically a new creation in themselves, with a sliding scale in between. But for "cover song" meaning something like Hend
37.
▲
by
kastnerkyle
9y ago
"Style transfer" also rarely works for object level transfer - it is more pattern based (high frequency content is often the "style" that is enhanced and transferred). Really nice transfers in practice sometimes require
38.
▲
by
kastnerkyle
9y ago
Agreed, although I generally think in terms of "obeying traditional music theory (to some extent) is necessary but not sufficient for a listenable melody". This also changes depending on genre and era - one reason for targeting ea
39.
▲
by
kastnerkyle
9y ago
Gabriel (as well as the others on the team) have definitely looked at these areas - if things were left out/not "featurized" it was likely done via an ablation test, or showed improvement over benchmarks, or maybe just to set
40.
▲
by
kastnerkyle
9y ago
https://arxiv.org/abs/1703.07469
41.
▲
by
kastnerkyle
9y ago
Stronger incorporation of hard constraints / rule or program learning in models for control and sequential decision making.
42.
▲
by
kastnerkyle
9y ago
I don't think mAP on VOC is indicative of real-world performance - especially given COCO mAP is just now cracking the 40s with similar architectures. But the gains in proposal based detection/instance segmentation from 2014 to now
43.
▲
by
kastnerkyle
9y ago
In addition to what /u/sotelo said, these types of models also have potential to detect fake audio - if something sounds good to the ear but is highly unlikely according to the model, it may indicate something that needs further
44.
▲
by
kastnerkyle
9y ago
There is a lot of interesting work going on in this area right now in the research community - using ML for things like program synthesis, meta-learning, and combining ideas from constraint satisfaction with ML approaches. The rules based a
45.
▲
by
kastnerkyle
9y ago
You mean something like this? [0] As far as I know from the mailing list, they are always working on more tools for integration so musicians can use them in different environments and different ways. As an amateur composer as well, I have f
46.
▲
by
kastnerkyle
9y ago
My take (as someone working in the area) on how/why we are encouraged by this progress: The model is unconditional i.e. any concept of form at all was learned directly from data, with no real structural hints to what it should learn.
47.
▲
by
kastnerkyle
9y ago
A bigger dataset of MIDI with velocity information and performance timing would be really, really great. High temperature versus low is tough to compare - I find that sometimes low temperature seems better, then I change the random seed and
48.
▲
by
kastnerkyle
9y ago
This is stunning! Great stuff. Since the input and prediction is a single sequence, did you experiment with beamsearch/stochastic beamsearch decoding (maybe with additional diversity criteria)? I found that even simple models (markov c
49.
▲
by
kastnerkyle
9y ago
Japan, as usual, is ahead of the game here! [0] [0] https://www.youtube.com/watch?v=pEaBqiLeCu0
50.
▲
by
kastnerkyle
9y ago
I tried similar approaches long ago (~2 years now?) with something related to RNN-RBM and it showed some slight glimmer of promise, and still think there might be some clever ways to combine concatenative methods and deep learning to avoid
51.
▲
by
kastnerkyle
9y ago
There is an older paper [0] and demo from [1] Alex Graves that inspired a ton of work around handwriting, and then speech. Previous work from Jose Sotelo et. al. (including me) called char2wav [2] is a close neighbor to Graves' approac
52.
▲
by
kastnerkyle
9y ago
This model is quite cool, but also quite a bit different than what lyrebird.ai is doing. NPSS has a lot of extra information in the control inputs about pronunciation and timing (the part-of-phoneme timer feature) - this means that most of
53.
▲
by
kastnerkyle
9y ago
I personally like this sample the most [0]. Note that these samples are not cherry-picked - having worked with very related algorithms [1] once it is trained well it pretty much "just works". There is a lot of room for DSP/ha
54.
▲
by
kastnerkyle
10y ago
You might be interested in minute 50 onward [0], or this recent paper from Facebook [1]. [0] https://www.youtube.com/watch?v=-yX1SYeDHbg&list=PLE6Wd9FR--... [1] https://arxiv.org/abs/1703.07684
55.
▲
by
kastnerkyle
10y ago
So in that case it has the same amount of parameters as GRU - did you compare with that?
56.
▲
by
kastnerkyle
10y ago
If you want to see the most basic form of this pipeline, I have a blog post on "bad speech synthesis" [0]. There are open source versions of WaveNet for TTS [1], but I have not run the code myself or seen the quality of the output
57.
▲
by
kastnerkyle
10y ago
Showing one of the top two samples from their blog (full prediction) along with the one you link (only acoustic model) would more clearly show what you are getting at in the explanation, since how the text inference works is most of the com
58.
▲
by
kastnerkyle
10y ago
On "...clearly wouldn't be used for real speech synthesis" - I think it depends what language you are looking at. For languages with many years of linguistic research (e.g. English), it will be hard to beat a good parametric
59.
▲
by
kastnerkyle
10y ago
This is no joke something I have considered - do you have a source on a pairing of "read speech" and "transcript" for this? I could process the movies myself but that seems... tedious...
60.
▲
by
kastnerkyle
10y ago
We will have a longer paper out soon with more details about the training process - but for now the code is there as well. There were a few small things in training that seem key to getting good results, we are analyzing those things and ho
More ›