Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kastnerkyle
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
18 ms
·
91.
▲
by
kastnerkyle
11y ago
The MRF loss is patch based, and adapted from CNNMRF [1]. Since the precursors to PatchMatch (as mentioned in the PatchMatch paper) were MRF + belief propagation based, I am pretty sure it could be done with some tweaking. These analogies s
92.
▲
by
kastnerkyle
11y ago
I have an easy wrapper to pocketsphinx here [1] - it has come in handy for me in the past. Another option is a gstreamer server with Kaldi [2]. [1] https://github.com/kastnerkyle/ez-phones [2] https://www.re
93.
▲
by
kastnerkyle
11y ago
If you use speech recognition systems on mobile (or a public web API), they are often handicapped due to space or processing constraints. Full on recognition models in clean environments (no background noise, other speech, no compression e.
94.
▲
by
kastnerkyle
11y ago
The fact that even two or three years ago (pre 2012 for sure) tagging anything with AI outside academia was a surefire recipe for ridicule should make anyone very wary of something that explicitly advertises as such. It is convenient bran
95.
▲
by
kastnerkyle
11y ago
This is mostly true until you need new data that is not a .mat or image, or to deploy to a server for production (licensing Matlab for multi-core servers is pretty rough), or deal with strings, or send stuff over the network, or one of the
96.
▲
by
kastnerkyle
11y ago
Parts of machine learning used to be focused on density estimation / manifold learning for exactly this reason. It seems to be the case that supervised tasks (for sufficiently complicated and diverse problems e.g. ILSVRC, MSCOCO) impli
97.
▲
by
kastnerkyle
11y ago
If we really want to get into "what is better" - the state of the art has recently moved significantly past this (paper [1] and blog [2]) - not to mention the recent submissions to ICLR in this realm [3][4]. GAN training is very
98.
▲
by
kastnerkyle
11y ago
This looks like a great resource! I have been using complex2real[1] for years, but this looks like another great link. Thanks. [1] http://complextoreal.com/tutorials/#.VlXqyMr88xU
99.
▲
by
kastnerkyle
11y ago
X. Zhang and Y. LeCun had an article [1] about using this technique for regularization very recently. We aren't talking specifically about overfitting in this case, but rather a lack of diversity / size in the training data, howev
100.
▲
by
kastnerkyle
11y ago
I don't think anyone is claiming TF won't get much, much faster (probably very soon). But claiming that right now - today - TF kills existing toolkits like Torch, Caffe, and Theano (which I have seen here and elsewhere - though
101.
▲
by
kastnerkyle
11y ago
Both Google and FB (from what I understand) have a ton of tricks to help this, but I expect Yann was speaking generally even with all these tricks, and he is right from what I have seen. The old paper on DistBelief talks about topping out a
102.
▲
by
kastnerkyle
11y ago
My point is more that even if they were on the same footing from benchmark timings, v2 is still far behind what is supported in Torch, Caffe, and Theano right now (v3 in all IIRC). Your comparison is very fair, and it is good insight!
103.
▲
by
kastnerkyle
11y ago
One case where I see 1 as being necessary is the softmax size / vocabulary boost in the seq2seq paper (8 GPUs IIRc, and 4 were dedicated to a softmax!). There are other ways to handle this (hierarchical softmax, sampled softmax, skip-t
104.
▲
by
kastnerkyle
11y ago
This benchmark basically shows that releasing TensorFlow with cudnn v2 backend support hurts - v2 is quite a bit slower than v3 (current) and v4 (upcoming). TF has announced that they will update to v4 support, which should help quite a bit
105.
▲
by
kastnerkyle
11y ago
Yann LeCun has been working on hardware implementations of convnets almost since the beginning. LeNet 5 (the check reader of ATT, circa early 90s) needed a dedicated and specialized hardware implementation IIRC. It could be kneejerk, but he
106.
▲
by
kastnerkyle
11y ago
For big LSTMs and long-ish sequences, the intermediate gradients can take up a huge amount of memory - often more than the model parameters themselves. In my experience it is mostly big LSTMs that need the 12GB+ GPUs. You can reduce the bat
107.
▲
by
kastnerkyle
11y ago
If this really works, the only options I see are that a) it works at a massively ineffective power transfer rate or b) it is dangerous if you accidentally get in the focused beam (or c) a combination of the two). Any kind of beamforming whi
108.
▲
by
kastnerkyle
11y ago
This is the point of the benchmarks in literature such as MNIST, CIFAR-10, ImageNet, SVHN, and so on. You can see a pretty comprehensive list here [1] that also shows papers and their reported performance. There are a lot of implementation
109.
▲
by
kastnerkyle
11y ago
I don't think overfitting will be an issue (at all!) on a 65TB dataset. A bigger CNN model should be more effective at this task than a DBN, (almost) regardless of the features added. If we can use convolutional generative models (fr
110.
▲
by
kastnerkyle
11y ago
If anyone is interested in a beginner introduction to what wavelets are generally, and how they are used - I did a blog post a while back using numpy: http://kastnerkyle.github.io/posts/wavelets/
111.
▲
by
kastnerkyle
11y ago
Just to be abundantly clear - the ladder network [1] outscores this by a moderate margin, and is also actually SOTA without data augmentation. In my mind at least, adding data rotations in the latent space is still different than a fully
112.
▲
by
kastnerkyle
11y ago
One more... http://www.robots.ox.ac.uk/~szheng/CRFasRNN.html
113.
▲
by
kastnerkyle
11y ago
This is really cool stuff - the network structure reminds me a lot of Graves' MDRNN[1] and Grid LSTM[2], as well as some work I helped with (ReNet [3]) I wonder if the structure over frequency/time is too "regular" - in
114.
▲
by
kastnerkyle
11y ago
The point of this type of modeling is to see how far a black box can get. I don't think anyone is claiming this LSTM is creating "state of the art" art.
115.
▲
by
kastnerkyle
11y ago
The literature is mostly image based, but as other users have said voice, text, and video have also seen substantial gains. It has also been used to great effect for taxi destination prediction (heterogenous information) [1], drug discovery
116.
▲
by
kastnerkyle
11y ago
For anyone who is interested in a python version with inferior plots I have a notebook here: http://kastnerkyle.github.io/posts/linear-regression/ Closed form is at the bottom, basically invalidating the whole res
117.
▲
by
kastnerkyle
11y ago
I disagree with the "great advance" dig. If you took a time machine 5 years ago and showed any of the recent advances in deep neural networks (without showing the algorithmic techniques) people would say "this is AI". Th
118.
▲
by
kastnerkyle
11y ago
One thing which works against the "cite everything" approach is that most of the major conferences have page limits of 8-10 pages with 1 page bonus for references. That means if you go over 1 page of references for (at least NIPS)
119.
▲
by
kastnerkyle
11y ago
I think the poster got mixed up. Quoc Le was the first author on "Distributed Representations of Sentences and Documents" aka paragraph2vec, so he has been involved in the x2vec scene. Just not the word2vec. And I would argue that
120.
▲
by
kastnerkyle
11y ago
This is pretty hard - we use raw data for speech [1, talked about in comment above] but it still needs some work to do really good synthesis. FFT is not really the way to go either - then you still need to deal with the problems of complex
More ›