Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kastnerkyle
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
14 ms
·
61.
▲
by
kastnerkyle
10y ago
You might also enjoy http://badsamples.tumblr.com/
62.
▲
by
kastnerkyle
10y ago
OP and others may be interested in this approach by Sony CSL from ~2012 using constraint Markov chains for matching lyrical style with explicit rhyme constraints, including (e.g.) Bob Dylan style with constraint of Beatles' Yesterday [
63.
▲
by
kastnerkyle
10y ago
Extrapolation in text space is not the same as more abstract movement - even simple things like word2vec can capture a lot of meaning through context. Probably not enough to construct wordplay like this yet, but it is not outside the realm
64.
▲
by
kastnerkyle
10y ago
Deeper meaning is inferred by the viewer/listener every bit as much as it is implied by the creator. I think the future is in empowering artists with better computational tools to iterate, create, and explore. That said, I think saying
65.
▲
by
kastnerkyle
10y ago
I think the man you are referring to is David Cope [0] [0] http://artsites.ucsc.edu/faculty/cope/experiments.htm
66.
▲
by
kastnerkyle
10y ago
There are a number of papers tackling a similar task [0][1][2] for anyone who is interested. There isn't enough information to tell exactly what is going on with Deepgram, but one way to approach this would be to construct a shared emb
67.
▲
by
kastnerkyle
10y ago
Re-implementation will be hard, several people (including me) have been working on related architectures, but they have a few extra tricks in WaveNet that seem to make all the difference, on top of what I assume is "monster scale train
68.
▲
by
kastnerkyle
10y ago
It would be a slow (but very efficient information-wise - only have to send text which itself can be compressed!) decompression process with current models / hardware due to sequential relationships in generation. I am sure people will
69.
▲
by
kastnerkyle
10y ago
I did a review for PixelCNN as a part of my summer internship, it covers a bit about how careful masking can be used to create a chain of conditional probabilities [0], which AFAIK is exactly how this "causal convolution" works (c
70.
▲
by
kastnerkyle
10y ago
Relatively, training is fast (due to parallelism / masking so you don't have to sample during training) but during generation sampling is a sequential process. They talk about it a bit in the previous papers for PixelCNN and Pixel
71.
▲
by
kastnerkyle
10y ago
At least in HF, this "dead air" is anything but dead (lightning, channels with lots of fading, things you only catch due to ionospheric effects, etc). Noise is basically any signal you don't care about in lower frequency band
72.
▲
by
kastnerkyle
10y ago
As vitovito mentions, wideband data in high resolution is incredibly expensive from a storage perspective. There are entities that do it, but this also has the same privacy concerns as saving every packet that comes over a wire - in fact
73.
▲
by
kastnerkyle
10y ago
OT: Can you send me a link to your 1999 Elsevier article? The old thread is locked due to age but the topic is very relevant to me... As to on topic discussion - wouldn't you only need to take out the comms? Then the fellow subs wouldn
74.
▲
by
kastnerkyle
10y ago
Sure, there are a lot of potential things to do - either by enumerating towards the full "brute force" loss partway as I think you are hinting toward, or treating as a subsequences, or something else [0, 1, 2]. The point is, IMO R
75.
▲
by
kastnerkyle
10y ago
Hybrid DNN-HMM type systems (or CNN-HMM) still roundly beat purely RNN based systems on public benchmarks (~5% across tasks such as Switchboard, WSJ, even a bit behind on TIMIT) and most of the best internal systems at the companies I kno
76.
▲
by
kastnerkyle
10y ago
RNN for discrete structured prediction tasks often use a technique called beam search during decoding (see for example papers on neural MT), which is about as close as you can get to a true Viterbi when you don't have independence assu
77.
▲
by
kastnerkyle
10y ago
Montezuma's Revenge, Castle Wolfenstein (not the shooter), and puzzle games in general have a problem of long term credit assignment and sparse reward. This "intrinsic reward" approach form the paper, based on pseudo counts s
78.
▲
by
kastnerkyle
10y ago
Note that DCGAN (or standard GAN generally) doesn't have an encode path (or any other way to easily condition generation) so it is not well suited to this particular task. The image quality is stellar though, and I linked to several
79.
▲
by
kastnerkyle
10y ago
Sure - that is always an option especially in deep learning right now. But this current crop of models (counting in DCGAN and LAPGAN/Eyescream) has really made a leap in my eyes from before "oh cool generative model" to "
80.
▲
by
kastnerkyle
10y ago
Yes, although the particular technique used here (variational autoencoder, or VAE) doesn't benefit from increasing the code space due to a particular penalty (KL divergence against a fixed N(0, 1) Gaussian prior) during training. So ev
81.
▲
by
kastnerkyle
10y ago
The future is now! [0,1] [0] http://richzhang.github.io/colorization/ [1] https://www.youtube.com/watch?v=cizgVZ8rjKA
82.
▲
by
kastnerkyle
10y ago
This technique is quite different from deep dream. An autoencoder is a model that is trained to compress then reconstruct an image. After seeing and trying to reconstruct many images, it learns how to make a "compressed representation&
83.
▲
by
kastnerkyle
10y ago
Thinking of "Watson" more as a catchall term for machine learning research at IBM is more useful than thinking of it as a unified platform (as the marketers try to sell it). This includes research efforts in speech recognition, NL
84.
▲
by
kastnerkyle
10y ago
As you say, if your only goal is ShoeBot this (scripts, functions, etc.) is a great way to go. One of the key points of deep learning approaches is that you have this abstract, powerful computational device that is trained to extract the ne
85.
▲
A guide to convolution arithmetic for deep learning
(arxiv.org)
6 points
by
kastnerkyle
11y ago
|
0 comments
86.
▲
by
kastnerkyle
11y ago
RNNs are optimizing a probability, at least with standard training to maximize likelihood. As you say, with a different training criterion sure it could do something different - but so could a Markov chain if it was built to optimize a diff
87.
▲
by
kastnerkyle
11y ago
RNNs (recurrent neural networks) can be seen as generalizations of Markov chains - or as I prefer think of it, Markov chains are very limited RNNs. There is an enormous amount of research happening around these sequential models (and differ
88.
▲
by
kastnerkyle
11y ago
For anyone interested in Markov chains (aimed at language) in Python, and their relationship to the larger world of language modeling (including "modern" RNNLMs and so on) this post by Yoav Goldberg is an excellent introduction, w
89.
▲
by
kastnerkyle
11y ago
It is a dense read, but you might have a look at [1]. This is how attention is implemented in Theano. Basically the key is going "3D" per timestep (where 2D per timestep is the norm when doing minibatch training), then taking a we
90.
▲
by
kastnerkyle
11y ago
This depends on implementation (unrolled RNNs vs true recurrence). Each minibatch needs to be the same length, but that is it. And that is even implementation dependent - if your core RNN had a special symbol for "EOS" it could
More ›