Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
spi
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
12 ms
·
61.
▲
by
spi
6y ago
Yes that might be the case. In my case I mostly trained big (tens to hundreds of millions of parameters) networks mostly made of 3x3 convolutions, and I think the V100 has dedicated hardware for that. Then as I mentioned you can get a furth
62.
▲
by
spi
6y ago
"Similar performance" still means 30%-50% slower [1] and half the RAM, not really that comparable. For much closer performance you should get a 2080ti, which should be roughly comparable in speed and have 11GB [edit: wrongly wrote
63.
▲
by
spi
6y ago
I fully agree with this. As a data scientist, I always think that this is a "natural" consequence of one of the main (if not _the_ main) metric used to evaluate machine translation algorithms, which is BLEU: https://en.
64.
▲
by
spi
8y ago
Perhaps more importantly than the limit on 200/2000 samples, the main difference from your code is that in the paper they only use the 10 first PCA (i.e. X is 10-dimensional instead of 784-dimensional). The fact that they do tends to c
65.
▲
by
spi
9y ago
Same disclaimer as v64: my background is also mathematics, but not in this field. As he says, what was proved so far is that: a_0 < a_1 <= p <= t <= 2^(a_0) The first inequality is by definition (a_1 is "the smallest cardin
66.
▲
by
spi
9y ago
Very nice article! I've been willing to delve into Bézier curves for a while now, and this looks like an awesome reference to start. I'd just like to make an annoying mathematician comment on the beginning (paragraph 3): you might