Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
psb217
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
16 ms
·
121.
▲
by
psb217
7y ago
Most of the statements they make regarding orderless autoregression, including statements about the "independence assumption" made by BERT, are misleading at best.
122.
▲
by
psb217
9y ago
If you read the AGZ paper closely, they actually use checkpoints during training. Specifically, during training they only perform updates to the "stable" set of parameters when the current "learning" set of parameters pr
123.
▲
by
psb217
9y ago
Training via bootstrapping (i.e. dynamic programming) does reduce the state space for search when working with function approximation. It represents a bias in what sorts of values the value function approximator should predict for each stat
124.
▲
by
psb217
9y ago
You are correct. There is no TD learning in AGZ. The value network is trained to directly predict the game outcome given the current state, and is not trained through "bootstrapping" based on the next state's value estimate.
125.
▲
by
psb217
9y ago
TD learning is, in some sense, a component of policy iteration. TD learning is about learning the value function for a given policy. In policy iteration you use a value function to decide how to update the policy for which the value functio
126.
▲
by
psb217
9y ago
Two additional points are (1) dataset collection is low variance relative to fundamental algorithmic advances, and (2) dataset collection relies less on having tip-top research talent (than fundamental algorithmic advances).
127.
▲
by
psb217
9y ago
Ms PacMan is difficult for current general-purpose models. Nonetheless, it's accurate to say that other groups haven't focused on developing a model or method specifically tailored to the game. The result in the linked article is
128.
▲
by
psb217
9y ago
There's some cool new work -- to be presented at ICML in August -- that claims practical improvements over existing compression methods by using neural nets for "adaptive compression". The authors are strong academic research
129.
▲
by
psb217
10y ago
AlphaGo relied heavily on (supervised) pretraining, and that seemed fairly successful.
130.
▲
by
psb217
10y ago
It's fp16 that they neglect on the consumer cards. On the Tesla p1xx variants they now have full double-speed fp16, which is a great feature for machine/deep learning users. Leaving this out of the consumer cards is certainly anno
131.
▲
by
psb217
10y ago
It would be great if cleaned-up demo code for many of these models/algorithms could be shared in a single "deep RL quickstart" repo. Various implementations (sometimes of dubious correctness) are already scattered around Gith
132.
▲
by
psb217
10y ago
The P100s have full support for half-precision (i.e. 16 bit) floating point ops. This can mean ~2x improvements in speed and memory usage in comparison to the Pascal TitanX, which is the top "consumer" card. This difference is sig
133.
▲
by
psb217
10y ago
There's a minor typo ("than" is written as "that") in the third block of main-body text on the linked landing page.
134.
▲
by
psb217
10y ago
I've been living in Montreal for ~7 years without a car. It's not difficult. Yes, it's anecdotal, but I would find a car burdensome.
135.
▲
by
psb217
10y ago
Should've trained to generate the output strings directly, using an LSTM decoder... That would be more end-to-end.
136.
▲
by
psb217
11y ago
I don't recall saying that it was obvious Amazon would destroy the brick-and-mortar book business, or that Amazon would come to control significant chunks of internet infrastructure. Your comment was about the fiscal viability of selli
137.
▲
by
psb217
11y ago
Uh, if the expectation when one arrives ~1hr late is that they will also stay ~1hr late, what is there to abuse?
138.
▲
by
psb217
11y ago
Books are almost the perfect item to sell from a catalog. They require no personalization, they have a pronounced long tail demand distribution (i.e. a small store often won't have what you want), they can be stored space-efficiently a
139.
▲
by
psb217
11y ago
The rockputer comprises both the rocks and the mechanisms for moving the rocks in response to input. If the rock moving mechanism is structured properly, then the rock movement patterns could adapt to changes in the inputs to the overall ro
140.
▲
by
psb217
11y ago
Very true. These are rates in the SF Bay Area and NYC. IDK what it's like elsewhere, though I've been told it's possible to get offers in this range in London too.
141.
▲
by
psb217
11y ago
Yes, but for graduating PhD students 175k-225k/year is roughly market (base) rate for people with proven understanding and ability in the most relevant areas (i.e. multiple publications in top conferences like ICML/NIPS/CVPR&
142.
▲
by
psb217
11y ago
This is an overly restrictive view of RL. Yann's claim about the potential utility of RL, taken at face value, is clearly false. Backprop on a deterministic computation graph is equivalent to deterministic policy gradient in the same g
143.
▲
by
psb217
11y ago
The minds are the Culture and the minds are anarcho-communists (at least as far as I can tell).
144.
▲
by
psb217
11y ago
This is just the problem of forming automorphism-invariant encodings of a graph, but extended to permit weighted graphs. First, given the Gram matrix for your vectors, interpret its entries (read from left-to-right and top-to-bottom) as def
145.
▲
by
psb217
12y ago
This is actually how (some) couriers worked in Neal Stephenson's Snow Crash. They latched on to the back of passing cars/trucks with magnetic harpoons and hitched a free ride until they needed to head another way.
146.
▲
by
psb217
12y ago
As the sibling comment suggests, multi-threading through Cython isn't as smooth as it could be. But, it doesn't seem too bad. I've used it in a rather rudimentary way to accelerate key computations for some NLP models that I&
147.
▲
by
psb217
12y ago
People tend to hold their phones in "portrait" mode, rather than "landscape". So, width would typically refer to the second largest dimension when talking about (most) phones. Additionally, given that phones are 3d objec
148.
▲
by
psb217
12y ago
FYI, some folks (see: https://spectrallearning.github.io/icml2014/ ) are certainly working hard to apply eigenvalue decompositions (well, more like generalized SVDs) to any inference problems they can get their convexit
149.
▲
by
psb217
13y ago
"Lossless compression requires identifying the pattern that produced an input as perfectly as possible": no, it requires identifying it absolutely perfectly. This includes all vacuous information as well. E.g., in the Wikipedia ex
150.
▲
by
psb217
13y ago
Lossy compression, yes. Lossless compression, not so much. And, the parent post seems to have been hinting at this. I.e. being able to usefully "tl;dr" all of Wikipedia would require a rather intelligent system. But, recapitulatin
More ›