Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
fchaubard
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
LLMs are NOT Turing Complete (at train time), we need "train time recurrence"
(fchaubard.github.io)
2 points
by
fchaubard
1y ago
|
0 comments
2.
▲
by
fchaubard
1y ago
exactly
3.
▲
Scaling RNNs to Billions of Parameters with Zero Order
(arxiv.org)
7 points
by
fchaubard
1y ago
|
3 comments
4.
▲
by
fchaubard
1y ago
Layman Abstract: Transformers keep around all previous tokens for each generated token, so they take up ENORMOUS gpu memory and cost during inference. But humans do not, we page in / out of our small, fixed-size "working memory&qu
5.
▲
by
fchaubard
2y ago
Yes. It’s more of a class spanning thing. I wanted batch composition across the two microbatches to be the same. So if you have class 1,2,3 in batch one and class 4,5,6 in class two I would fully expect the cosine distance to be orthogonal
6.
▲
by
fchaubard
2y ago
I think about designing your ideal solver. What do you want in a solver. I want my solver to squeeze all the juice out of the train that it possibly can and no more. If your problem is complete noise, I don’t want my solver getting 100% tra
7.
▲
by
fchaubard
2y ago
Yes! I think this a great area of research. If you think of the gradient values as a blame score for why you got the answer wrong, then you can have a lot of fun with exploring which weights light up for different problems. A note, in Ring
8.
▲
by
fchaubard
2y ago
I’ll do my best to answer here. > Do you expect instability between successive macrobatch gradients? That is, why are you comparing microgradients within a single batch, adding a whole bunch of serialization headaches, rather than compar
9.
▲
by
fchaubard
2y ago
Hey thanks! Ya we tried similar strategies to this and could not beat cosine distance < tau, average, else skip. It was too much to put in the paper and we may put it in the arxiv version but we tried Sign AND gating and zero’ing out if
10.
▲
by
fchaubard
2y ago
Yes it will allow stable training at much smaller batch sizes. Test it out and let us know if it works for your use case!
11.
▲
by
fchaubard
2y ago
Here too!