Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
shawntan
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
shawntan
2y ago
Not sure what you got out of the paper, but for me it was more spurring ideas about how to fix this in future architectures. Don't think anyone worth their salt would look at this and think : oh well that's that then.
32.
▲
by
shawntan
2y ago
Yes, but unfortunately that doesn't answer the question the title poses.
33.
▲
by
shawntan
2y ago
Yes, but the article doesn't really answer this question.
34.
▲
by
shawntan
2y ago
If we're trying to quantify what they can NEVER do, I think we'd have to resort to some theoretical results rather than a list empirical evidence of what they can't do now. The terminology you'd look for in the literatur
35.
▲
by
shawntan
3y ago
Not sure if this is the type of answer you're looking for, but RWKV is not really recurrent the same way RNNs are recurrent. This quasi-recurrentness allows it and its comrades to use algorithms like parallel SCAN to achieve log N comp
36.
▲
by
shawntan
3y ago
You're both kinda right. The type of computation that happens for that attention step that you refer to is parallel. I would say the thing that is "constant" is the computation graph depth (the number of sequential computati
37.
▲
by
shawntan
3y ago
It's probably covered many other machine learning books, but generally assumed to be common knowledge here: The final validation for what you need the application for should be done on your end. Need to use CRFs for NER? Test on your o
38.
▲
by
shawntan
3y ago
Exactly this. The tokens generated should always be valid, unless some post-processing layer between the model's output and the user interface detects for some keywords which it would prefer to filter out. In which case I suppose there
39.
▲
by
shawntan
3y ago
NaNs are not only possible by design, but are extremely common. Training of LLMs involve many tricks about how to deal with training steps that result in NaNs. Quantisation of LLMs also require dealing with huge outlier values.
40.
▲
by
shawntan
3y ago
This is a strange explanation. These models usually give as output the same set of vocabulary that was used as its input vocabulary. > the model isn’t trained on understanding the useRalativeImagePath token, and so it outputs something t
41.
▲
by
shawntan
3y ago
TSPs are not unsolvable. My point was that Transformers and neural networks as they are now are not Turing machines if you don't allow for the model to grow with the input size. That said, it has to grow in "depth" not just p
42.
▲
by
shawntan
3y ago
Something close exists: RASP https://arxiv.org/abs/2106.06981 python implementation: https://srush.github.io/raspy/
43.
▲
by
shawntan
3y ago
Short version of this without the caveats: It's not even Turing complete. I review a few papers on the topic here: https://blog.wtf.sg/posts/2023-02-03-the-new-xor-problem/
44.
▲
by
shawntan
3y ago
Would I be able to solve the Travelling Salesman Problem with a Transformer with the appropriately assigned weights? That would be an achievement. You'd beat some known bounds of the complexity of TSP.
45.
▲
by
shawntan
3y ago
Right, I see your point. Since Sudoku is fixed-size, you can always construct a Transformer with the worse-case depth. That makes sense. I was assuming given a trained Transformer, you wouldn't know how many effective "steps of co
46.
▲
by
shawntan
3y ago
great substack and podcast (love the guests you have on) i can't help but notice that Swyx (author of latent space) is accumulating a body of work exploring LLMs and covering AI progress. is this something that is analysis by an expert
47.
▲
by
shawntan
3y ago
I suppose you mean in order to give the answer to a Sudoku puzzle, you'd need a string of tokens anyway: [(x,y) grid coordinates], [digit]. I think if we're getting specific to this particular Sudoku example, the CoT would probabl
48.
▲
by
shawntan
3y ago
I recommend reading the theoretical work on the computational capabilities of Transformers: https://twitter.com/lambdaviking/status/1630581475425828864 References to other work can probably be found in that artic
49.
▲
by
shawntan
3y ago
I'm a little bit tired of the Bitter Lesson copypasta. It doesn't seem to me like point 1 to 3 have been proven to be true about ConvNets. Sure, scaling convnets further help, I don't think anyone would argue that training a
50.
▲
by
shawntan
3y ago
Right. So where I end up on this, given the examples of intuitions that DO work, is it's always the _right_ levels of prior knowledge that's needed. The intuitions on language (encoding basic grammar) didn't pan out, but the
51.
▲
by
shawntan
3y ago
The scope of "what to try" is large, we (as a community) should prioritise things that we think would work. If the criteria is not only "faster compute" it would seem "things that mimic human high level processes&qu
52.
▲
by
shawntan
3y ago
But "don't try to codify 'insight' into the process" seems to suggest "don't try different approaches". I'm not sure how people can at once trot out the "Bitter Lesson" and interpret i
53.
▲
by
shawntan
3y ago
Depending on your starting assumptions, Transformers, RNNs and LSTMS can either be a) Turing-complete, in which case they can all match unbounded brackets, or b) they are all not Turing-complete, LSTMs might be able to match unbounded brack
54.
▲
by
shawntan
3y ago
You should check out the work referenced in the abstract, the most recent here: https://twitter.com/lambdaviking/status/1630581475425828864 There are limitations for what a Transformer can compute if we do not all
55.
▲
by
shawntan
3y ago
Top-left graphic is Conway's Game of Life. Activates when you hit the Konami code ( https://en.wikipedia.org/wiki/Konami_Code )
56.
▲
by
shawntan
3y ago
> I thought the key takeaway is that nonlinearity is important. Multilayer perceptrons can be collapsed into an equivalent single layer if there isn't nonlinearity thrown in. multi-input XOR can be solved using a form of sin(x) func
57.
▲
by
shawntan
3y ago
Yeah I really like this project, and I've been meaning to dive into it.
58.
▲
by
shawntan
3y ago
Yeah it is. I've only learned it while looking up material for this post. I figured people know NAND is NOT(AND(.)), so NXOR might be more familiar counterpart to XOR.
59.
▲
by
shawntan
3y ago
To be clear, LSTMs and RNNs have no issue with the parity problem. As for humans, I think you could run into boredom / fatigue if you just told someone to sit there and 'maintain state' for PARITY, but I'm sure most of u
60.
▲
by
shawntan
3y ago
On the memory limitation: Yeah, lack of an infinite tape is a real bummer. On some level I think we're some kind of bounded-tape automaton, which still makes the current limitations of Transformers pretty severe. (Yes, yes, prompting,
More ›