Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
owl_brawl
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
owl_brawl
2y ago
For those who haven't read it, Rich Caruana's thesis on multi-task learning is beautifully written (the cited 1998 paper here). It's amazing to see how far the field has come, and, at the same time, how advanced the thinking
2.
▲
by
owl_brawl
3y ago
I would love answers to these questions too, particularly on the vocab size
3.
▲
by
owl_brawl
3y ago
Hi alekandreev, Any reason you decided to go with a token vocabulary size of 256k? Smaller vocab/vector sizes like most models in this size seem to be using (~16-32k) are much easier to work with. Would love to understand the technical
4.
▲
by
owl_brawl
3y ago
The "RAG" part (over the training set) is by far the smallest contribution to the performance gains reported (see the ablation study in section 5.2). I don't think the model is actually learning in-context from the selected s