Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
andy12_
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
151.
▲
by
andy12_
2y ago
This grant [1]. "SUPPORT SEXUAL AND REPRODUCTIVE HEALTH (SRH) CARE CLOSE TO THE DISPLACED POPULATIONS. [...] Place Of Performance //XGZ GAZA STRIP GAZA STRIP". Though I don't think that it will be used exclusively f
152.
▲
by
andy12_
2y ago
It's actually implied in the paper that the neural memory module M can be anything, and there's probably a lot of room to test different kinds of architectures for M. But in this paper M is an MLP of 1 layer (fig. 7 is an ablation
153.
▲
by
andy12_
2y ago
The eq. 12 is a loss function to associate a given key and value in the memory MLP using test-time training with gradient-descent. The eq. 15 is simply the operation to query a value that was previously inserted in previous tokens using eq.
154.
▲
by
andy12_
2y ago
Is Renfe really the second best by reliability? That makes me very surprised, as I have never taken a ride in Renfe that wasn't 5-10 minutes late.
155.
▲
Tokenformer: Rethinking transformer scaling with tokenized model parameters
(arxiv.org)
3 points
by
andy12_
2y ago
|
1 comments
156.
▲
by
andy12_
2y ago
Transformers have become the predominant architecture in foundation models due to their excellent performance across various domains. However, the substantial cost of scaling these models remains a significant concern. This problem arises p
157.
▲
Selective Attention Improves Transformer
(arxiv.org)
1 points
by
andy12_
2y ago
|
1 comments
158.
▲
by
andy12_
2y ago
Unneeded elements in the attention’s context degrade performance. We introduce Selective Attention, a simple parameter-free change to the standard attention mechanism which reduces attention to unneeded elements. Selective attention improve
159.
▲
The AdEMAMix Optimizer: Better, Faster, Older
(arxiv.org)
2 points
by
andy12_
2y ago
|
0 comments
160.
▲
by
andy12_
2y ago
This is so cool. I want to try to implement something with this right now.
161.
▲
by
andy12_
2y ago
a.k.a, the same work as Anthropic, but with less interpretable and interesting features. I guess there won't be Golden Gate[0] GPT anytime soon. I mean, you just have to compare the couple of interesting features of the OpenAI feature
162.
▲
by
andy12_
2y ago
At first I thoght that this was another Medusa-like paper, simply using more unembed heads for guessing subsequent tokes, but damn, not at all. This is amazing. And it doesn't even use extra parameters, it's just an auxiliary trai
163.
▲
by
andy12_
3y ago
> MoD transformers demonstrate the value of routing among different types of computations. In this work the types were either the conventional transformer block, or a null computation (functionally equivalent to multiplying by zero). How