Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
apophis-ren
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
2 ms
·
1.
▲
by
apophis-ren
1y ago
It's mentioned in the article. But for really, really long-horizon tasks, it might be reasonable that you don't want to have a small discount factor. For example, if you have really sparse rewards in a long-horizon task (say, a re
2.
▲
by
apophis-ren
2y ago
Flash attention is an implementation trick; you can implement MHA/GQA, for example, with flash attention.