Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
yvn1uo
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
yvn1uo
3y ago
I actually think the chip level HW-SW co-design is a good idea. It does open up more opportunities to mitigate communication issue than optimizing the mapping given a fixed chip and system design. For example, the number of GPUs per server
2.
▲
by
yvn1uo
3y ago
One possible explanation is that they hit the teraflops number during the prefill stage, where you can process all tokens at once, and are generally more operationally intensive, so you can use more compute. Utilization usually drops during
3.
▲
by
yvn1uo
3y ago
1.5 years is actually not that bad. In fact, all changes and improvements to LLMs since the original Transformer paper is just the size -- tensor dimension, layers, etc. GPT-3, which is still widely used today, was proposed more than 3 year