Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kraken12
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
kraken12
3y ago
The 4090 has half the memory bandwidth, so it could not get a 5X gain, it would actually run slower on a memory bound LLM like this.
2.
▲
by
kraken12
3y ago
Yeah, it is an architectural simulation study, this is what is usually done right at the beginning before resources are allocated to go deep on idea. So in that sense it is imaginary; but this is how new ideas get incubated.
3.
▲
by
kraken12
3y ago
Yeah, a preliminary architectural study to sanity check if an idea could potentially pay off.
4.
▲
by
kraken12
3y ago
Maybe they could do something like AMD's GPU memory stacking, that is good for scaling, and of course they are using many chips not one chip..
5.
▲
by
kraken12
3y ago
Seems like HN comments have determined that the cost number is not fudged..
6.
▲
by
kraken12
3y ago
Seems to me the 18 tokens per second from [1] is the throughput and includes the batch size, so I don't think they misread the Deepspeed inference paper. So the chiplet ASIC supercomputer paper would seem to show a decent performance&#
7.
▲
by
kraken12
3y ago
Yep, it's a research paper in comp arch, the initial proof-of-concept study before you go and spend real money on it.
8.
▲
by
kraken12
3y ago
They are using many chips and taking advantage of the way data flows in LLMs to make it work; so it would be cost-effective unlike Cerebras
9.
▲
Chiplet ASIC supercomputers for LLMs like GPT-4
(arxiv.org)
142 points
by
kraken12
3y ago
|
84 comments
10.
▲
Geocomputers and the Commercial Borg
(sigarch.org)
1 points
by
kraken12
10y ago
|
0 comments