Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
punkgenius
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
punkgenius
3y ago
This seems to be more achievable and cost-effective than Cerebras. Some comments mention Cerebras cost millions for each 'die'.
2.
▲
by
punkgenius
3y ago
I think it’s all about the performance-to-cost ratio. The reason you need a cache is because you want to reduce the latency and power accessing data. DRAM can also be thought of as the cache of disc drives, why dont people use cheap disc dr
3.
▲
by
punkgenius
3y ago
Nowadays GPUs have sacrificed some performance for better programmability. ASICs always trade programmability for better performance and energy efficiency, it's really about how 'specific' you want it to be. I guess for appli
4.
▲
by
punkgenius
3y ago
Maybe 74 tflops is the best they've achieved, but not all 16 GPUs can consistently hit that number? Just guessing.. The 211 tokens/sec throughput on GPU is just insane, it's even better than what TPU can do on PaLM 540B.
5.
▲
by
punkgenius
3y ago
Looks like figure 8 of paper [1] says it is 18 tokens/s
6.
▲
by
punkgenius
3y ago
~200MB to 1GB per ASIC, from Table 2 on page 10.