Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
txyx303
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
txyx303
8mo ago
Those are scribe lines where you usually would cut out chips which is why it resembles multiple chips. However, they work with TSMC to etch across them.
2.
▲
by
txyx303
1y ago
feels very low compared to claude/gpt for me
3.
▲
by
txyx303
2y ago
I don't think they rely on SRAM very much for training. https://cerebras.ai/blog/the-complete-guide-to-scale-out-on-... outlines the memory architecture but it seems like they are able to keep most of the storage
4.
▲
by
txyx303
2y ago
Seems like they support training on a bunch of industry standard models. I think most of the customers in the training space tend to be for fine tuning right? The P and T in GPT stand for pre-trained - then you tune for your actual specific
5.
▲
by
txyx303
2y ago
afaik they have the current SOTA language models for arabic
6.
▲
by
txyx303
2y ago
MLPerf brings in exactly zero revenue. If they have sold every chip they can make for the next 2+ years, why would they be diverting resources to MLPerf benchmarking? Artificial analysis does good API provider inference benchmarking and has
7.
▲
by
txyx303
2y ago
Batched inference will increase your overall throughput, but each user will still be seeing the original throughput number. It's not necessarily a memory vs compute issue in the same way training is. It's more a function of the au
8.
▲
by
txyx303
3y ago
That was more of a WSE-1 problem maybe? They switched to a new compute paradigm (details on their site if you look up "weight streaming") where they basically store the activation on the wafer instead of the whole model. For somet