4 ms·
How is this even possible?
by supernova8 1y ago
How is this even possible?
- kristopolous 1y agoIt's their own hardware : https://www.cerebras.ai/blog/cerebras-cs3 https://www.cerebras.ai/blog/cerebras-cs3
- unshavedyak 1y agoIncase i'm missing something, why wouldn't it be possible? Claude and Gemini have similar offerings for a similar/same price, i thought. Eg if Claude Code can do it for $200/m, why can't Cerebras? (honest question, trying to understand the challenge for Cerebras that you're pointing to) edit: Maybe it's the speed? 2k tokens/s sounds... fast, much faster than Claude. Is that what you're referring to?
- UnPerson-Alpha2 1y agoHe just wrote another way of making an exclamation, like "wow, incredible!".
- meepmorp 1y agoThey make frisbee-sized CPUs.
- sliken 1y agoIndeed. Pretty much all silicon today comes on 12" or so wafers, broken into chip sized pieces, and each chip is tested and the ones that failed are thrown away. Cerebras uses the entire 12" and builds in redundancy so that with current defect rates a large fraction of the wafers are usable. This allows a huge level of parallelism, a large amount of on board ram, and the removal of the need to move data on/off the wafer. So the available bandwidth is insane and inference is mostly bandwidth limited.