3 ms·
70b runs on 4x CS-3 estimated at $2-3m each, let’s say total system cost $10m, drawing ~100kW power. They don’t mention batch size, so let’s start with batch si
by moconnor 2y ago
70b runs on 4x CS-3 estimated at $2-3m each, let’s say total system cost $10m, drawing ~100kW power. They don’t mention batch size, so let’s start with batch size 1 and see where we get. At 100% utilisation for 3 years that’d be 42 billion tokens for a cost of $10m capital plus ~$0.5m power and cooling let’s say, or $250 per million tokens. They’re claiming they can sell their API access at $0.60/million. To break even on this they’d need a batch size of 420. I don’t know how deep their pipeline is but Llama 3.1 70b has 80 layers with 6 meaningful matmuls per layer so it’s not a crazy multiple of that.
A single A100 processes at 13 t/s/u for batch 32. That costs $10k to process 39 billion tokens over 3 years = $0.25 tok/s/u. If you have batch size 420 you can do it even cheaper.
TL;DR: Cerebras are certainly advertising at a loss-leading price and will only have a viable product if they can get extraordinarily high utilisation of their system at this price. I don’t think they can, so they’re basically screwed selling tokens. Maybe this is to attract attention in the hope of selling hardware to someone willing to pay a premium for very low latency, but I suspect it’s just a means of getting one more round of funding in the hope of reducing costs in the next version.
- amirhirsch 2y agoyour estimate of $2-3M per CS-3 is a price not a cost. It costs about $20K per wafer from TSMC and the price they charge reflects the NRE of designing the system and taping out the masks more than the additional costs to package up their wafers into a system. If this business scales, they can probably afford to lower the price by a factor of 10.
- mikewarot 2y agoMy math (and google) shows a 300 mm diameter wafer, and 300,000,000 transistors/mm^2 So for $20,000 (in quantity) you get somewhere around 10 trillion transistors? That's enough for about 50 4096x4096 multiply-accumulate chips. At a nice slow 1 Mhz clock rate, each would take about 3.5 watts, and give you 16 teraflops of performance. If you stepped up the power and cooling, you could likely get to 350 watts of power at 100 Mhz, and 1.6 Petaflops. 50 of those chips, for $20,000 --> $400 each
- jwan584 2y agobatch size by Q4 will be solid double digits (cerebras employee)
- moconnor 2y agoIs that e.g. batch 16/32 for each operation e.g. 16-row matmuls in a pipeline? Or a pipeline of vector-math ops that has 16/32 stages? Is the pipeline also double digits deep?