5 ms·
In their whitepaper they claim "with all model parameters in on-chip memory, all of the time", yet that entire 15 kW monster has only 18 GB of memory. Given th
by bmh 7y ago
In their whitepaper they claim "with all model parameters in on-chip memory, all of the time", yet that entire 15 kW monster has only 18 GB of memory.
Given the memory vs compute numbers that you see in Nvidia cards, this seems strangely low.
- Veedrac 7y ago18GB is huge! An NVIDIA V100 has 6MB of L2 memory. HBM is off-chip, and vastly (~100x) slower.
- baybal2 7y ago18GB of very fast memory will still be just as hard to keep fed with data as that 6MB cache
- Veedrac 7y agoThe idea is that the whole model resides in the fast memory, so you don't need to ‘keep it fed’.
- dmitrygr 7y ago44K/core is very little memory
- Veedrac 7y agoIndeed, but cores are only responsible for small fragments of the network, so don't need huge amounts of memory.
- dmitrygr 7y agoUnless you need to multiply large matrices, where you need access to very large rows and columns...like in...ML applications
- Veedrac 7y agoThat's what the absurdly fast interconnect is for. You send the data to where the weights are.
- dmitrygr 7y agoAbsurdly fast != Single cycle It will be physically impossible to access that much memory single cycle at anything approaching reasonable speeds. I suppose you could do it at 5Hz :)
- Veedrac 7y agoA core receives data over the interconnect. It uses its fast memory and local compute to do its part of the matrix multiplication. It streams the results back out when it's done. The interconnect doesn't give you single-cycle access to the whole memory pool, but it doesn't need to.
- dmitrygr 7y agoI think it is telling that in one sentence there is a claim that it is faster than nvidia, and in another, a claim it does tensor flow. I do not think this architecture could do both of these at once. It could not do tensor flow fast enough (not enough local fast mem) to compete even with a moderate array of GPUs
- deepnotderp 7y agoHey! You seem knowledgeable, mind emailing me at tapabrata_ghosh [at] vathys (dot) ai ?
- bmh 7y agoThat's true, but it doesn't match their claim of keeping all of the model on the chip. An 18 GB cache is huge for sure, but that's not what they claim.
- baybal2 7y agoCheck the C-suite track record... a pattern of making quick selling companies on a "wow effect" which then quickly turn defunct and valueless after sale. There were big red flags about Cerebras claims for a quite some time. Some say it is the Graphcore on steroids. Not so much about the tech side, stuff like that been tried before (without good results: the more is your reticle fill, the poorer is the exposure,) but the business wise side of their claims don't make sense. First, and the biggest one being the economics. It is completely impossible that a run as small as 100, or even 1000 wafers be more economical than that of mass market product, even if you deduct packaging costs. On top of that, just any process modification or "tweak" for a low volume run will destroy just any economy of scale. And as I understood, they pretty much brag about doing so. Lastly, some tech notes. Maybe they got the issue solved, maybe not: the bigger the chip, the more memory starved it is for a simple reason of geometry. With the "chip" the size of a wafer, it is gonna be extremely memory starved unless it has more IO than computing devices. Then, the thermal ceiling for CMOS is around 100W per cm², and it is a very hard limit. I see no point why they brag about beating it when they truly didn't: 20 watt per square cm² is quite low for HPC. I suspect they are indeed quite limited by thermals, if they had to backpedal on their original claims.
- simg 7y agowhat's wrong with Graphcore?
- blihp 7y agoThat's 18GB of static RAM accessible in one clock cycle... the memory on a GPU isn't in the same class of fast. Given the bandwidth and latency of this thing, you'd likely have to use a cluster of machines doing all sorts of pre-, post- and I/O processing just to keep this thing busy.