11 ms·
Cerebras’s giant chip will smash deep learning’s speed barrier
- Zenst 7y agoA chip that size, imagine the yield. Equally, cooling - has to be water based as a heatsink that size would be on par to a small anvil and the weight factor would be some serious issues. Though unsure as no pictures of it in-play alas and all they say is - "20 kilowatts being consumed by each blew out into the Silicon Valley streets through a hole cut into the wall", which does somewhat beg for a picture as just raises more questions. Why would they make a chip this big with AMD showing a chiplet design approach is cheaper and more scalable on so many levels. Let alone, yields. Equally, arms approach to utilising the back of the chip as a power delivery :- https://spectrum.ieee.org/nanoclast/semiconductors/design/arm-shows-backside-power-delivery-as-path-to-further-moores-law https://spectrum.ieee.org/nanoclast/semiconductors/design/ar... Then a wafer scale chip like this, using that approach, would save so much power. But again, yeilds will be a factor and can imagine this is not the cutting edge process node as you find as nodes mature, the yields improve. So an older node size would have a better yield and be more suitable for such wafer scale chips. But again, no mention of what is used. I have read in the past that it would use Intel's 10nm, but this article mentions TSMC. Another article I read that they used a 16nm node ( https://fuse.wikichip.org/news/3010/a-look-at-cerebras-wafer-scale-engine-half-square-foot-silicon-chip/ https://fuse.wikichip.org/news/3010/a-look-at-cerebras-wafer... ), which as mentioned above about node maturity, understandable.
- monocasa 7y ago> Why would they make a chip this big with AMD showing a chiplet design approach is cheaper and more scalable on so many levels. Let alone, yields. They're taking a radically different approach, and hoping that they'll be able to route around defects, unlike AMD where a defect in the uncore kills the whole chiplet.
- tedivm 7y agoA lot of the people involved in this actually come from AMD, so I imagine they're familiar with the issues AMD ran into.
- tedivm 7y agoI've seen a demo of the machine. It's about 17u in size, with the vast majority (like 15u) of that being for cooling. This was over two years ago so things may have changed. Right now I'm hosting some DGX's, and only one datacenter in the bay area had the ability to power a full rack of them. Power density is going to be a real issue for the these systems.
- Zenst 7y agoWow, that really does add some perspective upon the cooling and the aspect about power requirements datacenter wise really does highlight how out-there these type of systems are over the usual rack layouts. Equally, the cooling capacity of the datacenter comes into play with such systems. Given the power density, the amount of heat being generated would equally be above your normal rack output.
- tedivm 7y agoYeah- kind of tangental but it also plays along with how datacenters are transitioning from selling space to selling power. It used to be I'd just rent space by the rack or by the U, and then maybe pay extra for the network connection. Now the space itself is pretty cheap, and the network hookups are unbelievably cheap, but datacenters are actually paying attention to power consumption. In the case of the DGX-1 I've had datacenters tell me I couldn't put more than two in a rack. We ended up finding a datacenter the specialized in them (Colovore, who I can not recommend highly enough)- their power and cooling systems are some of the most impressive I've ever seen.
- luma 7y agoIn most cases the cooling capacity is in fact the actual limit you are running up against. Getting more power into a rack is a simple matter of running more cable. Getting more power _out_ of the rack is a much more complicated issue to resolve.
- Zenst 7y agoYes the whole getting more power into a datacenter is much easier to add than the extra cooling capacity to remove that power once it has transitioned into heat. But I'd imagine they would plan and monitor that aspect and may even have redundant cooling systems. But certainly a potential gotcha and one that would soon sort out the bad datacenters when they end up seeing all there hosting overheat and offline.
- deleted 7y ago[deleted]
- BooneJS 7y agoMany single-chip processors contain redundancy or ability to route around bad units. Yield isn’t an issue if it has programmable datapaths, even at this scale.
- gimmeThaBeet 7y agoI'm really curious about the benefits of their implementation. It's far beyond my grasp to make any serious criticisms and I don't really want to doubt them, it just seems a pretty radical departure from even the direction of innovation. The way they paint it sounds like they're putting in redundant cores to account for failure of what seems like what I would call the 'first line' cores, i.e. there's cores that are only used if some primary ones aren't working? But sort of intuitively that doesn't make a whole lot of sense given the parallel nature. Maybe they are just putting in 101% of specified cores, and if there's a ~1% hopefully uniform-ish core failure rate then it's all gucci? I guess my question is probably similar to yours, what are you giving up with yield-enhancing redundancy of a behemoth die vs integrating a bunch of confirmed working chiplets together?
- whatshisface 7y agoThe article claims that keeping everything on one die raises interconnect bandwidths and lowers latencies over what would be possible in a conventional supercomputing setup. Connections are made over the silicon that is normally left aside for cutting the chips apart. Apparently that is a special process that they had to collaborate with a partner in order to get working.
- phonon 7y agoThe CEO says 1-1.5%. "Cerebras approached the problem using redundancy by adding extra cores throughout the chip that would be used as backup in the event that an error appeared in that core’s neighborhood on the wafer. “You have to hold only 1%, 1.5% of these guys aside,” Feldman explained to me. Leaving extra cores allows the chip to essentially self-heal, routing around the lithography error and making a whole-wafer silicon chip viable." https://techcrunch.com/2019/08/19/the-five-technical-challenges-cerebras-overcame-in-building-the-first-trillion-transistor-chip/ https://techcrunch.com/2019/08/19/the-five-technical-challen...
- frankchn 7y agoChiplet designs means that you still have to route signals either onto an interposer or onto a PCB. If you have a silicon interposer you have the same issue of making a really large silicon die. If you route into the PCB, then you may need SerDes depending on what you do and bandwidth will be lower and latency will be higher due to signal integrity issues. Maybe something like Intel's EMIB technology where they have small interposers along edges of chips rather than having a giant interposer might help here. Yields are probably fairly good if they design for manufacturing by placing extra cores / wires to route around failures as I am sure they are.
- agoodthrowaway 7y agoThe future of these interconnects is to make them optical. Once the interconnects are optical lots of problems get solved. Chips don’t have to be in same enclosure, simplifying cooking etc.
- baybal2 7y agoI will dissent. Organic interposers are dirt cheap, and nearly as good unless all you want is density.
- ivalm 7y ago> A chip that size, imagine the yield From discussion at a demo the yield is good, since they are using a large node. Their hardware rerouting also mitigates defects on most chips.
- Accujack 7y ago>which does somewhat beg for a picture as just raises more questions. There's a picture in the article. >Why would they make a chip this big Did you read the article? >this article mentions TSMC. Another article I read that they used a 16nm node Yes, 16nm/TSMC.
- Zenst 7y ago> There's a picture in the article. Yes - hardly helpful ones as you get a picture of a wafer and a box, not breakdown beyond that - hence had look and found other articles with much more detail upon this that answers the questions I raised in relation to the lack of pictures - like the cooling aspect in which you snipped my quote and removed that lovely thing we call context. >Did you read the article? Yes and had you read what I said you would see that the article does not answer the aspects I was asking - see what you did there. >Yes, 16nm/TSMC Yes - I found that in another article - which I also linked, you're welcome.
- wolf550e 7y agoNot 100% of the chip is enabled, they disable defective parts and don't advertise a model that has 100% parts enabled, so they don't need magical zero defect wafers. Images of the whole computer were published, you can see the massive cooling system: https://www.tomshardware.com/news/worlds-largest-chip-gets-a-new-home-cerebras-launches-cs-1-system https://www.tomshardware.com/news/worlds-largest-chip-gets-a...
- 01100011 7y agoDid they come up with an architecture which can route around any defect? Probably not. Now, granted, 90% of their chip is probably dedicated to compute, but I'd bet there's some management infrastructure where they absolutely cannot tolerate a defect.
- kingosticks 7y agoThey'll simply have redundant copies of that logic. And they'll be physically located at areas of the wafer that yield well - some areas are much worse than others and I would imagine they'll make use of that.
- Zenst 7y agoInteresting so on a die, there are area's which are more prone to faults and they are able to factor that into the design? Though if there are known hotspots, wouldn't that point to the process node inducing them over silicon quality? Or is it a case of silicon production produces known hotspots that are predictable? FWIW, I'm currently learning towards process node over the silicon being the source of hotspots, given what I know about silicon production.
- kingosticks 7y agoWith normal-sized dies, at the die-level, I've not seen people design around this; other than the more obvious places e.g. the corners (bad power delivery, prone to mechanical issues, normally left vacant) and the middle (gets hotter, also sometimes bad power). However, there are many test structures placed across the die to measure/check variations and also design rules that constrain the relative placements of certain things. That also goes towards increasing yield. But at the wafer-level, yes. > wouldn't that point to the process node inducing them over silicon quality? I don't see why. I would only vaguely guess it's related to the manufacturing process they follow at that particular node. Maybe it's not even directly silicon related but something else. I'm not convinced it's worthwhile separating out the process node and the silicon quality, they are entwined when looking across a large sample size. Unfortunately, someone that actually knows why probably isn't allowed to share why.
- giacaglia 7y agoI've wrote about the challenges that Cerebras went through and what is next: https://towardsdatascience.com/why-cerebras-announcement-is-a-big-deal-6c8633ffc49c https://towardsdatascience.com/why-cerebras-announcement-is-...
- bcatanzaro 7y agoReminds me of that other great prediction of a GPU killer from IEEE Spectrum back in 2009: https://spectrum.ieee.org/computing/software/winner-multicore-made-simple https://spectrum.ieee.org/computing/software/winner-multicor...
- ajtulloch 7y agoFor the folks who are downvoting this comment, the author is absolutely a subject matter expert (and completely correct).
- deepnotderp 7y agoBut he also works at nVidia and Larrabee versus the WSE are two entirely different things. Larrabee was an architectural approach to more general purpose parallel hardware whereas the WSE is a more special purpose and physically different than a GPU.
- Google234 7y agoWhat did go wrong with Intel's MIC (Xeon Phi) project? I can't find a compressive account of this from HPC people. The idea seemed pretty sound: large die, simpler circuit, and much more parallelism, in the x86 line..
- raphlinus 7y agoYou'll probably find Tom Forsyth's blog on this to be interesting reading: https://tomforsyth1000.github.io/blog.wiki.html#%5B%5BWhy%20didn't%20Larrabee%20fail%3F%5D%5D https://tomforsyth1000.github.io/blog.wiki.html#%5B%5BWhy%20...
- liuliu 7y agoI vaguely remember that at the dawn of the deep learning (2013 to 2014), there were talks and hopes that Xeon Phi would smash the performance of Nvidia GPUs. However, the sample people got are too late (I believe it is at the end of 2014) and the performance figures are disappointing. It might be just the software was simply not there yet unfortunately. But then the wheels moved forward and everyone started to buy Nvidia chips in their datacenters.
- deleted 7y ago[deleted]
- gfodor 7y agoI’m a know-nothing when it comes to this area, but I shouted expletives at least twice when I read this article. This is crazy.
- green-eclipse 7y agoThe Cerebras chip really stands out in terms of the chip industry's relationship to Moore's law. Look at the graphs in this article for reference: https://medium.com/predict/cerebras-trounces-moores-law-with-first-working-wafer-scale-chip-70b712d676d0 https://medium.com/predict/cerebras-trounces-moores-law-with...
- ThrowawayR2 7y agoThat article is utter balderdash. Yes, it's obvious that you can fit more transistors on a "chip" if you make the chip be much, much larger than what we ordinarily think of as a chip. No, it does not mean that Moore's Law has been invalidated or some new "AI Moore’s Law" (quoting from the post) has come into being.
- dnautics 7y ago> Yes, it's obvious that you can fit more transistors on a "chip" if you make the chip be much, much larger than what we ordinarily think of as a chip. Without defending the article, it is however the case that simply scaling a chip size has nontrivial problems. For example, Will the piece of silicon warp or shatter if one side happens to get hotter than the other?
- ThrowawayR2 7y agoPossibly. Wafer scale integration has been investigated before though and there were even a couple of attempts at commercial products; it's not a brand new technology. Nevertheless, it might be interesting to examine Cerebras' patents to see if anything of significance relating to WSI is there.
- atq2119 7y agoThat article is hogwash. Sure, the Cerebras "chip" is impressive. But the idea that it will accelerate Moore's law and usher in the singularity is just nonsense. Nobody has even made serious efforts to use deep learning for physical design, and its scope for improving designs is limited at best even in theory. If this was trying to aim at solid state physics and materials research, then maybe one could be carefully optimistic about a genuine breakthrough via something like room temperature, standard pressure super-conducting. As it stands, I call blind hype.
- wbhart 7y agoFrom the perspective of an outsider, I can't see how a company like this could survive. They claim on the one hand to have done something really amazing and are at the stage where they are looking for customers. Normally, you'd expect them to be touting performance figures to secure such investment. Instead, they've decided to keep the performance secret. And they've managed to find some "expert" who says this is normal. Does anyone here have expertise in this area? Is this the model for a successful company in this area?
- tedivm 7y agoI got a demo of this two years ago, and honestly I don't think it matters that they aren't sharing these numbers. Any company that is going to consider this is going to want to benchmark it on their own models and systems, and as long as Cerebras allows that they aren't going to have trouble finding customers (assuming their claims line up with reality). Even if that doesn't work out most of the people on these time have built companies that were acquired by either AMD or another chip maker.
- justicezyx 7y agoMass market customers are just going to skip without benchmarks. Although, at this stage, Crebras does not care about mass market yet.
- deleted 7y ago[deleted]
- phonon 7y agoThere are probably only a few hundred prospective customers. (Some may buy several units). Each unit will cost millions. They can discuss the expected workloads/performance with each prospective customer individually.
- jandrese 7y agoKeeping the performance figures a secret is a red flag on the level of "run, don't walk, away from this company". At best their solution is on par with GPUs in a performance per watt/dollar sense. At worst they're scammers looking for a sucker.
- m0zg 7y agoThey did build some valuable tech, no question there, but be sure to account for the typical startup hyperbole. By the time you can get your hands on this (if that ever happens), the hyperbole will converge a bit closer to reality, the tradeoffs will become apparent, etc, and you'll discover that it is not, in fact, going to "smash" barriers of any kind in any practical sense. From TFA: "Cerebras hasn’t released MLPerf results or any other independently verifiable apples-to-apples comparisons." That's all you really need to know.
- mark_l_watson 7y agoI don’t know if this mega-chip will be successful, but I like the idea. Before I retired I managed a deep learning team that had a very cool internal product for running distributed TensorFlow. Now in retirement I get by with a single 1070 GPU for experiments - not bad but having something much cheaper, much more memory, and much faster would help so much. I tend to be optimistic, so take my prediction with a grain of salt: I bet within 7 or 8 years there will be an inexpensive device that will blow away what we have now. There are so many applications for much larger end to end models that will but pressure on the market for something much better than what we have now. BTW, the ability to efficiently run models on my new iPhone 11 Pro is impressive and I have to wonder if the market for super fast hardware for training models might match the smartphone market. For this to happen, we need a deep learning rules the world shift. BTW, off topic, but I don’t think deep learning gets us to AGI.
- corporate_shi11 7y agoIt's also my impression - from my modest exposure to DL over the past two years as a student taking courses - that deep learning must be overcome to reach AGI. Specifically gradient descent is a post hoc approach to network tuning, while human neural connections are reinforced simultaneously as they fire together. The post hoc approach restricts the scope of the latent representations a network learns because such representations must serve a specific purpose (descending the gradient), while the human mind works by generating representations spontaneously at multiple levels of abstraction without any specific or immediate purpose in mind. I believe the brain's ability to spontaneously generate latent representations capable of interacting with one another in a shared latent space is functionally enabled by the paradigm of neurons 'firing and wiring' together. I also believe it is the brain's ability to spontaneously generate hierarchically abstract representations in a shared space that is the key to AGI. We must therefore move away from gradient descent.
- mantap 7y agoDon't forget the human brain takes about 7 to 8 hours off every day to rejiggle itself, to use a scientific term. The brain's architecture is better than having a training stage but it's by no means able to continually learn without stops and starts.
- michelpp 7y agoThe members of the GraphBLAS forum have discussed this chip a couple of times. There's a lot of research on making deep neural networks more sparse, not just by pruning a dense matrix, but by starting with a sparse matrix structure de novo. Lincoln Laboratory's Dr. Jeremy Kepner has a good paper on Radix-Net mixed radix topologies that achieve good learning ability but with far fewer neurons and memory requirements. Cited in the paper was a network constructed with these techniques that simulated the size and sparsity of the human brain: https://arxiv.org/pdf/1905.00416.pdf https://arxiv.org/pdf/1905.00416.pdf It would be cool to see the GraphBLAS API ported to this chip, which from what I can tell comes with sparse matrix processing units. As networks become bigger, deeper, but sparser, a chip like this will have some demonstrable advantages over dense numeric processors like GPUs.
- rsp1984 7y agoThis fits perfectly into the narrative of yesterday's discussion on HN [1]. Deep Neural Nets are somewhat of a brute force approach to machine learning. Training efficiency is horrible as compared with other ML approaches, but hey, as long as we can trade +5% of classification performance for +500% of NN complexity and throw more money at the problem, who cares? I see a dystopian future where much better and much more efficient approaches to ML exist, but nobody's paying attention because we have Deep Neural Nets in hardware and decades of infrastructure supporting it. [1] https://news.ycombinator.com/item?id=21929709 https://news.ycombinator.com/item?id=21929709
- justicezyx 7y agoWell, if a better algorithm cannot beat DNN in a realistic product setting, then how can you say its better after all? If the algorithm is indeed better, how can DNN dominates and turn into a dystonia...
- someguyorother 7y agoWhat economists call path dependence. The alternative algorithm would be better than DNN if the same amount of effort was put into creating special-purpose hardware, libraries, and so on; but in the dystopia, it's not fully refined DNN vs fully refined alternative algorithm, but fully refined DNN vs alternative algorithm with hardware and software optimized for DNN. The alternative algorithm always looks unappealing because the playing field historically favors DNN, and so doesn't take off in the dystopia.
- sbierwagen 7y agoOne example would be ternary logic, which more efficiently represents numbers: https://en.wikipedia.org/wiki/Radix_economy https://en.wikipedia.org/wiki/Radix_economy
- justicezyx 7y agoYou are referring back to OP's own reasoning fallacy... DNN emerges out from being an underdog. Its superiority was proven by technology and economy. What you said is of course not wrong, but they can never be proven right. As immediately you switch the role, your argument then favors the other one.
- ZhuanXia 7y agoThem shunning benchmarks is pretty lame.
- baybal2 7y agoThe guy who runs Cerebas has history of quick selling companies that then went nowhere. He bets all on wow-effect, and sells to trend chasing suckers. Less than stellar benchmarks will ruin the "magic"
- varelse 7y agoI am far more excited by the underlying Wafer Scale Integration moonshot than I am by any AI benchmarks here. I know it's trendy to think there can only be one w/r to the AI Iron Throne but nope, not the case, everyone is writing bespoke code in production where the money is made. Well, almost everyone, Amazon seems to be the odd duck but they're a bunch of cheapskate thought leaders anyway (except for their offers to junior engineers in their desperate hail mary attempt to catch up with FAIR and DeepMind, but... I... digress...). Which is to say that graphs written to run specifically on Cerebras's giant chip will smash deep learning's speed barrier for graphs written to run best on Cerebras's giant chip. And that's great, but it won't be every graph, there is no free lunch. Hear me now, believe me later(tm). But if we can cut the cost of interconnect by putting a figurative datacenter's worth of processors on a chip, that's genuinely interesting, and it has applications far beyond the multiplies and adds of AI. But be very wary of anyone wielding the term "sparse" for it is a massively overloaded definition and every single one of those definitions is a beautiful and unique snowflake w/r to efficient execution on bespoke HW.
- foota 7y agoIsn't that similar to what AMD is doing with infinity fabric? Obviously not at such a large scale.
- jamesblonde 7y agoInfinity fabric is closer to Nvidia's NVLink - much lower interconnect B/W. PCI 4.0 will be interesting as a commodity alternative, particularly when paired with AMD Rome chips with huge numbers of I/O lanes - for distributed training. https://wccftech.com/amd-epyc-rome-zen-2-7nm-server-cpu-162-pcie-gen-4-lanes-report/ https://wccftech.com/amd-epyc-rome-zen-2-7nm-server-cpu-162-...
- LASR 7y agoThis is pretty much on the same level as a raw material change that happened when CPUs went from a bunch of DIP ICs connected together to a single silicon die. LSI was a big leap in optimizing compute/$. I had the same “hold the phone” reaction when I reading about AI but then saw the whole wafer in the marketing photo.
- geomark 7y agoThe article talks about a few things that they call inventions, like making interconnections across what would normally be scribe lines. But I personally worked on wafer scale integration about 25 years ago and we were already doing that. We called it inter-reticle stitching. The technology was ancient back then - 0.5 micron feature size on 4 inch wafers - but the wafer scale techniques are applicable to modern technologies. In particular, developing a yield model that informs your on-chip redundancy choices and designing built-in self test and selection circuitry so that you can yield large chips. The chip we developed was so large that only two would fit on a wafer. We got 50% yield on a line that was far from mature at the time. The company lacked the vision to do anything with what they had developed. To them it was just a chip for which there were few customers. The suits didn't know how to make bank with this methodology that could yield nearly arbitrarily complex chips in nearly any target process. Edit: There were a number of papers and conference proceedings published back then but not much shows up when searching Google. Here's one discussing the issues and results of field stitching https://fdocuments.in/document/ieee-comput-soc-press-1992-international-conference-on-wafer-scale-integration-589dfe172703a.html https://fdocuments.in/document/ieee-comput-soc-press-1992-in... From 1992, so yeah, field stitching is not a recent invention.
- DaniFong 7y agoGreat post, but I would like to add that the critical question for whether an invention because a useful innovation is usually not "is this novel" but rather "is there a currently viable project here with people who care about the thing and genuine motivation and persistence and adequate resources." In other words, "how is this effort new to the universe?" I would say it's certainly at a different scale and a different time. And we should be super thankful that the commercial interest is such that we can try out new chip designs in a different domain now; you can really imagine a rethink for the kinds of things that are possible once you're really at scale here.