3 ms·
Wow. It has 15.3 billion transistors. It's amazing we can buy something with that many engineered parts. Even if the transistors are the result of duplication a
by tomkinstinch 10y ago
Wow. It has 15.3 billion transistors. It's amazing we can buy something with that many engineered parts. Even if the transistors are the result of duplication and lithography, it's an astonishing number. Creating the mask must have taken a while.
Does anyone know what the failure rate is for the transistors (or transistors of a similar production process)? Do they all have to function to spec for a GPU, or are malfunctioning transistors disabled or corrected? What does the QC process look like?
- djcapelis 10y agoExact failure and bin rates for most semiconductor companies are considered deep dark internal trade secret. Other than pure scale, yield rates are one of the biggest factors in semiconductor cost and profit margin. And the answer is it depends. If you lose some transistors, you expect to lose the entire chip. But the vast majority of the transistors on each chip are part of cache or many many duplicate GPU cores, which if they fail to pass tests, can be disabled or downclocked and then the chip is binned into the appropriate product line. With GPUs this is much easier than other types of chips, because the level of functional duplication that exists allows a lot of flexibility. If a core is bad, you use a different one, and GPU cores are small enough they'd be stupid not to put some spares on each chip. Same with memories. Generally one can safely assume: * Most chips that come off the line are binned into a lower category and do not function at max spec for everything, which is why the price jump is so high at the extreme upper end of a hardware series. * With ASIC lithography most transistor malfunction isn't correctable, you mostly have to either downclock (some types of faults) or disable (the rest) that piece. * Rates of transistor malfunction is still incredibly fucking amazingly phenomenally low. Like with 15B transistors on a chip, you have trouble affording a failure rate of even one in a billion. So your line has to be, as the kids say: on fleek.
- Salgat 10y agoIs there an approximation you can give, such as a magnitude? Is it around 1 transistor per million, per billion, per thousand?
- djcapelis 10y agoI don't have direct knowledge of current yield rates, so this is speculation. That said, I did give what I think is a reasonable order of magnitude in the above comment when I said a company would have trouble affording a rate of 1 in a billion transistors for unclustered defects. I meant that literally. Some chips just probably can't be profitable with a rate that high. Nvidia might be able to make yield on that rate since GPUs have enough functional duplication, but I'd expect Intel's rate to be under one in a billion and over one in a trillion. Also note that clustered failures are different. Some whole wafers might be junk if alignment is off, or if there's a bubble somewhere, a series of adjacent chips would be destroyed. If you throw a whole wafer away, that thankfully doesn't mean you have to produce a billion perfect wafers to make up for it. So the yield rate above would only need to apply for the parts where the process is otherwise dialed in. If you have a bubble or alignment issue, it really doesn't matter at that point whether you kill tens of transistors or a few billion, any chips where the bubble is are likely just gonna get marked and tossed. And it is sometimes routine on some types of lines that if the chip yield is low enough on a wafer the whole thing is tossed since it's not economical to cut and further test and package up any working ones. Semiconductor economics are pretty nuts. It's actually more common to specify error rate in terms of defects per mm^2, because the exact number of transistors involved in the defect is mostly irrelevant if it's a defect that is wide ranging.
- ckozlowski 10y agoSpot on. I used to work for an OEM, and the Intel and AMD engineers would quietly explain to us how this worked on a number of occasions. The AMD X3 chips I think were the best example of this being done. These were quad-core parts that AMD was manufacturing at the time, but had defect that made one core faulty. So that core was disabled, and sold as triple-core part. http://www.zdnet.com/article/why-amds-triple-core-phenom-is-a-bigger-deal-than-you-think/ http://www.zdnet.com/article/why-amds-triple-core-phenom-is-...
- mud_dauber 10y agoI would expect that partially defective chips are repaired during probe or (more likely) final test by blowing fuses. The chip's yields, and therefore cost, will depend on the foundry's natural defect rate per area and the design quality.
- Retric 10y agoAn interesting note is processes improve over time, so most companies end up binning processors much more conservatively over time. Price discrimination means company's want to restrict their highest bin even if they could sell it for less. https://en.wikipedia.org/wiki/Price_discrimination https://en.wikipedia.org/wiki/Price_discrimination
- tcas 10y agoI do not have the answers for your questions (and I don't think anyone can share actual failure rates), but I would direct you this video which goes over a lot of modern chip fabrication techniques, circa 2009: https://www.youtube.com/watch?v=NGFhc8R_uO4 https://www.youtube.com/watch?v=NGFhc8R_uO4 It's crazy stuff. There are wafer test machines which will interface with the wafer directly and do some testing (which are $$$$), JTAG type tests, which access parts of the chip out of band, and functional testing. Some products, like SD Cards actually have a microcontroller on board that will provide the test routines and error correction without the need of an expensive machine. Design for test is extremely important. I'm by no means an expert however, I mostly deal with JTAG and functional tests.
- cottonseed 10y agoI just read that the Xilinx XCVU440 FPGA has >20B transistors, and that's one generation old (20nm, UltraScale+ is on 16nm finfet). Insane.