8 ms·
New Intel Instructions for Alder Lake, Also BF16 for Sapphire Rapids
- chx 7y agoThis doesn't mean ADL will actually see a wide release. I am still dead convinced we will never see a wide release of a 10nm desktop chip. It will be delayed then cancelled in favor of 7nm. Server wise, same, limited releases, more smoke and mirrors etc etc Intel will limp along with 14nm until 7nm. 10nm is utter broken and they can't fix it.
- reitzensteinm 7y agoI think that's true if 7nm progress is independent of 10nm, and the mistakes made with 10nm aren't also delaying 7nm. My understanding is that Intel's 14->10nm shrink was the most aggressive in the industry, promising to yield a greater increase in density than the geometry would imply, when usually there's a loss factor and the density increase isn't as good as you'd naively expect. Even after 10nm was delayed by quite some time, Intel pointed to this and declared they weren't lagging the industry as the shrink was closer to a 1.5 node shrink. If the 10nm delay was the result of this aggressiveness and Intel was well in to developing 7nm in a similar fashion before it became obvious how 10nm was going to turn out, this may not be a matter of skipping over a single bad apple.
- chx 7y ago> I think that's true if 7nm progress is independent of 10nm, and the mistakes made with 10nm aren't also delaying 7nm That is exactly what's happening
- keanebean86 7y ago7nm was planned for 2017 as of 2014. Ouch. https://www.tweaktown.com/news/41582/intel-to-hit-10nm-in-2016-with-7nm-cpus-arriving-in-2018/index.html https://www.tweaktown.com/news/41582/intel-to-hit-10nm-in-20...
- chx 7y agoThat's just early optimism, as far as these things can be known from the outside, the process is fine very much unlike 10nm which was known to be broken for years even before Cannon Lake.
- justAnotherNET 7y ago7nm intel is EUV, 10nm intel is quad patterned DUV. Massive difference.
- scottlocklin 7y agoIs there a EILI5 on why Intel borked 10nm and why others apparently didn't?
- chx 7y agoELI5? Not really. 1. The metal pitch targeted (36nm) requires Self-Aligned Quad Patterning (SAQP) apparently is very very hard to do without EUV. TSMC is doing 40nm metal pitch and Samsung does 36nm BUT both of them are using EUV. 2. Cobalt https://fuse.wikichip.org/news/525/iedm-2017-isscc-2018-intels-10nm-switching-to-cobalt-interconnects/ https://fuse.wikichip.org/news/525/iedm-2017-isscc-2018-inte... 3. COAG https://www.reddit.com/r/AMD_Stock/comments/dk2pw7/why_it_is_so_difficult_for_intel_to_catch_up_the/f4ml1k4 https://www.reddit.com/r/AMD_Stock/comments/dk2pw7/why_it_is...
- Symmetry 7y agoI think the ELI5 is "Intel tried to do some things that were cleverer than they could actually do and messed up."
- scottlocklin 7y agoI remember them developing an EUV lithography rig at the ALS 11 bend magnet in the 90s... guess it never worked out. Looking at the website: dude's still there! Thanks for the explanation; close enough to ELI5
- throw0101a 7y agoPossibly a meta question on instructions: is there such a thing as having 'too many'? I know transistors are cheap, but at what point is it diminishing returns on new instructions? How many different use cases need to be handled? Certainly there are new situations that need to be handled (e.g., H.264/5/6, AV1 video), but will there ever be a point where we can say "this is enough"?
- vbezhenar 7y agoWhen you can put more transistors, you have to use them for something. You can either put more cores, but that requires for software to be written in a way to utilize those cores and that's not always easy task or even possible. Or you can use those transistors to make some common operations faster and that potentially will increase performance with little rewrites (or even with no rewrites for libraries). "This is enough" will be said when we wouldn't be able to put more transistors. I think that it won't happen in the next 10 years, but, of course, it'll happen eventually.
- throw0101a 7y agoCaching? More I/O? I would think that the speed that which most cores and instructions work at is often rarely the bottleneck nowadays for a good portion of workloads.
- derefr 7y ago> When you can put more transistors, you have to use them for something. I mean, it’d honestly be really interesting to know the performance charateristics (TDP et al) of an Intel 8088 if it were shrunk to 14nm and then driven at a modern clock speed. Maybe less is more?
- JoeAltmaier 7y agoHa, no. It didn't have cache, pipelining, separate bus and instruction subsystems, a wide bus. It would run one instruction every N clock cycles instead of N every 1 clock cycle, and hang on every memory read or write. Multi-byte math would take many instructions. It would be pathetic.
- throw0101a 7y agoIs there a rationale behind Intel's codename scheme? With (e.g.) Ubuntu they go through the English alphabet, so you can get some idea of the order things have/will come out, but it seems that Intel is using a random word generator for their X Lake names.
- auvi 7y agoI think all the X Lake names are taken form real lakes in the state of Oregon.
- throw0101a 7y agoSure, but how are the names chosen? If Ubuntu et al it is alphabetical order, and you can generally tell the timeline of releases. But how can one really tell what the current product from Intel is, and what came before, and what are the upcoming releases? It's not like they're going through Oregon lakes in some kind of order, or are they? By discovery, by size/volume, other?
- barkingcat 7y agoAs in most code names there is no order, if you want an order, maybe in terms of favourite to least favourite parks/lakes of the Intel staff. What you are looking for are monotonically increasing version numbers, which code names are definitely not. If you are looking for details about generations of chips, https://ark.intel.com/ https://ark.intel.com/ is your friend.
- throw0101a 7y ago> What you are looking for are monotonically increasing version numbers, which code names are definitely not. They are not, but given that Intel seems to put them in their public information / marketing material, it seems like the "codenames" are being used as version numbers. If they were strictly internal-to-Intel I could see that POV, but that doesn't seem to be happening. And the Intel's model numbers / SKUs also seem to be created from a random number generator. :)
- throwaway_pdp09 7y agoThe article says "LBRs (Last Branch Recording) in order to speed up branches". Reading the PDF seems to be more about recording it (for profiling?). If it really is about speedups can someone summarise how, and how that differs from using branch predictors.
- chrisseaton 7y agoI think it's for tooling like profilers, not directly for performance. But you could use that tool to improve performance in your code, so maybe that's what they mean.
- zbjornson 7y agoThat's correct. Info on LBR in general: https://lwn.net/Articles/680985/ https://lwn.net/Articles/680985/
- gok 7y agoIs it clear to anyone if the BFloat16 support in these chips means they have extra-low precision multipliers internally, or are the dot products still implemented with the single precision guts? Does it actually increase compute over FP32 or is it just a bandwidth win?
- jcranmer 7y agoThe new instructions are, I believe: * VCVTNE2PS2BF16 — Convert Two Packed Single Data to One Packed BF16 Data [i.e., float32 -> bfloat16 conversion] * VCVTNEPS2BF16 — Convert Packed Single Data to Packed BF16 Data [ditto] * VDPBF16PS — Dot Product of BF16 Pairs Accumulated into Packed Single Precision [i.e., multiply two v2bf16 with each other, and then sum the two results into a f32.] The last instruction is the only computation instruction, and that can be implemented by breaking the 23-bit multiplier of an FMA unit into two 7-bit multipliers, one for each bfloat16, as well as duplicating the normalization logic beforehand.