7 ms·
Everybody seems to view this as AMD mimicing Intel when it acquired Altera. (That acquisition has not born visible fruit.) My contrarian speculation is that th
by gvb 6y ago
Everybody seems to view this as AMD mimicing Intel when it acquired Altera. (That acquisition has not born visible fruit.)
My contrarian speculation is that this is a move driven by Xilinx vs. Nvidia given Nvidia’s purchase of Arm and Xilinx’ push into AI/ML. Xilinx is threatened by Nvidia’s move given their dependence on Arm processors in their SOC chips and their ongoing fight in the AI/ML (including autonomous vehicles) product space. My speculation is that this gives Xilinx an alternative high performance AMD64 (and possibly lower performance & lower power x86) "hard cores" to displace the Arm cores.
Interesting times.
- gmueckl 6y agoWhy not license other softcores, e.g. from SiFive?
- gvb 6y agoThe performance of soft cores are significantly lower than hard cores. Xilinx already has a RISC soft core in their MicroBlaze architecture so they don't have a pressing need for a low power, reasonable performance RISC soft core. Ref: https://en.wikipedia.org/wiki/MicroBlaze https://en.wikipedia.org/wiki/MicroBlaze AMD has high performance CPUs being fabbed by TSMC (same foundry as Xilinx), so (theoretically) AMD CPUs can be grafted onto the Xilinx FPGA as a hard core. With AMD and the MicroBlaze, they have the high performance and low power processor spectrum covered with no need for 3rd party licensing costs.
- brandmeyer 6y agoFPGAs are to ASICs as interpreted languages are to compiled languages. I don't mean that literally, but I do mean it in the performance sense. At the same process node, an FPGA is over 50x the power and 1/20th the speed of a dedicated ASIC and it isn't getting any better.
- dragontamer 6y agoThat's a decent analogy. Note however that Xilinx has a dsp slice (UltraScale) which is a prefabbed adder / multiplier. This would be PyTorch in the analogy. FPGA LUTs cannot compete against ASICs, so modern FPGAs have thousands of dedicated multipliers to compete. The LUTs compete against software, while the dedicated 'Ultrascale DSP48 Slices' competes against GPUs or Tensors. -------- It's not easy, and it's not cheap. But those UltraScale DSP48 units are competitive vs GPUs. It's still my opinion that GPUs win in most cases, due to being more software based and easier to understand. It is also cheaper to make a GPU. But I can see the argument for Xilinx FPGAs if the problem is just right...
- TomVDB 6y agoIn my experience, the 20x performance number is after taking the DSPs into account.
- dragontamer 6y agoJust looking at raw FLOPs: the 7nm Xilinx Versal series tops out at 8 32-bit TFlops (DSP Cores only), plus whatever the CPU-core and LUTs can do (but I assume CPU-core is for management, and LUTs are for routing and not dense compute). In contrast: the NVidia A100 has 19 32-bit TFlops. Higher than the Xilinx chip, but the Xilinx chip is still within an order of magnitude, and has the benefits of the LUTs still. ----- It should be noted that Xilinx Versal "AI engine" is a VLIW SIMD-architecture: https://www.xilinx.com/support/documentation/white_papers/wp506-ai-engine.pdf https://www.xilinx.com/support/documentation/white_papers/wp..., effectively an ASIC-GPU hardwired into the FPGA.
- davrosthedalek 6y agoThe question is, how much does your algorithm get from the 19 TFlops for a GPU, and how much from the 8 from the Versal. I'm sure many algos fit GPUs fine, but some don't, and might get more out of an FPGA.
- banjo_milkman 6y agoRaw FLOPs is completely misleading, which is why Nvidia focus on it as a metric. The GPU can't keep those ops active - particularly during inference when most of the data is fresh so caches don't help. It's the roofline model. In my experience FPGA>GPU for inference, if you have people who can implement good FPGA designs. And inference is more common than training. Much of this is due to explicit memory management and more memory on FPGA.
- dnautics 6y agoI kind of love this crude analogy.
- tails4e 6y agoI agree with the sentiment, but the numbers are off. It's about 10x the power worst case (maybe 5x for some dsp heavy apps) and also around 5 to 10x for speed. An FPGA can easily run at 100s of MHz, up to 500 with good design pipelining, so suggesting an ASIC could do 500x20 times the speed is 10Ghz, so definitely beyond most ASICs, so I think 5x is more reasonable.
- brandmeyer 6y agoMy experience is that to get those "high" clock frequencies that the work per cycle has to be extremely small. If you normalize to total circuit delay in units of time than you still end up many times worse, because you need many extra pipeline cycles to get the Fmax that high.
- tails4e 6y agoMy day job is ASIC design and we do some prototyping on FPGAs, so the exact same RTL is used as an input. We always benchmark power, performance, etc between ASIC and FPGA, so this is based on some real deigns. A 5x reduction in power is fair for most of what I've seen and the FPGA is actually better at achieving FMAX than you'd expect - control paths do need a lot more pipelining than ASIC, but compute intensive (DSP) datapaths are pretty good with a few tweaks. I think sometimes people throw code at them and get 100 MHz and say we'll FPGAs are slow so it's expected, but in my experience with a little tuning you can get most datapaths to run at 500MHz. You do pay the power penalty vs dedicated ASIC, but the performance is very good.
- brandmeyer 6y agoI think it depends a great deal on what you're doing. A fully pipelined double-precision floating-point fused multiply-add in FPGA tech will reach well over 500 MHz on current parts, but takes almost 30 cycles of pipelined latency to deliver each result. On the same process node, a well-optimized CPU will run at 6-8x the clock frequency and only require 4 cycles of latency to deliver each result. Is this flow filled with divide-and-conquer algorithms with very low work per step? Yes. Is that particularly ill-suited to FPGA logic? Yes. Is it unfair to the FPGA? Not in my opinion. I stand by my claim: If you normalize a general circuit's speed in units of time instead of cycles, then you'll find that ASICs come out much much farther ahead.
- rjsw 6y agoI'm guessing you didn't mean to write "softcores", licensing a SiFive design to be a hard core connected to the FPGA fabric would be one option.
- gmueckl 6y agoArgh, you are right! Thanks for pointing it out. I did pick the wrong wording.
- duskwuff 6y agoXilinx already has a number of devices with ARM hard cores, though (like the Zynq series). There's no compelling reason for them to switch away from that.
- rjsw 6y agoUnless NVidia gives Xilinx a compelling reason to switch away from ARM.
- ansible 6y agoAnd for their products that include hard cores, maybe they will switch to RISC-V like with the MicroSemi PolarFire. I'm still debating on getting the Icicle development kit.
- jl2718 6y agoI don’t think NVidia/ARM would affect Xilinx much. Given what the bulk of FPGAs are doing in the data center, I think AMD was looking more at NVidia/Mellanox and of course Intel/Altera, but for networking, not compute. For Xilinx, this gives a path to board-level integration with x86.
- andy_ppp 6y agoOr package level integration...
- gumby 6y agoPossible, but I suspect heat and area would be problems, at least for the CPUs. A smaller AMD core could be supplied as a hard core on the Xilinx part but would that really be worth it?
- state_less 6y agoIt would be an interesting to get a cpu, gpu and fpga in one package. ML abstractions could be sent to the most sensible implementation. Maybe handheld mobile SDR transceivers could benefit from such an arrangement too. AMD has had some great low power parts and maybe they’ll drive for that here?
- gumby 6y ago“Lowmpower” is not what comes to mind when I think “FPGA”.
- pclmulqdq 6y agoIntel had some heat problems when they tried this. The FPGAs weren't able to use their heat budget dynamically, and as a result, the whole SiP had bad performance.
- lallysingh 6y agoAMD chiplets?
- person_of_color 6y agoWhat is the difference between hard and soft cores?
- acallan 6y agoA soft core is a CPU that is programmed into an FPGA instead of a "regular" core that is made of discrete components.
- KSteffensen 6y agoSoft cores use the configurable logic matrix of the FPGA. You can choose to implement them or not, depending on your use case. They can also be tuned to the use case, adding or modifying CPU instructions, cache structure, e.t.c. This involves writing RTL code, with all the design, verification and backend synthesis work that comes with that. Tools like Synopsys ASIP Designer tries to help with this effort. Hard cores are not part of the configurable logic matrix but are separate resources on the FPGA. That means they can't be tuned to the use case in the same way as a soft core. The trade-off is that they typically are better optimized with regards to clock frequency and power consumption since the components are made to be a CPU and not generic configurable logic. One example of an FPGA with a hard core CPU would be the Xilinx Zynq devices.
- dreamcompiler 6y agoAn FPGA is (to a very crude approximation) just a bunch of static RAM organized in an unusual way. If you think of normal static RAM as "address wires go in one side and data wires come out the other", in an FPGA there are no "address wires" -- it's all data in/data out. The memory cells are still just memory cells; what we label the wires is merely a matter of engineering perspective. In a Xilinx memory cell we choose labels for the wires typically used for logic gates. Anyway in a Xilinx chip the bits of data you put in the memory cells determine what logic function gets executed. That works because in general any particular stored memory -- in any computer, anywhere -- is (conceptually) just a logic function, and conversely all logic is implementable with the stuff we conventionally call memory. But we typically don't do that, because "real" logic made of fixed-function transistors is much faster than logic built with changeable memory cells. However, there's a market for fully-changeable logic--even if it's slower--and that's what Xilinx chips are. Every CPU is just a bunch of registers and logic. If you hand me a few million discrete NAND gates, I can use them to build an X86, a RISC-V, and ARM, or whatever. It will be the size of a house and it will be very slow, but it will run the binary code for that processor. With a Xilinx chip, you have a few million NAND gates (or NOR gates or inverters or whatever you like) at your disposal and they're all on one chip and you can wire them up however you want with nothing but software. Bingo: You can build an X86 out of pure logic, and it's all on one chip rather than being the size of a house. That's a soft core. The nice thing about soft cores is that you can build whatever CPU functions you want and leave off the functions you don't need. If you want to change the design, you just download a bunch of new bits to the Xilinx memory cells. Thus you can change an ARM into an X86 in an instant, without changing any hardware. Soft cores are very flexible, but they're also slow, because implementing logic with static RAM cells is slower than doing it with dedicated transistors. That's where hard cores come in: A hard core is a dedicated area of silicon on the Xilinx chip carved out to only implement an ARM chip or a PowerPC or other CPU with fixed-function transistors. So it's fast. The downside is you can't change its functionality on-the-fly. If you decide you'd rather have a PowerPC than an ARM chip you have to change the whole chip. In both types of cores, you still have a bunch of memory cells left over that you can program to do whatever kind of logic you like.
- ohazi 6y agoAlso, the advantage Altera supposedly got after being acquired by Intel was better fab integration with what was then the best process technology available (High-end FPGAs genuinely need good processes). 1. That's no longer the case, so sucks for Altera / Intel 2. AMD doesn't have a fab, so any advantages are necessarily on the design / architecture / integration side.
- dogma1138 6y agoIntel got their 3D/2.5D stacking tech from Altera, the FPGA in Xeon sockets also is doing as well as it can considering the niche market.
- andromeduck 6y agoMy read on the Altera acc was that Intel needed to shore up fab volumes in the face of their foundry customers jumping to TSMC first chance they could. As the capital required per node continues to rise exponentially, they need more and more volume to amortize that over. This is also why they're trying to get into GPUs again.
- samps 6y agoTo slightly refine this, Intel didn't have many "foundry customers" before Altera. Via Wikipedia (https://en.wikipedia.org/wiki/Intel#Opening_up_the_foundries_to_other_manufacturers_(2013) https://en.wikipedia.org/wiki/Intel#Opening_up_the_foundries...), the need to fill up the manufacturing lines was engendered by poor x86 CPU sales around ~2013, not poor third-party fab runs. In 2013, Intel was still ahead of TSMC with 22 nm.
- andromeduck 6y agoDidn't they also drag down Panasonic or something? I remember there being a lot of rumors of them being an extremely bad partner at the time, basically no technical support, unreliable capacity due to by core business and not giving two shits about the success of the venture in general.
- 6y ago
- baybal2 6y agoI do not believe it makes sense to spend so much money for a niche, in a niche product like AI/ML chips. And I believe AMD are good with using calculators.
- datameta 6y agoBeing early on the ML hardware acceleration boat is going to pay off astronomically. Embedded inferencing is going to be a society defining technology by the time we hit mid-decade. It's already being used for predictive maintenance of machinery in IIoT with huge payoffs via decrease in unforeseen total machine failure or need of heavy overhauls.
- hacknat 6y agoIt really blows my mind how many people are still bearish on ML. It’s fair to argue timelines (although even that is becoming less true), but I think the evidence is firmly on the side of the bulls now.
- m0zg 6y agoI think you're onto something here. AMD is likely seeing the end of the road for their CPU business within the next decade, since it will run up against physics and truly insane cost structures that will come after 5nm. At the same time we're far past the practical limit wrt ISA complexity (as evidenced by periodic lamentations about AVX512 on this site). The only real way to go past all of that right now is specialized compute, reconfigurable on demand, deployment of which is hampered by the fact that it's very expensive and not integrated into anything, so the decision to use it is very deliberate, which in practice means it rarely ever happens at all. Bundle a mid-size FPGA as a standardized chiplet on a CPU, integrate it well, provide less painful tooling, and that will change. Want hardware FFT? You got it. Want hardware TPU for bfloat16? You got it. Want it for int8? You got it. Think of just being able to add whatever specialized instruction(s) you want to your CPU. I'm not sure this is worth $35B, but if Lisa Su thinks so, it probably is. She's proven herself to be one of the most capable CEOs in tech.
- WWLink 6y ago> My speculation is that this gives Xilinx an alternative high performance AMD64 (and possibly lower performance & lower power x86) "hard cores" to displace the Arm cores. I'm trying to imagine an x86-based Ultascale+ style processor. Hopefully AMD can help fix the mess known as vivado and the petalinux tools. lol.