46 ms·
The History, Status, and Future of FPGAs
- wwarner 6y agoThis is really interesting. If a cpu hardware vulnerability like spectre could be repaired by patching an fpga on the SOC that would be incredible. That type of functionality would overtake the entire cloud market in about 3 days.
- glitchc 6y agoIt would also open up new attack vectors.
- thehappypm 6y agoThat's the real nightmare. Now all of a sudden, you can program the CPU itself if you can access the update mechanism. CPUs being non-programmable is a feature as well as a bug.
- deelowe 6y agoCPUs are already "programmable" via microcode updates.
- cwzwarich 6y agoPretty much every new non-x86 CPU doesn't have updatable microcode, so that's a very x86-centric problem.
- gtsteve 6y agoMicrocode is loaded when the OS starts though right? At the very least it's not persistent.
- deelowe 6y agoBIOS or OS
- pjmlp 6y agoAnd have been since ages, that was one of the themes regarding RISC Vs CISC design.
- rwmj 6y agoI'm afraid it doesn't work like this. That would only be possible if the chip was using an FPGA fabric for the relevant parts of the design. For example if the L1 cache was implemented as an FPGA you could in theory patch around L1TF. But they wouldn't do that because it would be far slower/larger than implementing it directly as an ASIC. Or you might imagine a chip that has an FPGA on the side (I expected Intel would ship this after acquiring Altera, but it never happened). But the FPGA would somehow have to have access to the paths that caused the vulnerability, which is highly unlikely, and would also be really slow compared to what they actually do which is hacking around it by microcode changes.
- duskwuff 6y ago> Or you might imagine a chip that has an FPGA on the side (I expected Intel would ship this after acquiring Altera, but it never happened). They did: https://www.anandtech.com/show/12773/intel-shows-xeon-scalable-gold-6138p-with-integrated-fpga-shipping-to-vendors https://www.anandtech.com/show/12773/intel-shows-xeon-scalab... But I get the sense this part was aimed at a few very specific customers. It required some PCB-level power delivery changes, so you couldn't even drop it into a standard server motherboard.
- jeffreyrogers 6y agoFPGAs are too slow for that. I think you can get the clock rate up to about 600Mhz, but that is only for very small portions of the chip. Otherwise you run into timing issues. The clock speed for most of the chip will be significantly lower.
- rcxdude 6y agoYup. If you just want a CPU, use a CPU. an FPGA is a terrible substitute, and generally you only want to embed a CPU on them if you are either developing a CPU or you want a not very fast CPU as an addon to a design which is already using an FPGA (and generally for this nowadays the vendors make FPGAs whith a CPU on the same die, because it's so common and frees up quite a lot of the FPGA fabric and power budget).
- rustybolt 6y agoAmazon already has FPGA's on the cloud: https://aws.amazon.com/ec2/instance-types/f1/ https://aws.amazon.com/ec2/instance-types/f1/ I don't think they are very popular though. Maybe they are used sometimes for machine learning?
- rwmj 6y ago> Intel, AMD, and many other companies use FPGAs to emulate their chips before manufacturing them. Really? I'm assuming if this is true it can only be for tiny parts of the design, or they have some gigantic wafer-scale FPGA that they're not telling anyone about :-) Anyway I thought they mainly used software emulation to verify their designs.
- k0stas 6y agoThe largest FPGAs were reticle-busters when I used to work on them. Today I think the largest FPGAs use chiplet-style integration. Even with the inefficiency of an FPGA, many smaller chip designs can still fit on the largest FPGA. Also, there are prototyping boards specifically built for emulation that integrate multiple FPGAs, although this does introduces a partitioning problem that has to be solved either manually or via dedicated emulator software.
- jcranmer 6y agoThe FPGA emulator for a chip I was working on involved an entire rack of FPGAs... for a single core.
- variaga 6y agoOf the half-dozen semiconductor- designing companies I've worked for, all of them used FPGAs for emulation. - modern FPGAs are huge. - when an asic design won't fit in a single FPGA, it's usually possible to partition the design into multiple FPGAs - software emulation/ simulation is not guaranteed to be "more accurate". FPGAs can interact with a real-world environment in ways that simulation simply cannot - simulations run 1000s of times slower than FPGAs. Months of simulation time can be covered in minutes on the FPGA Edit: to be clear, they all use simulation too, but FPGAs are used to accelerate the verification process
- GeorgeTirebiter 6y agoIs that still true in 2020? Or is the simulation getting good enough to skip the FPGA prototyping phase?
- d_silin 6y agoI wonder if it is possible to add a (small) FPGA to a personal computer that could accelerate any specific software tasks (video/audio encoding, ML algorithms, compression, extra FPU capabilities) on user demand.
- jeffreyrogers 6y agoThe problem with this will be the overhead of transferring data to/from the FPGA, which once accounted for often causes doing the computation on the CPU to make more sense. It's obviously not a show-stopper, since GPUs have the same problem, but are still useful, but it's hard to find a workload that maps well to this solution.
- not2b 6y agoThis is normally handled in emulation by putting the inner parts of the testbench (the transactors) onto the FPGA as well, to minimize the amount of data that has to be transferred between the CPU and the FPGA. If the FPGA is to be used as a peripheral, again a division of labor needs to be found that minimizes the amount of data that needs to be communicated. But if there is FPGA logic on the same chip as the CPU cores, the overhead can be greatly reduced, and we're seeing more of that now.
- derefr 6y agoIn a DAW, accelerating a heavy VST plugin might make sense. But often those are amenable to being translated to GPGPU code already. I guess the one place where GPGPU-based solutions wouldn't work, is when the code you want to accelerate is necessarily acting as some kind of Turing machine (i.e. emulation for some other architecture.) However, I can't think of a situation where an FPGA programmed with the netlist for arch A, running alongside a CPU running arch B, would make more sense than just getting the arch-B CPU to emulate arch A; unless, perhaps, the instructions in arch-A are very, very CISC, perhaps with analogue components (e.g. RF logic, like a cellular baseband modem.)
- deelowe 6y agoI assumed this was kind of intel's plan when they purchased Altera. I this issue with this is the amount of time it takes to load the bitstream, but I thought I saw some things recently where progress was being made on this front.
- Koshkin 6y agoI wonder what would be the advantages of using an FPGA to test a CPU design - compared to relying on a (presumably more accurate) computer-based simulation. (I understand the reasons one might want to implement a CPU in an FPGA.)
- dbcurtis 6y agoThis idea is more than 30 years old. It has been done, and one upon a time companies were built around this idea. First off, mapping an entire CPU to an FPGA cluster is a design challenge itself. Assuming you can build an FPGA cluster large enough to hold your CPU, and reliable enough to get work done on it, you have the problem of partitioning your design across the FPGA's. Second problem: observability. In a simulator, you can probe anywhere trivially, with an FPGA cluster, you must route the probed signal to something you can observe. (I am not even going to talk about getting stimulus in and results out, since with FPGA or simulator, either way you have that problem, it is just different mechanics.) The big problem is that an FPGA models each signal with two states: 1 and 0. A logic simulator can use more states, in particular U or "unknown". All latches should come up U, and getting out of reset (a non-trivial problem), to grossly oversimplify, is "chasing the U's away". An FPGA model could, in theory, model signals with more than two states. The model size will grow quickly. Source: Once upon a time I was pre-silicon validation manager for a CPU you have heard of, and maybe used. Once upon a time I was architect of a hardware-implemented logic simulator that used 192 states (not 2) to model the various vagaries of wired-net resolution. Once upon a time I watched several cube-neighbors wrestle with the FPGA model of another CPU you have heard of, and maybe used. Note: What would 3 state truth tables look like, with states 0,1,U? 0 and 1 is 0. 0 and U is 0. 1 and U is U -- etc. You can work out the rest with that hint, I think. Edit to add: Why are U's important? They uncover a large class of reset bugs and bus-clash bugs. I once worked on a mainframe CPU where we simulated the design using a two-state simulator. Most of the bugs in bring-up were getting out of reset. Once we could do load-add-store-jump, the rest just mostly worked. Reset bugs suck.
- jacquesm 6y ago> Reset bugs suck. Indeed they do. And even if you have working chips you get the next stage: board level reset bugs. A MC68K board I helped develop didn't want to boot, some nasty side effect of a reset line that didn't stay at the same level long enough stopped the CPU from resetting reliably when everything else did just fine. That took a while to debug.
- jcranmer 6y agoAs a bit of a counterpoint: One of my prior projects involved working with a lot of ex-FPGA developers. This is obviously a rather biased group of people, but I saw a lot of feedback around that was very negative about FPGAs. One comment that's telling is that since the 90s, FPGAs were seen as the obvious "next big technology" for HPC market... and then Nvidia came out and pushed CUDA hard, and now GPGPUs have cornered the market. FPGAs are still trying to make inroads (the article here mentions it), but the general sense I have is that success has not been forthcoming. The issue with FPGAs is you start with a clock rate in the 100s of MHz (exact clock rate is dependent on how long the paths need to be), compared with a few GHz for GPUs and CPUs. Thus you need a 5× performance win from switching to an FPGA just to break even, and you probably need another 2× on top of that to motivate people going through the pain of FPGA programming. Nvidia made GPGPU work by being able to demonstrate meaningful performance gains to make the cost of rewriting code worth it; FPGAs have yet to do that. Edit: It's worth noting that the programming model of FPGAs has consistently been cited as the thing holding back FPGAs for the past 20 years. The success of GPGPU, despite the need to move to a different programming model to achieve gains there, and the inability of the FPGA community to furnish the necessary magic programming model suggests to me (and my FPGA-skeptic coworkers) that the programming model isn't the actual issue preventing FPGAs from succeeding, but that FPGAs have structural issues (e.g., low clock speeds) that prevent their utility in wider market classes.
- lnsru 6y agoIt’s not the speed, that holds FPGA adaptation back. It’s development process/time. While one can start with GPU immediately, there is a need for FPGA to develop whole PCIe infrastructure and efficient data movers. One is done with GPU while FPGA developers just start with algorithms. As long as one does not need real time capability, GPU is an obvious choice. My 200 MHz design outcompetes every CPU and GPU out there with very narrow data processing window, but development time is 5x compared to regular software.
- kyboren 6y agoGPUs work great for accelerating many applications, and it's true that that reduces interest in FPGAs. For applications that map well to GPUs, you're absolutely correct that the higher clock speeds (and greater effective logic area) make GPUs superior as accelerators. However, some applications do not map well to GPUs. Particularly those applications with a great deal of bit-level parallelism can achieve enormous speedups with bespoke hardware. For those applications where it doesn't make sense to tape out an ASIC, FPGAs are beautiful--even if they only operate at a few hundred MHz. I think the "programming model" is actually the biggest barrier to wider adoption. Your comment is suffused with what I believe is the source of this disagreement: The idea that one programs an FPGA. One designs hardware that is implemented on an FPGA. The difference may sound pedantic, but it really is not. There is a massively huge difference between software programming and hardware design, and hardware design is downright unnatural for software developers. They are completely different skill sets. On top of that add all the headaches that come with implementing a physical device with physical constraints (the article complains about P&R times but this is far from the only burden) and it becomes clear that FPGAs are quite frankly a massive pain in the ass compared to software running on CPUs or GPUs.
- lnsru 6y agoI am working right now on bare metal websockets implementation on Xilinx Series 7 FPGAs. Currently it’s ZynQ SoC, but final product will probably have Kintex 7 inside, so no Linux. The tools make me cry, no examples, application notes from 2014 with ancient libraries. I hope, vendors will fix tooling. But I see, Xilinx has released Vitis, so their scope is elsewhere, no interest in old crap. Using Git with Vivado is already enough pain. So I keep my text sources in Git and complete zipped projects as releases. Ouch!
- tails4e 6y agoI posted this elsewhere, there are a lot of good resources and examples for the tools: https://github.com/xupgit/FPGA-Design-Flow-using-Vivado/tree/master/slides https://github.com/xupgit/FPGA-Design-Flow-using-Vivado/tree... https://www.xilinx.com/support/university.html https://www.xilinx.com/support/university.html https://www.xilinx.com/video/hardware/getting-started-with-the-vivado-ide.html https://www.xilinx.com/video/hardware/getting-started-with-t... There are others thst cover the SDK side of things, but the HW side/Vivado is well documented.
- mindentropy 6y agoHave you looked at open source solutions? Tim Ansell is managing some great projects on open source solutions. Check out Symbiflow, LiteX, Yosys etc.
- lnsru 6y agoAre these mature already? It took some time for KiCad to get to current usable state and I don’t want to be early adopter. In fact, I want to have my private hardware MVP next year with current tools. On the other hand I can’t imagine my slacker colleagues using anything else than Vivado. Learning Vivado for them was already mission impossible.
- IshKebab 6y agoI wouldn't say KiCad is usable yet. I've made multiple attempts to use it and it just is fundamentally user hostile. Unfortunately the devs see any attempt to improve user friendliness as "dumbing down". Fortunately there is (finally!) an open source PCB design program that doesn't suck: Horizon EDA. I've only made one PCB with it but honestly it was pretty great and the author fixed every usability bug I reported in a matter of hours, which is an insane difference from KiCad's "you're holding it wrong". The only think I don't like about it is it has an unnecessarily powerful and confusing component system (there are modules, entities, gates, etc.). But really it is the best by far. Anyway, on FPGAs, I think the tools are only vaguely mature for iCE40 and even then you basically need to already be an expert unfortunately.
- retro_guy 6y agoMaybe you will find this article about Large-Scale Field-Programmable Analog Arrays [FPAAs] interesting as well: https://hasler.ece.gatech.edu/FPAA_IEEEXPlore_2020.pdf https://hasler.ece.gatech.edu/FPAA_IEEEXPlore_2020.pdf
- PanosJee 6y agoinaccel.com is making lots of steps to bring FPGA to 2020 Spark/k8s integration Abstraction of popular cores Python APIS Serverless deployments Etc
- justicezyx 6y agoFPGAs are good at nothing in the scale that can challenge non-configurable silicons... They are good at a lot of things that are in a smaller scales. Like general prototyping/testing/simulation, telecom, special-purpose real-time computing etc. The behind-scene logic is that FPGAs can never make things as flexible as software. And flexible software always offset the inefficiency in a non-configurable chips. Just comparing FPGAs and CPUs/GPUs will never teach FPGAs vendors the reality, or they choose to ignore after all...
- GeorgeTirebiter 6y agoI believe you are incorrect. A counterexample to your claim is the increasing use of FPGAs in the datacenter. And various AI engines are FPGA-based. You'll do better for a CPU in Real Silicon; but a full-featured MPU w/standard peripherals + FPGA for unusual & must-be-fast functions is hard to beat.
- justicezyx 6y agoTell me how much users are using FOGAs and why xillinx is just a fraction of nVidia's market cap. 5 years ago, nvidia was 2x of xillinx in market cap, now it's 10x.
- bsder 6y agoThe problem that FPGAs have is that they are only good for low-volume solutions that require flexibility and have no power constraints. That's a really narrow market. Telecom equipment and lab equipment, basically. If I need volume, I need at least an ASIC. If I need to manage power, I need a full custom design.
- GeorgeTirebiter 6y agoMicroSemi (now part of Microchip) makes some low-power FPGAs. Xilinx has made the coolrunner CPLDs for years that are mighty low-power (they're not huge, but often are big enough for some needed extra logic.). (Another not care too much about power is Military.)
- m3kw9 6y agoThe thing with FPGA is that companies when faced with cash and time crunch will opt to use a FPGA instead of designing ASICs. The tools suck but companies will hire someone that will do it. FPGA fit a very particular constraint and still solves very specific problems efficiently
- andromeduck 6y agoIMO the next big application for FPGAs is going to be to serve as a programmable DMA-engine of sorts. Have some a bunch of hard logic like ALUs and/or IO/s strewn about. Like for hw accelerated sql queries, malloc/free, data-specific compressors and the like.
- inaccel 6y ago2 are the main challenges of the FPGA utilization: - The first one is the FPGA programming. Now using OpenCL and HLS is much easier compared to VHDL/verilog to design your own accelerators. - The second one is the FPGA deployment and integration. Until now it was very difficult to integrate your design with applications, to scale-out efficiently and to share it among multiple threads/users. The main reason was the lack of an OS_layer (or abstraction layer) that would enable to treat FPGAs as any other computing resource (CPU, GPU). This is why at inaccel we developed a unique vendor-agnostic orchestrator for FPGAs. The orchestrator allows much easier integration, scaling and resource sharing of FPGAs. That way we have managed to decouple the FPGA designer from the software developer. The FPGA designer creates the bitstream and the software developer just call the function that wants to accelerate. No need to define the bitstream file, no need to define the interface or the memory buffer allocation. And the best part: It is vendor and platform agnostic. The FPGA designer creates multiple bitstream for different platform and the software developer couldn't care less. The developer just call the function and the inaccel FPGA orchestrator magically configure the right FPGA for the right function.