6 ms·
Two thoughts: * Most of the algorithms that we want to work with in this domain are doing arithmetic operations on ints and floats. This isn't super difficult
by ianhowson 7y ago
Two thoughts:
* Most of the algorithms that we want to work with in this domain are doing arithmetic operations on ints and floats. This isn't super difficult to do in an RTL, but it's like implementing C++ objects in assembler. You can do it, but you need to think harder than you should.
* FPGAs make you worry about timing. This is a massive shift in thinking for software people. It's also not a value-add; I don't want to care about timing. And it enforces chip-wide dependencies (you can have separate clock domains, but not many of them).
If you simplify the model to "pipelines of arithmetic ops" and then provide an abstraction that eliminates timing (e.g. all ops run in a fixed number of clock cycles and the compiler automatically pipelines them where necessary) then I think you'd have something usable. But this is basically a GPU with a lot of SRAM. Such a constrained problem would run extremely well on any modern GPU or SIMD machine, without the power and cost and obscurity constraints of FPGAs.
- daphreak 7y agoI feel like this is the approach that Mathworks [1] (Matlab/Simulink -> FPGA) and Intel [2] (OpenCL -> FPGA) use. [1] https://www.mathworks.com/solutions/fpga-asic-soc-development.html https://www.mathworks.com/solutions/fpga-asic-soc-developmen... [2] https://www.intel.com/content/www/us/en/software/programmable/sdk-for-opencl/overview.html https://www.intel.com/content/www/us/en/software/programmabl...
- jakear 7y agoHaven’t touched it in a while, but I recall working with BSV in college to be quite different from normal Verilog. Removed a lot of the annoyances with Verilog and put all the functionality behind a Haskell front end. No need to worry about timing, etc. because each component was designed to follow some “ready bit” based interface.
- Dayshine 7y agoYup, wrote a toy processor in Bluespec. It was wonderfully easy to do, and typing was pretty handy too.
- vvanders 7y agoThe big difference for FPGAs from GPUs is that you pay a power cost for that flexibility(both in SRAM, and LUTs). Heck, even most modern FPGAs include a fixed function for multiply operations and the like. Unless your doing something with strong timing constraints and you need it to be very wide they just don't make sense before you even get to the HDL question.
- ianhowson 7y agoRight, and this why I believe that FPGAs aren't competitive with GPUs for computationally-bound tasks. FPGAs are weak here because (a) routing costs power, (b) Most ALUs need to be implemented from LUTs, (c) high-order data flow needs to consider timing of all dependencies, manually, and getting this wrong is very wasteful of resources. GPUs have an apparent weakness where their ALUs are overly general for most tasks. If you're doing primarily INT8 tasks then all of that floating point hardware goes to waste, and it appears that an FPGA has an advantage (INT8 is cheap and predictable there.) However, ALU volume isn't the constraint on GPUs -- it's memory bandwidth. Every prosumer GPU make in the last decade has had an excess of ALUs and a shortage of memory ops per second. This is where we're seeing specialized TPUs, deep learning accelerators and image processing accelerators that provide the right ops paired with the right types of memory in the right places. They're seeing large performance and power improvements over GPUs, at the cost of being further specialized. FPGAs have an advantage for low-latency applications and where predictable timing (nanoseconds) is required, but I don't see them displacing GPUs for compute-bound tasks.
- shaklee3 7y agoThe Volta and higher architecture share ALUs with the floating point units, so nothing is really going to waste. Same with the tensor cores.
- imtringued 7y agoThe weakness of GPUs is that the minimum amount of work is incredibly high for acceleration to make sense. If you have a batch with less than 10000 units of work then you shouldn't even think of writing a GPU kernel. GPUs will never used for stream/network processing workloads because no one wants to wait for a buffer to fill up. A FPGA could in theory offer the same throughput but with a much lower latency.
- justaaron 7y agoI DO want access to timing as most of what I want to use an FPGA to implement directly relates to timing. Think: real-time audio.
- ianhowson 7y agoSure. Then why not use a DSP? Then you've got lower price, easier programming and can still achieve cycle-accurate timing if you desire.
- nitrogen 7y agoThe DSP is great for the computation side of audio, but unless that DSP has every bus you want to interface with, you also need an FPGA for I/O routing to PCI, Ethernet, ADCs and DACs, etc.
- reitzensteinm 7y agoIs this really the case? The timing parent is talking about is fractions of a cycle, maybe five orders of magnitude above audible frequencies. How could you take advantage of exposed timing?
- _iiu1 7y ago1) multiplexed audio. Sure, your frame clock may only be 44.1khz, but the bit clock on a mere 8 channels of this will be 11.2896mhz, to say nothing of oversampling, to say nothing of processing 2) low latency processing of the above may require in-situ pipelining in which the pass-through buffer itself IS the processing buffer, etc. imagine being able to eat the pot you boil your pasta in. 3) why WOULDN'T any well-designed elegant system have every single tick as a function of that bit clock? It makes no sense to deliberately place spanners in your own path, particularly as pertains to jitter etc. It's not like you are going to pause your wavelet transform and check your email in process...
- justaaron 7y agopoint being: Control. Precise control over timing is required for deterministic temporal activities. Removing precise control over timing from the language stack one uses to program FPGA's with is removing a desirable feature for many of their uses cases. If one is interesting in glossing over all this abstraction, why is one wishing to use an FPGA at all? I will reverse the question and say: "in which scenarios is someone hoping to avoid addressing precise timing constructs in FPGA programming?" Obviously I'm not referring to clock propagation delay or quantum entanglement etc LOL I mean the intentional macro stuff wrt "timing"
- patrick5415 7y agoI disagree that worrying about timing is not a value add, at least for real time [1] applications. Worrying about the timing is absolutely crucial for applications like servo loops. Making timing guarantees seems a lot more straightforward with an fpga than an RTOS. Controls folks and software people tend to have very different conceptions of what “real-time” means. I’m talking about a loop that must execute once every (say) microsecond exactly or things start physically breaking.
- ip26 7y agoI don't want to care about timing If you don't care about timing at all, that suggests you don't really care if it's fast- in which case, why are you using an FPGA?
- gugagore 7y agoTiming here refers to a design consideratiob, e.g. how fast can the clock go before a logic bit at the output of one unit gets misinterpreted as the opposite level on the input of another unit. That depends on the layout of the design, how many inputs are connected to that output, etc. They don't want to care about timing. Not that they don't care about how fast it runs.
- deleted 7y ago[deleted]