4 ms·
Interestingly, the same result has been achieved by a single person, Christian Jäkel of TU Dresden: https://arxiv.org/abs/2304.00895 https://arxiv.org/abs/2304
by jayhoon 3y ago
Interestingly, the same result has been achieved by a single person, Christian Jäkel of TU Dresden:
https://arxiv.org/abs/2304.00895 https://arxiv.org/abs/2304.00895
- HappyPanacea 3y agoAlso, He used 5311 Nvidia A100 GPU hours and the team in the article used 47000 Intel Stratix 10 GX 2800 FPGA hours.
- lgats 3y agoIntel® Stratix® 10 devices deliver up to 23 TMACs of fixed-point performance and up to 10 TFLOPS of IEEE-754 single-precision floating-point performance. Nvidia A100 GPU Double-Precision Performance: FP64: 9.7 TFLOPS, FP64 Tensor Core: 19.5 TFLOPS From my naive perspective, it would appear Christian's model was almost an order of magnitude more optimized?
- mafuy 3y agoThere are no floating point operations involved here though
- VonTum 3y agoHi, I'm Lennart, the author of the FPGA paper. You can't really compare them on a FLOPS basis. Firstly because our and Jäkel's algorithms are completely different. In fact within the FPGA accelerator itself we don't even use a single multiply, all our operations are boolean logic and counting. Whereas Jäkel was able to exploit the GPU's strong preference for matrix multiplication. All his operations were integer multiplications. In fact, in terms of raw operation count, Jäkel actually did more. There appears to be this tradeoff within the Jumping Formulas, where any jump you pick keeps the same fundamental complexity, with even a slight preference towards smaller jumps. It is just that GPU development is several decades ahead of FPGA development, thanks to ML and Rendering hype, which more than compensates for the slightly worse fundamental complexity. As a sidenote, raw FLOP counts from FPGA vendors are wildly inflated. The issue with FPGA designs is that getting this theorethical FLOP count is nigh-impossible, because getting all components running at the theorethical clock frequency limit is incredibly difficult, compare that with GPUs, where at least your processing frequency is a given.
- versteegen 3y agoHello, thanks for taking the time to reply here! My own Master's thesis was also about optimising and implementing algorithms to count certain (graphical) mathematical objects, but you picked a much more famous problem than me. I'm very surprised I didn't know the definition of Dedekind numbers, although it's related to things I touched on. I'm not too familiar with FPGAs but hope to have a use for them some day. Measuring their performance in FLOPs seems strange. How close to those theoretical limits does one typically get? Are there are a lot of design constraints that conspire against you, or is it just that whatever circuit you want can't be mapped densely to the gate topology?
- bazzargh 3y agothe van Hirtum paper now refers to that one as confirmation of its result https://arxiv.org/abs/2304.03039 https://arxiv.org/abs/2304.03039 beaten to the punch by a couple of days!