3 ms·
How many Rasperry Pis have the equivalent processing power to an original Cray-1? (I would not be at all surprised if a single RPi were as powerful as twenty C
by billpg 3y ago
How many Rasperry Pis have the equivalent processing power to an original Cray-1?
(I would not be at all surprised if a single RPi were as powerful as twenty Cray-1s.)
- mechagodzilla 3y agoA Cray-1 ran at 80 MHz, and with careful coding could sustain about 2 double-precision floating point operations per cycle - so 160 MFLOPS. Looking at linpack benchmarks, it looks like it's right around a raspberry pi 2 running at 1 GHz (169 DP MFLOPS), and a little worse than a Raspberry Pi 3 at 180 MFLOPS.
- dfox 3y agoThe 160MFlOps of Cray-1 is the theoretical maximum of the implementation (as are the whatever FlOp/s numbers quoted by GPU vendors), while the numbers for Linux capable Broadcom SoCs in RaspberryPis are results from LINPACK. The Linpack benchmark is not exactly a representation of what you would compute on a supercomputer, but even this kind of semi-synthetic benchrmark will cause the Cray-1 to score significantly lower than the theoretical 160MFlOp/s.
- AquaLineSpirit 3y agoA Raspberry Pi 4B has 13.5 GFLOPS[0] while a Cray-1 has 160 MFLOPS[1] so you need about 85 Cray-1s :) Couldn't find any numbers for a Pi Pico. 0: https://web.eece.maine.edu/~vweaver/group/green_machines.html https://web.eece.maine.edu/~vweaver/group/green_machines.htm... 1: https://en.wikipedia.org/wiki/Cray-1 https://en.wikipedia.org/wiki/Cray-1
- thrtythreeforty 3y agoWikipedia says the Cray-1 was capable of 160 MFLOPS. As a rule of thumb, modern scalar pipelines can sustain one ALU op per cycle, and you see nearly all Linux-capable CPUs quoted in GHz. So we should expect gigaflops, minimum. And indeed [1] suggests that the Pi 4 is capable of 13.5GFLOP, so about 84 Cray 1's. (The further speedups come from the fact that ARM also has NEON vector instructions, and from multiple cores.) The Pi Pico, on the other hand, does not have a floating point unit. So it emulates it in software (soft float). The C SDK docs [2] suggest 13.8kHz (!) operation for single-precision add. I'll be generous and suggest that the 2x cores could double this performance. So then, it'd achieve 0.0086% of the Cray-1's performance. Oof. If you're willing to do integer arithmetic, things look much better for the Pico, of course - it runs at 125MHz and the above scalar rule of thumb applies. [1]: https://web.eece.maine.edu/~vweaver/group/green_machines.html https://web.eece.maine.edu/~vweaver/group/green_machines.htm... [2]: https://datasheets.raspberrypi.com/pico/raspberry-pi-pico-c-sdk.pdf https://datasheets.raspberrypi.com/pico/raspberry-pi-pico-c-...
- ta988 3y agoAre you talking about Pis or picos like mentioned in this article.
- crest 3y agoThe original Cray 1 had already had a very powerful vector execution unit and low memory latency relative to the CPU frequency (by todays standards). The RP2040 has even lower memory latency, but only ¼ MiB RAM and no hardware floating point support or SIMD (neither packed-SIMD nor "true" Cray style vectors). At least the BootROM includes hand optimised Soft-FP code and the single-cycle I/O block even includes a memory mapped hardware integer divider and two interpolators per CPU core. These can help a lot with fixed point DSP workloads, but are totally different from what made the Cray 1 special at its time. Todays better MCUs may match certain performance numbers of early supercomputers, but their designs have more differences than similarities. There have been attempts to recreate early Crays in FPGAs, but so little software for them has been preserved (in a way accessible to the public) that it's difficult to judge how good the recreations are and even the later Cray 1 based designs need a really big FPGA to reimplement all relevant pieces at least as fast as the original. Cray didn't just build fast floating point adders and multipliers, but the main memories and interconnects between them to make them useful. Good luck finding a way to attach 100s of interleaved DRAM banks to your FPGA, because your internal block RAM won't be enough as main memory and memory access timings have improved the least over time.
- ignite 3y agoAnd what's the relative power consumption? :-)
- adestefan 3y agoAdd in cooling requirements, too.