5 ms·
I find it incredibly dishonest of Adapteva to equate it to a "theoretical 45 GHz CPU". There are much better ways to talk about the performance level of their h
by rys 13y ago
I find it incredibly dishonest of Adapteva to equate it to a "theoretical 45 GHz CPU". There are much better ways to talk about the performance level of their hardware than that metric, especially given the rest of the text in their Kickstarter pitch is aimed at people who need to inherently understand the hardware's execution model in order to program it effectively.
The computing industry has established language and metrics to discuss computing performance and, while the waters often get muddied when the hardware is wide, that's a step too far.
- mtrimpe 13y agoThis board should deliver about 90 GFLOPS of performance, or — in terms PC users understand — about the same horse-power as a 45GHz CPU. That doesn't seem too outrageous to me. Edit: They state the real fact and then give another figure explicitly stating it's an attempt to translate this into a metric the average user can somewhat relate to. According to http://en.wikipedia.org/wiki/FLOPS#Computing http://en.wikipedia.org/wiki/FLOPS#Computing it seems that they're off by a factor of two, but I'm guessing that's just an honest mistake. Second edit: I was under the impression that this was the result of dumbing down by a journalist, however it seems it's from Parallela itself. That is a bit disingenuous indeed.
- Xcelerate 13y agoMe neither. The only catch is that you can't get serial computation that fast, but I assume anyone buying something called "Parallella" would realize that already.
- Tuna-Fish 13y agoA single Ivy Bridge core has 8 Flops/MHz of computing power. 45GHz Ivy Bridge would be able to do 360GFlops.
- helpbygrace 13y agoIf you are correct for the first clause(8 Flops/MHz), 45GHz of Ivy Bridge core has 360k Flops (8 Flops/MHz * 45GHz ==> 8 Flops * 45k).
- danbruc 13y agoIt should read 8 FLOPS per cycle double precision. So a 3 GHz 4 core Ivy Bridge processor could theoretically peak at 96 GFLOPS double precision, 192 GFLOPS single precision.
- deleted 13y ago[deleted]
- stephencanon 13y ago8 double-precision flops/cycle/core is the correct figure for Ivy Bridge and Sandy Bridge. With Haswell adding FMA, that figure doubles again(!)
- mrb 13y agoHum, no. Sandy/Ivy Bridge can only execute 4 double-precision instructions per cycle per core, in the form of two SSE instructions per cycle (one instruction doing adds, the other doing muls, executed by different units). Doing 8 double-precision instructions per cycle would translate to either four 128-bit SSE instructions, or two 256-bit AVX instructions per cycle, which is not possible (unless I did not keep track of the latest AVX capabilities).
- kryptiskt 13y agoThe only reason it isn't a big lie is that it's an utterly meaningless statement. In any case it is a misrepresentation of what modern CPUs are capable of.
- phoyd 13y agoIt is 90 GFLOPS at less that 5 Watts. That's not too shabby. Adapteva claims that the Epiphany cores have a GFLOPS/W ratio of 50. See here: http://streamcomputing.eu/blog/2012-08-27/processors-that-can-do-20-gflops-watt/ http://streamcomputing.eu/blog/2012-08-27/processors-that-ca... Also, the boards for the backers feature a ZYNQ-7020 SOC by XILINX which sports a 1.3M Gate FPGA, available to the user. This ain't bad either.
- Uchikoma 13y agoMost others - like my MacBook here [1] - are advertising GHz / core not "summing" the cores. So the statement "PC users understand" is false. [1] http://store.apple.com/us/browse/home/shop_mac/family/macbook_pro http://store.apple.com/us/browse/home/shop_mac/family/macboo...
- marshray 13y agoBecause all of us dumb PC users measure performance in terms of "horse-power". :-)
- amalag 13y agoThey do answer that on their kickstarter page: http://www.kickstarter.com/projects/adapteva/parallella-a-supercomputer-for-everyone http://www.kickstarter.com/projects/adapteva/parallella-a-su... Why do you say the Parallella is a 45GHz computer? We have received a lot of negative feedback regarding this number so we want to explain the meaning and motivation. A single number can never characterize the performance of an architecture. The only thing that really matters is how many seconds and how many joules YOUR application consumes on a specific platform. Still, we think multiplying the core frequency(700MHz) times the number of cores (64) is as good a metric as any. As a comparison point, the theoretical peak GFLOPS number often quoted for GPUs is really only reachable if you have an application with significant data parallelism and limited branching. Other numbers used in the past by processors include: peak GFLOPS, MIPS, Dhrystone scores, CoreMark scores, SPEC scores, Linpack scores, etc. Taken by themselves, datasheet specs mean very little. We have published all of our data and manuals and we hope it's clear what our architecture can do. If not, let us know how we can convince you.
- dekhn 13y agoThey are clearly wrong. The purpose of higher clock rates is to produce a given answer in a smaller amount of time (latency). The purpose of adding more processors (cores) is to produce more answers in a given time (throughput). They are free to report their results using any standard measurement of throughput. Their answer is weasely. But then, clock rate is irrelevant to latency and throughput. Really matters how much more work per cycle you get done, how fast you can move IO, and the cost of throughput per watt (and whether the system can meet your requirements at all).
- vidarh 13y agoThey are "clearly wrong" when talking to geeks about specific types of problems. For most most people this means nothing, and multiplying it is fine. And for a lot of situations where you are considering batch jobs, multiplying it is fine as a quick illustration. It is not as if the raw numbers tell you anything anyway, since the characteristics of the system are so unusual. They miscalculated how people would interpret it, and got burned. But they've been clear about what it is they actually mean the whole time.
- TallGuyShort 13y agoI don't see a big difference between this and "petaflops" measurements that are the de-facto standard in bragging about supercomputers. You really only hit that peak performance for embarrassingly-parallel problems, but unless you have a specific workload or benchmark to talk about, it's the best you have and is a fairly well-accepted practice in the industry. On a related note "petaflops" would be a great name for a pet bunny.