13 ms·
Parallella, a $99 Linux Supercomputer
- rys 13y agoI find it incredibly dishonest of Adapteva to equate it to a "theoretical 45 GHz CPU". There are much better ways to talk about the performance level of their hardware than that metric, especially given the rest of the text in their Kickstarter pitch is aimed at people who need to inherently understand the hardware's execution model in order to program it effectively. The computing industry has established language and metrics to discuss computing performance and, while the waters often get muddied when the hardware is wide, that's a step too far.
- mtrimpe 13y agoThis board should deliver about 90 GFLOPS of performance, or — in terms PC users understand — about the same horse-power as a 45GHz CPU. That doesn't seem too outrageous to me. Edit: They state the real fact and then give another figure explicitly stating it's an attempt to translate this into a metric the average user can somewhat relate to. According to http://en.wikipedia.org/wiki/FLOPS#Computing http://en.wikipedia.org/wiki/FLOPS#Computing it seems that they're off by a factor of two, but I'm guessing that's just an honest mistake. Second edit: I was under the impression that this was the result of dumbing down by a journalist, however it seems it's from Parallela itself. That is a bit disingenuous indeed.
- Xcelerate 13y agoMe neither. The only catch is that you can't get serial computation that fast, but I assume anyone buying something called "Parallella" would realize that already.
- Tuna-Fish 13y agoA single Ivy Bridge core has 8 Flops/MHz of computing power. 45GHz Ivy Bridge would be able to do 360GFlops.
- helpbygrace 13y agoIf you are correct for the first clause(8 Flops/MHz), 45GHz of Ivy Bridge core has 360k Flops (8 Flops/MHz * 45GHz ==> 8 Flops * 45k).
- danbruc 13y agoIt should read 8 FLOPS per cycle double precision. So a 3 GHz 4 core Ivy Bridge processor could theoretically peak at 96 GFLOPS double precision, 192 GFLOPS single precision.
- deleted 13y ago[deleted]
- stephencanon 13y ago8 double-precision flops/cycle/core is the correct figure for Ivy Bridge and Sandy Bridge. With Haswell adding FMA, that figure doubles again(!)
- mrb 13y agoHum, no. Sandy/Ivy Bridge can only execute 4 double-precision instructions per cycle per core, in the form of two SSE instructions per cycle (one instruction doing adds, the other doing muls, executed by different units). Doing 8 double-precision instructions per cycle would translate to either four 128-bit SSE instructions, or two 256-bit AVX instructions per cycle, which is not possible (unless I did not keep track of the latest AVX capabilities).
- kryptiskt 13y agoThe only reason it isn't a big lie is that it's an utterly meaningless statement. In any case it is a misrepresentation of what modern CPUs are capable of.
- phoyd 13y agoIt is 90 GFLOPS at less that 5 Watts. That's not too shabby. Adapteva claims that the Epiphany cores have a GFLOPS/W ratio of 50. See here: http://streamcomputing.eu/blog/2012-08-27/processors-that-can-do-20-gflops-watt/ http://streamcomputing.eu/blog/2012-08-27/processors-that-ca... Also, the boards for the backers feature a ZYNQ-7020 SOC by XILINX which sports a 1.3M Gate FPGA, available to the user. This ain't bad either.
- Uchikoma 13y agoMost others - like my MacBook here [1] - are advertising GHz / core not "summing" the cores. So the statement "PC users understand" is false. [1] http://store.apple.com/us/browse/home/shop_mac/family/macbook_pro http://store.apple.com/us/browse/home/shop_mac/family/macboo...
- marshray 13y agoBecause all of us dumb PC users measure performance in terms of "horse-power". :-)
- amalag 13y agoThey do answer that on their kickstarter page: http://www.kickstarter.com/projects/adapteva/parallella-a-supercomputer-for-everyone http://www.kickstarter.com/projects/adapteva/parallella-a-su... Why do you say the Parallella is a 45GHz computer? We have received a lot of negative feedback regarding this number so we want to explain the meaning and motivation. A single number can never characterize the performance of an architecture. The only thing that really matters is how many seconds and how many joules YOUR application consumes on a specific platform. Still, we think multiplying the core frequency(700MHz) times the number of cores (64) is as good a metric as any. As a comparison point, the theoretical peak GFLOPS number often quoted for GPUs is really only reachable if you have an application with significant data parallelism and limited branching. Other numbers used in the past by processors include: peak GFLOPS, MIPS, Dhrystone scores, CoreMark scores, SPEC scores, Linpack scores, etc. Taken by themselves, datasheet specs mean very little. We have published all of our data and manuals and we hope it's clear what our architecture can do. If not, let us know how we can convince you.
- dekhn 13y agoThey are clearly wrong. The purpose of higher clock rates is to produce a given answer in a smaller amount of time (latency). The purpose of adding more processors (cores) is to produce more answers in a given time (throughput). They are free to report their results using any standard measurement of throughput. Their answer is weasely. But then, clock rate is irrelevant to latency and throughput. Really matters how much more work per cycle you get done, how fast you can move IO, and the cost of throughput per watt (and whether the system can meet your requirements at all).
- vidarh 13y agoThey are "clearly wrong" when talking to geeks about specific types of problems. For most most people this means nothing, and multiplying it is fine. And for a lot of situations where you are considering batch jobs, multiplying it is fine as a quick illustration. It is not as if the raw numbers tell you anything anyway, since the characteristics of the system are so unusual. They miscalculated how people would interpret it, and got burned. But they've been clear about what it is they actually mean the whole time.
- TallGuyShort 13y agoI don't see a big difference between this and "petaflops" measurements that are the de-facto standard in bragging about supercomputers. You really only hit that peak performance for embarrassingly-parallel problems, but unless you have a specific workload or benchmark to talk about, it's the best you have and is a fairly well-accepted practice in the industry. On a related note "petaflops" would be a great name for a pet bunny.
- yatsyk 13y agoHow many megahashes this hardware computes?
- wmf 13y agoNot enough. Don't bother.
- helpbygrace 13y agoI agree with you. This might be good for bitcoin mining. :)
- lucb1e 13y agoThe real question is: How much power does it consume compared to its hashing power?
- Tuna-Fish 13y agoMuch, much less than the specialized ASIC platforms do.
- helpbygrace 13y agoBut when we consider the consuming power of ASIC platform, I think this board has strength. They said this board consumes 5 watt for typical jobs. https://en.bitcoin.it/wiki/Mining_hardware_comparison https://en.bitcoin.it/wiki/Mining_hardware_comparison
- IanCal 13y agoIt's nowhere near fast enough. My 7970s can push out about 1.3Ghash/s and combined they are capable of around 7 TFLOPs. When (/if) they release the BFL Jalapeño it'll run at 5 Ghash/s and be powered by USB. 90 GFLOPs is equivalent to a decent processor, but nowhere near powerful enough for bitcoin mining.
- Tuna-Fish 13y ago> But when we consider the consuming power of ASIC platform, I think this board has strength. No, it can't possibly have. SHA is half bitshifts-by-constants. On an ASIC platfrom, those essentially refactor to no-ops. There is no way, no how general-purpose hardware could ever possibly get anywhere near even a piss-poor special purpose ASIC for this task. If you think otherwise you simply don't understand the domain. Those 600-watt ASIC systems contain multiple chips and run at tens of GHashes/s. That 5-watt chip, if it's very, very good, might maybe break 40MHash/s.
- Xcelerate 13y agoAs someone who uses supercomputers, I'm not sure I entirely understand the market of this product. It's really cool and I'd love to have one to tinker with, but due to its high parallelization, I see no benefit of using this over a graphics card. I'm not sure if $99 can get you a GPU that reaches 90 GFlops though... perhaps that's where the benefit lies. EDIT: After reviewing their website, I notice they state > One important goal of Parallella is to teach parallel programming... In this respect, I can see how this is useful. Adapting scientific software to GPUs can be difficult and isn't the easiest thing to get into for your average person. This board, with its open-source toolkit and community could make this process a lot easier.
- amalag 13y agoComparing it with a GPU is a natural discussion to have. I believe there is more information on their website to answer the question. But maybe someone else can explain the programming paradigm difference. http://www.adapteva.com/introduction/ http://www.adapteva.com/introduction/
- mkl 13y agoEach of the 64 cores can independently run arbitrary C/C++ code, which is much more flexible than a GPU. Each core has 32KB of local memory, which can also be accessed by the other cores, and there's 1GB of external memory too. Specs: http://www.adapteva.com/products/silicon-devices/e64g401/ http://www.adapteva.com/products/silicon-devices/e64g401/ Architecture: http://www.adapteva.com/wp-content/uploads/2012/10/epiphany_arch_reference_3.12.10.03.pdf http://www.adapteva.com/wp-content/uploads/2012/10/epiphany_... SDK docs: http://www.adapteva.com/wp-content/uploads/2013/04/epiphany_sdk_reference.4.13.03.301.pdf http://www.adapteva.com/wp-content/uploads/2013/04/epiphany_...
- jareds 13y agoI wonder how much work it would take to port cpuminer to this platform for use in mining bitcoins and litecoins. Probably not worth it for bitcoins with the new FPGA hardware but could be interesting for litecoin depending on the memory access speed to main memory.
- lucb1e 13y agoSo how fast is this really? It doesn't sound like much of a supercomputer to me. If it were so super for $99, it'd have been hyped everywhere already and gamers would not buy desktops anymore. It rather sounds like a platform to practice multithreading on.
- Leszek 13y agoSupercomputer =/= "a fast computer".
- lucb1e 13y agoThen what is it? Wikipedia also seems to say it's "a fast computer": "a computer at the frontline of current processing capacity, particularly speed of calculation" https://en.wikipedia.org/wiki/Supercomputer https://en.wikipedia.org/wiki/Supercomputer
- Nursie 13y agoI've heard folks say before (and I include university profs in that) that it's a computer design focused on parallel calculation and low-latency I/O. Not sure if this machine qualifies under that, but that is at least one competing definition to just "cutting edge and very fast"
- VLM 13y agoA supercomputer is a machine that's I/O bound instead of CPU bound, at least as a first approximation. You'll get lots of specs thrown at you like in mid 2013 a supercomputer means using X, Y, and Z technologies. But that is just a longer format version of the above. A pessimist usually warps the definition to a machine that's primarily programmer limited rather than CPU or I/O limited, LOL. Over the decades as parallelism has been popular its drifted more toward being financially limited more so than anything else, in the long run this is probably going to be the new definition, a overall system who's performance is solely limited economically. You might think thats all computers, not so, there's plenty which are inherently limited by architecture to low performance, or limited by programming to single core / single thread tasks. The biggest bummer of supercomputers in the parallel era is no one is doing anything about latency. That's nice that your 2000 processor design with 20 deep pipelines eventually after enormous latency can really churn stuff out, but the olden days pursuit of low latency resulting in speed was pretty interesting technologically. Hilariously you'll even get noobs who don't even understand the difference between latency and speed or claim there isn't one.
- trotsky 13y agoCan anyone explain what the practical differences between something like this and a gpgpu approach? It doesn't sound particularly performant compared to modern gpus otherwise. Maybe they add in some more general purpose instructions for a little more flexibility?
- dagurp 13y agoIf nothing else it's very small and uses very little power
- vidarh 13y agoA typical GPU can execute a small number of threads on a large number of streams of data carefully laid out in memory. Every time you want to do something conditionally on just one data stream, you waste a lot of capacity. In contrast, the Epiphany chips can execute individual threads on each core in parallel on data either local to the core (fastest), on any other core, or in separate main memory. The current Epiphany chips aren't too spectacular, since the core count is "low". They can "only" execute 16 individual instruction streams in parallel. But that's on a chip the size of your finger nail, and their roadmap is aiming for 1024 core chips. They're effectively aiming for people to find ways of making effective use of simple, small, power efficient cores for problems that are not "data parallel" enough to be efficiently done on GPU's.
- UnoriginalGuy 13y agoThis might be a really stupid question: How difficult would it be in the practical sense to keep all of the cores on something like this "fed" with enough information to get benefits from its concurrency? I mean to "feed" all 64 cores enough data/code so they can all "do something" concurrently is one hell of a job all on its own!
- vidarh 13y agoDepends a lot on the type of problem, and I think that's going to be what makes or breaks them. They have some good examples, but you're right, it's a hard problem and one of the reasons it's so important for them to get these dev boards out.
- andyjohnson0 13y agoPrevious discussions: [1] https://news.ycombinator.com/item?id=4635618 https://news.ycombinator.com/item?id=4635618 [2] https://news.ycombinator.com/item?id=4705487 https://news.ycombinator.com/item?id=4705487
- swalsh 13y agoThere's so much negativity in this thread, wasn't the whole idea that these guys had plans and an architecture to scale up to the order of a terra-flop by 2014, and 20 by 2022? And look! they're shipping. This first chip may not be impressive, but I'll welcome a new player to the market who has big plans to innovate.
- DanBC 13y agoSmall cheap cards for learning are a great idea. But then they've included some really weird wording. They know that people are hostile to that wording, yet they chose to continue to use it. When you're educating people it's important to be clear about terminology. Having said that, I think it's neat, and I wish them luck. I think they're missing one of the main points of the RPi's success - it is dirt cheap. At $35 people will take a risk. At $90 people need to think about it. That might sound odd on HN where people tend to have a lot more disposable income.
- rayiner 13y agoI don't get the negativity either. If you look at the architecture manual, this is like a cheap Tilera. It's an interesting programming model (lots of cores in a shared memory SMP with weak memory ordering), and the CPU's are pretty vanilla RISC architectures. For $99, it's a great way to play with something that has the properties of the kinds of CPU's you might see in a future supercomputer.
- new299 13y agoI wrote my notes up here last time this was doing the rounds: http://41j.com/blog/2012/10/my-take-on-the-adapteva-parallella/ http://41j.com/blog/2012/10/my-take-on-the-adapteva-parallel... I'm pretty skeptical, having played with the Tilera I'm not sure it gives you enough of a benefit to warrant the extra effort. The Parallella also looks a lot like a Tilera, I do wonder if there might be IP issues there down the line. I also still think our best bet for this kind of thing is multicore ARM systems.
- m_mueller 13y agoNot to take away anything from Parallela, but you can also get to play around with something that's used in today's supercomputers by buying any Geforce 5xx and start programming in CUDA. A fully documented ARM architecture as an accelerator chip is certainly interesting though. It will take time until the software tooling catches up, but the initial buzz in the HPC community about those ARM newcomers is certainly there. I'd give them a good chance in the long run to outrun Intel MICs and catch up to NVIDIA Tesla. What I'd like to see next is a PCI express expansion card using this technology. See, one of the great benefits about Tesla cards is that you can swap them out in your supercomputers just like you do it with RAM - and you get the newest chip architecture, as long as your PCI bus can handle the load. For multi purpose systems you often still like to have a good number of x86 cores in there however.
- mas921 13y agoa supercomputer is cluster of machines connected by high throughput, low latency interconnect. Hundreds of servers connected together with 1Gigabit is still a "grid cluster" .. you need at least 10Gigabit Ethernet (over iWARP) or infiniband (RDMA) to be considered a supercomputer. This is marketing B.S.! this B.S. is "emphasized" by the 90GFLOPS = 45Ghz thing. 90GFLOPS is by a single 45GHz "ALU" (perhaps an ALU doing Multiply-Add - MADD op.) not a full fledged CPU (like the i7 or Xeon, which has 4-8 cores with each core having 3 ALU's) as the readers might imply. For example the i7 3770K does 121.6GFLOPS @ "only" 3.5Ghz (ref> table page 2 http://elrond.informatik.tu-freiberg.de/papers/WorldComp2012/PDP2833.pdf http://elrond.informatik.tu-freiberg.de/papers/WorldComp2012...) measuring performance with Ghz is soooo Penitum III! the whole thing is very misleading, and I don't like that! Supercomputer? not even funny! Its a Super-"Raspberry Pi". That's it!
- micheljansen 13y agoWow, I remember seeing the original Kickstarter for this and thinking "this will never see the light of day", yet here it is. I still find it a bit of an odd product; neither for hobby or business, but it sure is cheap.
- vidarh 13y agoIt's a developer board. The product is the chips, not this board. This board is there mainly to get a dev board in the hands of people who might want to build cool stuff with it. That they've actually managed to get it price competitive with a lot of cheap ARM computers, despite sporting a Zynq (ARM SoC with built in FPGA) is amazing.
- sliverstorm 13y agoCan't help but wonder if they are in fact taking a loss, backed by Adapteva.
- qdog 13y agoThey seem to actually have support from some of the hardware manufacturers. From Update #31 "Much gratitude goes out to the component manufacturers who really “got it” (Xilinx, Analog Devices, Intersil, Micron, Microchip, Samtec all deserve special thanks). Without their help we would be losing $100 per board!" So, the backers are getting a Very Good Deal, with the hopes that a successful launch will make demand high enough to make the $99 viable with volume.
- melling 13y agoI backed this project simply because it's a great idea to build a very small highly parallel computer that runs on very little power. Maybe this one won't hit it out of the park but it might give other people ideas. Building the first one of anything is always hard. Add a little serendipity and we might get an entirely new use for computers. Just saying that I could do more with a $99 graphics card sort of misses the point.
- st8ic 13y agoThese guys are completely dishonest. I saw their kickstarter video where they said that for $99 you could have "a computer many times faster than anything on the market ZOMG". Yeah, maybe it's faster for all those times during the day when you calculate matrix chain products. But for largely single-threaded tasks, like EVERYTHING you do on a day to day basis, it's going to be significantly slower than your average dual-core i3.
- vidarh 13y agoI backed them on kickstarter, and I don't remember seeing any claim like what you claim to have seen. To me it was always clear that the current models are not particularly fast. They may be fast "per watt", and if they succeed in their roadmap, then their future 1024 core chips may be fast for the subset of problems that they are suitable for. In the meantime, the kickstarter page is/was careful to focus on this as a stepping stone, and developer platform for playing with the technology first and foremost, and not as being about delivering some incredibly fast computer for end users. If anything, they've provided an extreme amount of data, down to cycle counts for memory accesses and the instruction set, and they've dumped a lot of code in our laps, including drivers etc., and the final unit actually comes with a faster version of the Zynq SoC than what they promised, after Xilinx apparently gave them an amazing deal.
- phaet0n 13y agoI'm really disappointed about how shallow the discussions about Adapteva are, and have been, on HN. To remind everyone, the H = hacker. This device is a godsend, as far as I'm concerned. For the first time ever I get fully documented access to compute array on chip. No the architecture wasn't designed for anything specific, like graphics, but that means I don't get bogged down in details I don't care about, like some obscure memory hierarchy. The chip is plain, simple, low-power, and begging for people to have an imagination again. Stop asking what existing things you can do with it, ask what future things having something like this on a SoC would enable. Also, you should really be thinking about the chip at the instruction level, writing toy DSL to asm compilers. Thinking along the lines of, oh yeah I'll use OpenCL so I can be hardware agnostic, is never going to allow you to see what can be possible with it. If you read the docs you'll see what a simple and regular design it is, perfect for writing your own simple tooling. It's been a long time, but I feel like a kid again. Like when I first discovered assembly on my 8086. Finally a simple device I can tinker with, play, and wring performance out of. Hallelujah! :)
- amalag 13y agoGood reminder! Do you see applications in an embedded sense, or are you looking at it to augment a regular computer's capability?
- phaet0n 13y agoI'm actually thinking Adapteva has a lot of future in present areas of growth. 1) On the mobile side, you can have Epiphany, their compute fabric, as a unit directly on the mobile SoC. You can do codec offload, like WebP, WebM, SILK/Opus. You can do basic computer vision for augmented reality applications, or image recognition. Or perhaps physics, integrate gyro output, position the device in absolute three space. I dunno, the point is the compute is open, there for exploitation. It's not like OpenCL where I have to beg the drivers to be available, correct, or performant. Nor is it like Qualcomm's Hexagon, where who knows if I can use it, and I sure as hell won't without signing an NDA. 2) As far as cloud and heterogenous compute goes, again I see an embedded Epiphany being useful. Everybody whines about various things, like for example missing double-precision. Firstly, it's not like the architecture can't be extended in future. But more importantly they miss little details. Each node in Epiphany can branch and do integer. You can see it doing wire-speed protobuff de/coding and other parallel data shuffling of long-living data, that could be compressed, or interleaved somehow. I'm more of a low-power, cloud kind of guy. So that's what I'll be playing with the most when I get my hands on the kit. That and maybe some parallel graph rewriting. Who knows, the sky's the limit.
- pmorici 13y agoCan it mine BitCoin competitively?
- tempaccount9473 13y ago> Can it mine BitCoin competitively? Prior to the popularity of mining using GPUs, it would have been the shizzle. Today's ASIC-based systems will hash circles around it.
- qdog 13y agoIt's only pulling 2W, so it really depends on the performance per W. Maybe that'll be my first project...
- pratik661 13y agoHmm I wonder if there is a way to bypass your graphics card and use this as a GPU?
- dsdjung 13y agoIt is not easy finding people with good parallel programming skills. Hopefully, this will help things along.
- wmf 13y agoOf course if you learn on Adapteva then your knowledge may not translate to the "worse" architectures that are used in the real world. If you want to learn parallel programming, the computer you already have supports threads, CSP, actors, OpenMP, OpenCL, etc.
- 6ren 13y agoNote: the $99 version has 16 cores, not 64 cores. http://www.kickstarter.com/projects/adapteva/parallella-a-supercomputer-for-everyone#faq_40886 http://www.kickstarter.com/projects/adapteva/parallella-a-su... (+ 2 ARM cores)
- nbdbvcrea 13y agoCool. I don't know any other cheap way to experiment with optimization for 64 cores.
- kingmanaz 13y agoCould you imagine a Beowulf cluster of these?
- api 13y agoI can think of some amazing uses for this. I'm tempted to get one just to port this old hack of mine to it: http://adam.ierymenko.name/nanopond.shtml http://adam.ierymenko.name/nanopond.shtml
- jasonkolb 13y agoThis is really cool. Since it's linux I assume it can run the JVM, correct? That's incredibly powerful, as even GPU programming requires bridge libraries. And what, $99? That's incredible. I'm going to get one...
- wmf 13y agoIIRC Linux does not run on Adapteva. Linux runs on the ARM which is next to the Adapteva chip.
- umsm 13y agoI'm not familiar with this so I have a question: Can you interconnect a few of these boards to create a more powerful unit? I notice they have "expansion connectors"...
- mmanfrin 13y agoCan someone explain to me how a $99 computer can have 45ghz of processing power, but an i7 costs 3x that for 1/10th that clock speed? What does this $99 miss out on that my i7 has the capability of doing?
- pekk 13y agoThis is kind of a novelty. Your i7 has way, way more power for jobs which only use a few cores. Most normal jobs are like that, so unless you have specific requirements, the i7 is going to give you much better performance.
- qb45 13y agoFirst of all, this 45GHz figure definitely isn't valid for modern x86 chips - thanks to multiple cores and SIMD instructions they reach few dozen GFLOPS at stock frequencies. Furthermore, x86 chips pack all of their performance in low number of cores, what makes them much more useful for common scalar code. And if 20 times higher scalar performance isn't enough to convince you to pay premium, the complexity required to achieve this level of scalar performance definitely is enough to discourage Intel from selling you i7s for $99.
- joeblau 13y agoI only want to know one thing. How fast can it mine Bitcoins?! I feel like that's the new "...but can it run Crysis?"
- NewAccnt 13y agoI wonder how those in the performance computing sector feel about running a proprietary supervisor with built in DRM on each and every CPU? Raspberry users might not care when for just hobbyist applications, but I doubt any serious scientist is going to overlook that. http://www.arm.com/products/processors/technologies/trustzone.php http://www.arm.com/products/processors/technologies/trustzon...
- trotsky 13y agoIntel platforms have a very similar risk via SMM and the platform code & controller. It's less advanced, but it can easily exert full control over the system without the os allowing it, minus access to some registers and on die cache. It could DMA in or out of the gpu memory as well. Whether your soc vendor forces a secure supervisor to load is up to them, and i'd be surprised if an HPC builder had trouble finding vendors to supply parts with a totally controllable boot chain. I'm sure there are ways to obscure it, but there are just as many ways on x86 platforms, the only real difference being that you could pull the eprom and reflash it and inspect the other board components. There's also plenty of evil things you can put in a soc without relying on trustzone. Bottom line is you have to trust your vendor. If you want a soc integrated and fab monitored by a business/state that is politically aligned with yours it is probably just a matter of paying a premium.
- qb45 13y agoThe hardware cost of TrustZone is rather low and vendors of "compute SoCs" have no reason to ship hypervisor software on their chips. And Raspberry Pi probably doesn't run any secure mode hypervisor as well.
- lgeek 13y agoTrustzone is just a set of hardware features. Most ARM devices don't come with a proprietary supervisor. In fact, Linux used to run in the secure world on some development devices.
- mrb 13y ago"this board should deliver about 90 GFLOPS of performance, or --in terms PC users understand-- about the same horse-power as a 45GHz CPU." This is wrong. A 4-core 3.0 GHz x86-64 processor delivers more GFLOPS than the Parallela: 96 GFLOPS with SSE instructions, because each core can execute 8 single precision instructions, 4 adds and 4 muls, each cycle. And yes, when Parallela claims 90 GFLOPS, they mean single-precision. For example, for the same price as Parallela, you can get a $100 Phenom II X4 965 (4-core, 3.4 GHz, 125W) delivering 109 GFLOPS. Count $200 to include minimal mobo/RAM/PSU (if all you care about is raw GFLOPS). The main advantage that Parallela has with their exotic architecture over x86-64 is a better GFLOPS/Watt metric. But if you care about this metric you should consider GPUs, which beat Parallela: http://parallelis.com/parallela-supercomputing-for-all-of-us/ http://parallelis.com/parallela-supercomputing-for-all-of-us... Parallela may not beat anything on GFLOPS/Watt and GFLOPS/$, but if they can maintain ease of development (x86-64's stronghold) while doing not too bad on these 2 metrics (dominated by GPUs), they may be a good compromise and may have a shot at succeeding in the HPC market.
- m_mueller 13y agoExactly right. ARMs lure isn't really the current performance for supercomputing, it's rather the expectation that they will hit the next big performance wall much later than x86 because of its simple architecture that's suitable for the maximum amount of cores per die space. Give it 2-3 years and we might have the big step in supercomputing architecture at hand.
- protomyth 13y agoIt might be interesting to have a go at writing a version of Connection Machine Lisp for it.
- backprojection 13y agoHow does this compare the the new Intel MIC (Xeon Phi) co-processor boards? I think they claim 1TFLOP. Can we think of this as a low-powered alternative? http://en.wikipedia.org/wiki/Intel_MIC http://en.wikipedia.org/wiki/Intel_MIC
- qb45 13y agoThe general idea is similar - lots of cores with distributed SRAM memory and some shared DRAM, all sitting on 2D mesh network. The main difference is that Epiphany is made of custom simple RISC cores, while Xeon Phi uses 1st gen Pentiums with huge SIMD FPUs slapped on for higher FP throughput (and TDP).
- backprojection 13y agoInteresting. It looks like (from info on wikipedia pages) the Xeon Phi 3100, gets about 3.3 GFLOPS/WATT, whereas the Epiphany E64G401 manages about 50 GFLOPS/WATT. So something like 10 of these might compare to 1 xeon phi, and still be cheaper in terms of hardware, and much cheaper in terms of power consumption.
- lgeek 13y agoOff the top of my mind (sorry, I don't have the time to double check now): Phi has shared GDDR and distributed caches. Phi cores and caches are connected through a bidirectional ring interconnect, not a 2D mesh network. Still similar, but not as much.
- D9u 13y agoI want one, or two, maybe more. I'm totally fascinated with parallel computing.
- zmmmmm 13y agoI wish these things had just a bit more memory. Most of the interesting algorithms I work with (bioinformatics) really want 4G of memory. A lot of them you can squeeze down to 2G but 1G is just out of the question.
- w34 13y agoWhile I find this quite exciting from a pure developer perspective, it also reminded me that I haven't had anything I'd call a Desktop box in quite some time. If I were to ever get a Desktop machine again, it would have to be cheap and light, definitely don't want anything clunky, otherwise a laptop seems preferable to me. There do not seem that many products that would fill that gap, Intel's NUC is too expensive, the Raspberry PI too slow. Apple's mini Mac seems like the best proposition in this segment. I wonder if the Parallela could not only be used as development center, but also as a Desktop computer? It won't run any fancy games, that's clear, but it may actually be usable for browsing, watching videos and office duties.
- gatehead 13y agoCan it run XBMC?
- iso8859-1 13y agoYes.
- madsravn 13y agoSo where do I buy one?
- iso8859-1 13y agoSign up on the site and they'll mail you when you can order. If you only need FPGA, you can get the Mojo (see link elsewhere in thread) from May.
- iso-8859-1 13y agoEpiphany Architecture Reference: http://www.adapteva.com/wp-content/uploads/2012/10/epiphany_arch_reference_3.12.10.03.pdf http://www.adapteva.com/wp-content/uploads/2012/10/epiphany_...
- dharma1 13y agousb 3.0 would have been nice
- dharma1 13y agoi think this is cool, but wouldn't learning OpenCL be more future proof for someone wanting to get into parallel processing? Seems like there is more drive behind GPU development than specialist hardware like this
- peripetylabs 13y agoThis is perfect for numerical computing applications like software-defined radio or image processing, which can now be done on embedded platforms. I'll definitely be ordering a board when they're available.