8 ms·
A Tiny Chip That Could Disrupt Exascale Computing (2015)
- Nomentatus 10y agoThis is a 2015 story that I remember reading, then. Google news search shows only a couple articles this year about Rex Computing and only one tiny bit of news, that they're at tapeout. That's probably par for the course for a startup creating product (or prototype) one. http://semiengineering.com/power-centric-chip-architectures/ http://semiengineering.com/power-centric-chip-architectures/ also a speaking engagement: http://insidehpc.com/2016/01/call-for-papers-supercomputing-frontiers-in-singapore/ http://insidehpc.com/2016/01/call-for-papers-supercomputing-... and a comment elsewhere that mentions another approach: the "Mill CPU of Mill Computing" As I recollect (perhaps quite wrongly) Itanium (VLIW) failed because compiler-writers couldn't really be bothered or couldn't mount the learning curve. So I'm most curious about what progress is being made on the compiler side.
- smaddox 10y agoYeah, something like this is very much needed, but it's not the hard part. The software is the hard part. The software is the reason we have the multiple levels of cache we have now. Without solving the software challenges, there can be no challenger for the existing architectures. It's interesting to note that convolutional neural nets (CNNs) are one solution to the software challenge. It's an imperfect solution, in the sense that CNNs are not as general purpose (at the same efficiency) and have strict data requirements for training, but it is a solution, and the big N are investing heavily to the point of designing ASICs. Eventually, though, we need to solve the software problem. That will require rethinking programming languages.
- Nomentatus 10y agoTrue. I shoulda added "so to speak", since this is a still more extreme approach and might simply break any compiler/language combination we have, as you say.
- trsohmers 10y agoWhile we have been exploring some ideas on how to have better programming approaches to address the unique features of our architecture, we have from the beginning though that we would be required to have some level of portability for existing applications. As of right now, we support standard C/C++ that runs through our Clang+LLVM backend, with the ability to support any language that has a LLVM frontend. Personally, I find the actor model to be the easiest existing way to take advantage of things like our network on chip and having hard time guarantees on memory movement. That being said, right now our focus is on C and C++ along with our API and custom library ports.
- dnautics 10y agoHaving written programs for this iteration of the REX Neo architecture, the architecture is not so dramatically different that programming languages will have to be rewritten. I'm not the smartest programmer in the world and I was able to figure out the assembly language fairly easily. Some concepts, like how to manage concurrent data processing and thread communications, need to be handled carefully, but that's more at the level of 'standard library' than the compiler. There is a clear pathway to getting C working on the architecture, and a reasonable direction (that will need some fleshing out) to getting performance-enhancing optimization of something like LLVM IR.
- white-flame 10y agoI wouldn't expect the assembly language level to be too far off from the common paradigms. Where I'd expect the software challenges to be would be in managing large amounts of memory, if the application programmer must manage shuffling data between the local scratchpad, specific locations in foreign scratchpads that must be (manually?) DMA'd around, and DRAMs.
- trsohmers 10y agoOur whole goal, as talked about in the software section of our website (and the ACM paper linked in it), is to have the scratchpads be entirely automated by our toolchain. While we want to allow for especially adventurous programmers to have full freedom with the scratchpads, existing and future programs written in C/C++/other languages supported in the future will handle memory allocation identically (from the programmers perspective) as existing architectures. One other thing to point out is that our actually addressing of a cores local scratchpad, as well as "foreign" scratchpads of other cores on the same chip and/or any other attached chip is handled exactly the same. All memory operations are handled through the exact same load/store instructions as part of a global flat address map that is the same for all cores in a system (one or multiple chips interconnected).
- Quequau 10y agoI recall an interview with someone formerly in upper management for the Itanium development project where he acknowledged that the most significant factor in the demise of Itanium was the exclusionary pricing structure Intel imposed on them.
- trsohmers 10y agoYou are correct that we have already taped out, though we haven't made any announcements yet, though will be talking publicly about it in the future with a big focus on the "magic" on the software side. You can read my comments on the Mill architecture elsewhere on HN (not a fan of stack machines), but my biggest disappointment in them is the fact that they have been working on Mill for ~10 years with a team ranging of 5 to 20 (from what I have heard) and have yet to get to silicon, while we have gone from a complete custom architectural idea to tapeout in ~11 months from closing our first seed funding. The big technical failure point for Itanium (in my opinion) is the fact that Intel took the relatively pure VLIW research by Josh Fisher @ HP Labs and tried to add a ridiculous number of features (and attempted x86 compatibility) that impacted the ability to statically schedule instructions. The resulting bastard architecture Intel called "EPIC" (rather than VLIW) had a very difficult job in getting the compiler to generate instruction parallel code since Intel added a huge amount of indeterminism into the architecture that goes against the original VLIW tenets. If your compiler has to assume the worst case latency for all instructions and memory operations, you are going to have a bad time.
- white-flame 10y ago> while we have gone from a complete custom architectural idea to tapeout in ~11 months from closing our first seed funding. To my understanding, the Mill project is not financed. They're enthusiasts working for sweat equity, and are likely going to seek (non-controlling?) investment to finally hit silicon when they're ready. For the scope of what they're doing, I think it's a defensible enough approach. It's not something that can be created in evolutionary stages; all designs of all parts need to be working together properly for there to be benefit from any part, and it's quite complex while also trying out tons of novel designs. (and the Mill isn't stack-based or stack-related. It's basically a crossbar of recent ALU/Load results being fed into further ALU/Store inputs in parallel. The belt is just some way to represent the set of recent results.)
- gpderetta 10y agoItanium failed for the same reason every other VLIW failed as a general purpose CPU: there just isn't enough information a compile time to model the dynamic properties of a program. In fact many of Itanium additions (strange instruction packing, alias disambiguation hardware) were attempts at overcoming this issue. The only moderately successful general purpose VLIW are Conroe and the related Denver, and they use a runtime translation layer to collect the required dynamic informations.
- trsohmers 10y agoFounder of REX here, and surprised to see this posted here. Happy to answer any questions, and you can check my comment history for some of my prior posts on REX. We've had some really great progress that we hope to share in the near future, so stay tuned. EDIT: Since this article is over a year old, we have made a lot of progress, and have recently taped out our first chip. We haven't officially posted a job opening, but we are very shortly going to be looking for software engineers that would love to work on our architecture. Feel free to shoot me an email if you're interested!
- skynetv2 10y agohave you published any white papers detailing any of the following: architecture, instruction set, software availability, benchmarking / application porting and performance etc. I read a couple of times that you got funding from various govt agencies. Most of these funding agencies publish rfp responses or slide decks unless you insisted on an NDA and was approved. I couldnt find any documents talking in depth about your work. I am in the HPC space (academic, research) and am genuinely interested in learning more about your work.
- trsohmers 10y agoWe'll be releasing a whitepaper by September covering the architectural basics, which will coincide with a public release of a SDK. We did have a paper[0] at last years Memsys conference that goes over some of the basic ideas of our compiler, though it is pretty vague (due to our reluctance to share prior to having patent protection at that time). [0]: http://dl.acm.org/citation.cfm?id=2818966 http://dl.acm.org/citation.cfm?id=2818966
- 6d65 10y agoIs it possible to get a Developer Kit for it? It would be great if there would be some raspberry pi like distribution with a Chip included. I think this could speed up the adoption.
- trsohmers 10y ago
- amelius 10y ago> there is no virtual memory translation happening, which in theory, will significantly cut latency (and hence boost performance and efficiency). This means that there is one cycle to address the SRAM, so “this saves half the power right off the bat just by getting rid of address translation from virtual memory.” In protected mode (i.e., what the kernel is using), will an Intel processor not also disable virtual memory lookup? Couldn't we just recompile scientific software to a protected mode environment to get those same benefits? Also, I think it is more useful and fair to compare against a GPU than a general purpose CPU. (As an aside, I don't see where the reduced latency gives such a big advantage. There will be latency anyway, so in any case your software has to deal with waiting in an efficient way (doing useful stuff in the mean time). Shaving off some latency will only help if your software design was bad to begin with.)
- SeanDav 10y agoIt would just be great to get in a decent chip that does not have built in, and unblockable, back-door hacking, like those on Intel, AMD and probably ARM.
- mvdwoord 10y agoI see a good opportunity for government to make this a reality. Not per se a fan of gov regulation for many things but I don't see this moving forward very fast. There are initiatives left and right (e.g. Talos) but if a significantly large government body (EU?) would make it a requirement, that might change the game. Lobbyists would probably convince them otherwise (you need closed HW to catch terrorists... etc).
- dewster 10y agoFrom the article: “Caches and virtual memory as they are currently implemented are some of the worst design decisions that have ever been made,” Sohmers boldly told a room of HPC-focused attendees at the Open Compute Summit this week. As a lay processor designer, I couldn't agree more. I don't like VLIW, but this architecture makes a lot of sense. I think it took up to this point for compiler technology to catch up with what is possible in hardware. Almost all the good ideas in computing were mined out long ago, the trick I think is to get the computing world to give up on those which are holding things back (cold dead hands if necessary).
- ridgeguy 10y agoI'm curious about the thermal issues. From the article, the power density is (4 W)/ (0.1mm^2), or 40W/mm^2. Intel's Haswell chip has a TDP of ~ 65W, an area of 14.7mm^2, for a power density of 4.4W/mm^2. Is this power density a cooling challenge?
- trsohmers 10y agoFirst note: The article is ~16 months old, so is outdated on some measures. I've corrected the numbers below, but in either case, you seem to have been confused between the size of a core and the size (and power) of an entire chip consisting of multiple cores. After tapeout of our first test chip, the final size for one of our cores is 0.27mm^2 (including the SRAM that makes up the scratchpad memory) on TSMC's 28nm process. We actually came in using less gates than originally anticipated, and our size without SRAM is a little less than 0.01mm^2. Now, for just going by what is on the linked article: The diagram comparing sizes are for single cores (0.1mm^2 estimate back then for a Neo core, 14.5mm^2 for a single Intel Haswell core). The power numbers in the table below that are for entire chips. You are quoting 65W for a single core, which is incorrect... The 65W Haswell chip I believe you may be referring to is the 4770S, which is 4 cores @ 65 watts, and looks like it has a die size of 177mm^2. Calculating this out using our current numbers, our planned full 256 core chip has changed a bit (doubled the performance since last year, doubled the power due to adding more stuff) and we estimate the TDP to now be 8 Watts and ~100mm^2, which gives us a power density of 0.08W/mm^2. Intel would then have 65W / 177mm^2 = 0.367W/mm^2. As would make sense in the case where we are claiming lower power operation, our power density is also lower.
- ridgeguy 10y agoThanks very much for clarifying. That the 4W didn't apply to a single core fell through a cognitive crack. The power density is impressively low, indeed. Looking forward to more info in Sept.
- gpderetta 10y agoThis chip was discussed on RealWorldTech a while ago: http://www.realworldtech.com/forum/?threadid=151566 http://www.realworldtech.com/forum/?threadid=151566 Let's say it wasn't well received.
- KKKKkkkk1 10y agoThere is nothing to disrupt. Exascale computing is a haux perpetrated on the US government by unscrupulous hardware vendors. Kudos to Rex for grabbing a piece of that action.