117 ms·
I see this posted every time Mill gets linked. Could someone simply explain the significance of the Mill architecture?
by RyanMiller 12y ago
I see this posted every time Mill gets linked. Could someone simply explain the significance of the Mill architecture?
- willvarfar 12y ago(Mill team, so biased; thanks for the chance to pitch ;) We are a DSP that can run general purpose code. Traditionally, to run general purpose code fast you needed an out-of-order superscalar architecture, as all the x86 and RISC cores are these days. DSPs have substantially better performance and substantially better efficiency, but have traditionally been ineffective executing general purpose code (such as the web browser you are using to read this). The Mill is a synergy of lots of small breakthroughs that together deliver significant improvements to general purpose single threaded code. Its been held that cores have stopped getting faster. We're faster. And we have similar as yet not filed improvements for multicore too.
- rrmm 12y agoCan you say anything about where you are on the path towards silicon implementation or is it all still under wraps?
- willvarfar 12y agoWe are working towards the FGPA. Its very early days, but the HW team is very experienced.
- thesz 12y agoWow. Finally. I've seen this Mill occdasionally pitching and always have been asking the same question: are your results from simulation or from FPGA? Now I know the answer. The most suspicious thing in Mill is that belt thing. To produce operands for N operations you need N*3 (2 for reads, one for write) ports of RAM. For even two operations that means 6 ports. No FPGA allow that out-of-the-box. Given that, you have to implement that in registers and logic, wasting FPGA resources. (AFAIK, silicon fabs also does not have such RAM blocks. you have to build them themselves, either from registers and logic (and make them slow) or using transistors (which make development process slow). this is THE source of relative slowness of Itanium and Elbrus thing from Russia.) If you want an advice, go for Tabula. You'll need many R/W ports per block of RAM, they seem to have those (12 ports RAM blocks). Maybe your design won't be as slow as I think it will.
- Rusky 12y agoIf I understand correctly from what they've said so far, there is no RAM/register file with N*3 ports, it just reuses the outputs of the functional units. That's also nowhere near the source of slowness in the Itanium.
- thesz 12y agoOkay. The reuse of output from functional units was done in TTA CPUs (Transport Triggered Architectures). Guess how they fare if you probable never heard of them? Guess also how fast or slow they are compared to regular OOO CPUs. You will be right if you guess that they are not that good in terms of raw performance and they are not that fast in terms of operating frequency. They are not fast in either way precisely because they use crossbar as a operand delivery network. They also can have FIFOs as the switching network or as an another functional unit dedicated to spreading information, but most often it is not used.
- marcosdumay 12y agoYou don't put a RAM block in the CPU pipeline. RAM is much slower, thus you keep is several subsystems away.
- cfallin 12y ago"RAM" here means SRAM (static RAM). That just means "array of storage elements (latches) connected by bitlines and wordlines", which is way more efficient than "random latches we scattered throughout the chip". SRAMs are used extensively for indexed storage such as physical register files, queues, predictor arrays, etc in modern microarchitectures.
- marcosdumay 12y agoAnyway, unless you have a very small array, even addressing is enough to make it slow. And arrays were already big at the time I programmed FPGAs, I can only imagine they are much larger now.
- 12y ago
- RyanMiller 12y agoCheers for the pitch. He wasn't kidding about being fascinating. Could probably lose that synergy, though. Can't wait for some more real competition in the architecture space. Are you being funded by any major silicon giants or are you being backed at all? How many years do you think before you reach some sort of manufacturing or are you still in the "When it's done" phase?
- igodard 12y agoNo funding by giants; not a public company. The SEC rules prevent us from talking further (we're too busy to go to jail for breaking the securities regs), but if you are interested in the business side of the company then you can sign up at MillComputing.com/investor-list; it's a low-traffic mailing list where we announce opportunities. In heavy semiconductor you don't really move out of "when its done" until the FPGA proof-of-principle is working. That's over a year plus "when it's done" :-)
- tomjen3 12y agoNormally I am a cynically old bastard, but if that is true then you really are the most important thing happening in hardware right now (or possible HPs the Machine if it isn't vaporware). One crucial question - are you compatible with X86?
- ema 12y agoThe mill is not compatible with x86, but the goal is to not require more than a recompile.
- vonmoltke 12y agoThat would be awesome, if accomplished. It will not be easy, though. My experience with TI DSPs and the POWER6 (the most recent major in-order processor) taught me that we are currently a long way from that with existing compilers. Even x86-64<->POWER required some platform-specific code for performance-sensitive blocks.
- alex-g 12y agoThe Mill certainly raises a lot of interesting code generation and optimization issues. I'm sure there's plenty of scope for figuring out good optimization strategies, as a lot seems to depend on the ability of the compiler to make good choices about instruction scheduling and belt slot allocation. Sure, that's the case for traditional architectures as well, but there's more prior art there too. There may also be ingenious algorithms which work better on the Mill architecture specifically. I'd love to know if there's any theory on the hardness of allocating positions on the belt, compared to traditional register allocation.
- willvarfar 12y agoActually, as its co-designed by a compiler writer (read the bio we paste with the talks: > Ivan Godard has designed, implemented or led the teams for 11 compilers for a variety of languages and targets, an operating system, an object-oriented database, and four instruction set architectures. He participated in the revision of Algol68 and is mentioned in its Report, was on the Green team that won the Ada language competition, designed the Mary family of system implementation languages, and was founding editor of the Machine Oriented Languages Bulletin. He is a Member Emeritus of IFIPS Working Group 2.4 (Implementation languages) and was a member of the committee that produced the IEEE and ISO floating-point standard 754-2011. ), its actually designed to be easy to write a compiler for. It's he polar opposite of the "sufficiently smart compiler syndrome" :) I keep suggesting we do a "sufficiently dumb compiler syndrome" talk, but it'd contain nothing novel; the art is well established by all the VLIW machines that have come before.
- hahainternet 12y ago> And we have similar as yet not filed improvements for multicore too. Goodbye mutexes?
- igodard 12y agoMill uses optimistic concurrency, similar to the IBM and Intel versions. From that you can build mutexes if you are willing to put up with the drawbacks of locking.
- Keyframe 12y agoThat's very interesting! What are your thoughts on RISC-V?
- thesz 12y agoRISC-V looks nice. It avoids raising exceptions wherever possible, which I like EXTREMELY. This saves space, allows for faster hardware and makes life of systems/compilator programmer easier. It should be praised for that matter alone.
- gizmo686 12y agoVery impressive sounding work. If the Mill came out today, it would be the first time I considered buying a computer component just because I was interested in it. If you can answer this, what is the bussiness plan for the Mill? Who do you expect to buy it when it comes out. Has any company expressed interest is using it?
- willvarfar 12y agoIvan talks business models in this hackaday interview: http://hackaday.com/2013/11/18/interview-new-mill-cpu-architecture-explanation-for-humans/ http://hackaday.com/2013/11/18/interview-new-mill-cpu-archit... Hope this helps!
- badsock 12y agoLately I've been looking into having CPUs run untrusted machine code (e.g. Google's Native Client). Hope you don't mind a completely off-topic question: is it anywhere on your team's radar to have provisions for that? So far it's been essentially chance whether the architecture has useful features to make this happen (e.g. x86 segment register abuse).
- willvarfar 12y agoIt's very much on our radar. The team have a lot of experience with capability systems, and although the Mill is not a capability based architecture, it is much finer grained than mainstream e.g. x86. There is a security talk http://millcomputing.com/topic/security/ http://millcomputing.com/topic/security/
- badsock 12y agoAwesome, great to hear and thank you for the link!
- Guvante 12y ago> And we have similar as yet not filed improvements for multicore too. Any chance you have a solution to the cache dilemma? Unified cache's are brilliant for simplifying implementations but can often lead to stalling.
- igodard 12y agoMill multicore has fully sequentially consistent cache coherency; there are no barrier operations. Sorry, how it's done is still NYF (Not Yet Filed). We expect a talk on the subject this fall.
- deleted 12y ago[deleted]
- ansible 12y agoWe've got Mill team members in this thread, but anyway here's the short, short version: It's about using as much of the die area on the CPU chip for actual computation, rather than supporting an instruction set with outdated design. The programmer's model of computation today has little correspondence to the realities of current semiconductor process technology in terms of what's fast, and what's easy to implement in hardware. The Mill is a bottom-up redesign that takes into account many of the design constraints with current technology, and attempts to design a good architecture that can maximize actual computational throughput.