6 ms·
I might be wrong. But if they automated the flow from RTL to GDS, the timing might not be optimal. I understand since they have lack of resources so that this i
by codebook 10y ago
I might be wrong. But if they automated the flow from RTL to GDS, the timing might not be optimal. I understand since they have lack of resources so that this is unavoidable but in normal chip design flow, the backend timing ECO is critical to achieve high frequency for all timing corners.
- Coffeewine 10y agoI agree completely, it's still impressive to me that they presumably managed a competitive offering with such a system. I imagine having it be a highly homogeneous design also helped.
- adapteva 10y agoDesign symmetry and regularity was the key. Harder to achieve that with a heterogeneous architecture.
- adapteva 10y agoYes, we are leaving 2X on the table in terms of peak frequency compared to well staffed chipzilla teams. Not ideal, but we have a big enough of a lead in terms of architecture that it kind of works.
- sitkack 10y agoThat seems analogous to human assembly optimization vs a compiler. But the time to market is greatly reduced, designs can be vetted and a 2.0 that is optimized for frequency can be shipped later.
- yellowapple 10y agoIIRC, human assembly optimization is unlikely to be better than a modern compiler nowadays. Same thing could very well happen for this "automated flow" if it starts incorporating its own optimization techniques.
- maccard 10y agoCompilers are smart at some things and not so smart at others. I can beat the compiler in tight inner loops almost every time, but it will also do insanely clever things that id never think of!
- wolf550e 10y agoThat is a myth. Most developers can't beat LLVM. LLVM can't beat the handcrafted assembly in libjpeg-turbo or x264 or openssl or luajit by compiling the generic C alternative.
- pcwalton 10y agoIf it's not at least able to match handcrafted assembly using intrinsics, you should file bugs against LLVM. There is no theoretical reason why compilers shouldn't be able to match or beat humans here: these problems are extremely well studied.
- robryk 10y ago"using intrinsics" is a cop out: you are essentially doing the more complicated part of translating that sequence of generic C code into a rough approximation of a sequence of machine instructions and leave the compiler to do the boring and simpler parts, like register allocation, code layout and ordering of independent instructions.
- pygy_ 10y agoYou may want to read this Mike Pall post about the shortcomings of high level language compilers regarding interpreters: http://article.gmane.org/gmane.comp.lang.lua.general/75426 http://article.gmane.org/gmane.comp.lang.lua.general/75426
- aseipp 10y agoSometimes consistency is desirable, as well as performance. Compilers are heuristic. They evolve and get better, but they can mess up, and it's not always a fun time to find out why the compiler made something that was performance sensitive suddenly do worse, intrinsics or not -- from things like a compiler upgrade, or the inlining heuristic changes because of some slight code change, or because it's Friday the 13th (especially when it's something horridly annoying like a solid %2-3 worse -- at least with %50 worse I can probably figure out where everything went horribly wrong without spending a whole afternoon on it). This is a point that's more general than intrinsics, but I think it's worth mentioning. Sure, I can file bug reports in those cases, and I would attempt to if possible -- but it also doesn't meaningfully help any users who suddenly experience the problem. At some point I'd rather just write the core bit a few times and future proof myself (and this has certainly happened for me a non-zero amount of times -- but not many more than zero :)
- nickpsecurity 10y agoThe comment above said you couldn't release the info due to the EDA vendor. However, people like Jiri Gaisler have released their methodologies via papers that just describe them with artificial examples. Others use non-manufarable processes and libraries (like NanGates) so the EDA vendors feelings don't get hurt about results that don't apply to real-world processes. ;) So, if you have a 16nm silicon compiler, I encourage you to pull a Gaisler with a presentation on how you do that with key details and synthetic examples designed to avoid issues with EDA vendors. Or just use Qflow if possible.
- adapteva 10y agoI'll pass for now...Gaisler is in the business of consulting, we survive by building products. I am happy to release sources, but it's completely up to the EDA company. [edit: was thinking of the wrong Gaisler, still will pass]
- nickpsecurity 10y agoDamnit. No promises but would you consider putting it together if someone paid your company to do it under an academic grant or something? Quite a few academics trying to do things like you've done with small chance that one might go for that.
- nickpsecurity 10y agoBtw, your site is down right now.
- gonzalocasas 10y agoIt's pretty ironic that parallella.org is down on an article about high parallelism because -apparently- it cannot take HN-front-page-load-levels.
- nickpsecurity 10y agoThat's concurrency, throughput, and load-balancing of web servers connected to pipes of certain bandwidth. It's not the same as parallel execution of CPU-bound code on a tiled processor. You could know a lot about one while knowing almost nothing about the other.
- runeks 10y agoThe interesting question, to me at least, is how much cheaper this chip is - with its suboptimal maximum clock rate - compared to a chip from a non-automated flow. If peak clock rate is one half, but cost is one hundredth, I'd say it's a spectacular achievement. 100th in costs and one half in performance is, granted, wishful thinking on my part. But I believe the important point is that with a sufficient productivity gain, this technology can reduce the old, non-automated way to something akin to writing software libraries in assembly. Writing software libraries in assembly is useful, but few bother to do it because they'd rather just buy more hardware. Chugging out twice a many chips, once you have your design finished, isn't really that much more expensive, as I understand it.