12 ms·
The “high-level CPU” challenge (2008)
- greenyoda 10y agoThe comments following this article (which span a period from 2008 to 2015!) are also very interesting.
- kevinwang 10y agoAnd interesting comments from 2008 and 2015 here: https://hn.algolia.com/?query=The%20“high-level%20CPU”%20challenge&sort=byDate&dateRange=all&type=story&storyText=false&prefix&page=0 https://hn.algolia.com/?query=The%20“high-level%20CPU”%20cha...
- yosefk 10y agoAuthor here. I think today that apart from making my case in a bit of an obnoxious tone, I also somewhat overstated it: while it's true that many "high-level" constructs do have a cost that will not magically go away due to any logic built into hardware, at least not fully, it is also ought to be true that a lot can be done in hardware to make software's life easier given a particular HLL programming model, and I'm hardly an expert on this. My true interests are in accelerator development so starting at the GPU and further away from the CPU and so lower level and gnarlier than C in terms of programming model. I will however say that the Reduceron and in general the idea of doing FP in hardware in the most direct way are a terrible waste of resources and I'm pretty sure it loses to a good compiler targeting a von Neumann machine on overall efficiency. The way to go is not make a hardware interpeter, that is no better than a processor with a for loop instruction added to better support C. The trick is to carefully partition sw and hw responsibilities as in the model to which C+Unix/RISC+MMU converged to.
- alain94040 10y agoJust curious, what did you mean by this: If your architecture meets these requirements, I'll consider a physical implementation very seriously (because we could use that kind of thing), and if it works out, you'll get a chip Fabbing someone else's idea sounds expensive. What did you have in mind?
- abainbridge 10y agoPeople put experimental digital blocks in ASICs all the time, it isn't necessarily that expensive. Its a bit like taking an extra pair of shoes on holiday - in general, I'm a bit concerned about reaching the airline's baggage weight limit. If you ask me to add your shoes to my bag before I start packing, I'm going to say no. If you ask at the end, and I've got some space left, then fine. However, he's probably talking about an FPGA implementation. That'd be sufficient to prove the concept. Once you've gone that far, you can normally do some simulations to predict the energy consumption on a real chip.
- _yosefk 10y agoI was at the time of writing and still am an accelerator architect and I'd gladly use someone's idea in a mass market product (ASIC) if they didn't mind. However, it is also true that working on any real product means that many valid ideas useful in some contexts will not be useful for me, and I guess this is true for many ideas for speeding up higher-level programming models, and perhaps it was misleading of me to fail to point this out. (As I said I don't love the tone of that article, it is unfortunately very effective as my articles written in that tone around 2008 tend to resurface more often than articles written in nicer, more balanced tone and with way more technical details from around say 2012-2013. What is my takeoff wrt future writing I'm still not quite sure.)
- justin66 10y agoAlan Kay takes part in discussion here occasionally. He seems like a pretty easygoing guy (but you might want to reign it in a bit) so you could probably just email him...
- _yosefk 10y agoAlan Kay likes to fail at least 90÷ of the time (shows you aim high enough) and says the industry is too dumb to digest good ideas. I like to succeed at least 90÷ of the time so that the dumb industry keeps employing me. I'm afraid we have irreconcilable differences. (And, this is me reigning in A LOT right here. Don't get me started...)
- grzm 10y ago90÷ To be clear, you mean 90%, correct? If so, the ÷ symbol (which I just learned is called an obelus) is typically used for division—I've never seen it used to mean percent. Is this a locale difference? A keyboard issue? I've seen that Android users will sometimes mistype ℅ for % due to their proximity on a certain keyboard, for example.
- andai 10y agoMaybe they were holding their phone at an angle?
- deleted 10y ago[deleted]
- comboy 10y agoCan you elaborate a bit on why do you think Reduceron or using FPGA along with CPU is not a good idea? I thought that since the clocks aren't gonna be much higher, that is the future. That maybe compilers will start generating some kind of VHDL that can make the app you spend your most CPU time on much faster (theoretical possibilities seems great with big enough FPGAs).
- adwn 10y agoSpeaking as someone who programs FPGAs for a living, they are good for three things (I'm simplifying a bit here): * interfacing with digital electronics, implementing low-level protocols, and deterministic/real-time control systems * emulating ASICs for verification * speeding up a small subset of highly specialized algorithms Of those, only the last one would apply in the context of this thread. However, caused by their structure and inherent tradeoffs, they are completely incapable of speeding up general purpose computation. As for specialized computation, if they heavily rely on floating-point ops, a GPU will nearly always be faster and cheaper.
- _yosefk 10y agoI think the new Stratix might well beat GPUs, not? The Reduceron specifically tries to quickly perform application of lambda expressions that GHC will try to avoid generating in the first place. The Reduceron speeds things up using several memory banks etc. but it still does things that shouldn't be done at all and the overhead is there at least in area and power.
- simias 10y agoI think FPGAs are too expensive to be used for general purpose computing. If on top of the chip price you add the development time it's just not cost effective. A high-end FPGA will cost you thousands of dollars and you won't be able to easily convert software code to HDL. A very high end GPU will be cheaper and easier to develop for. There are situations where a FPGA is better suited of course (very low latency real time signal processing for instance) but for general purpose computing FPGAs are not exactly ready for primetime IMO.
- nostrademons 10y agoI'm curious whether you think the ideal boundary between SW/HW might've shifted in the last ~40 years, as the things we use computers for have drastically changed? I know basically nothing about hardware, but I know the software layer from the OS/compiler up through the UI. There's a fair bit of evidence that things we've traditionally assumed belong in the kernel actually belong in userspace, and they're being reinvented in userspace as a result. For example, most modern languages & frameworks put some form of scheduler in the standard libs - we're reimplementing the abstraction of a thread as promises or fibers or async/await or callbacks. Many big Internet companies disable virtual memory in their production servers, because once the box begins swapping you might as well count it as down. Many common business apps program to a database, not a filesystem, and then the database uses block-based data structures like B-trees and SSTables but then has to implement them on top of filesystems. At the same time, the classic OS protection boundary is the process, but the unit of code-sharing in the open-source world is the library. As a result, the protection mechanisms that OSes have gotten very good at are largely useless at preventing huge security violations from careless coding in a library dependency. Most of these came from computers being used outside of the original domains that the system software developers assumed, eg. nobody in the 1970s would've imagined 10 million GitHub users of widely different skill levels all swapping code. Knowing what we do now about the big markets for computation, are there additional operations we'd want to put in hardware, or things currently done in hardware that should be moved to software?
- hueving 10y agoNit: disabling swapping is not the same this as disabling virtual memory. Virtual memory is just something that allows swapping, but does not require swapping.
- qznc 10y agoThere is a fascinating talk by Cliff Click "A JVM Does That?" At the end he shares some opinion about what should be done by JVM or OS and what should change. video: https://youtu.be/uL2D3qzHtqY https://youtu.be/uL2D3qzHtqY slides: http://www.azulsystems.com/blog/wp-content/uploads/2011/03/2011_WhatDoesJVMDo.pdf http://www.azulsystems.com/blog/wp-content/uploads/2011/03/2... I remember a talk, where he was also talking about hardware. It does not seem to be this one. For example, a time register would be useful. Syscalls like clock_gettime are too slow. CPU info like cycle counts fail with dynamic frequency scaling.
- BuuQu9hu 10y agoActor-based dynamic language author here. (Doesn't matter which one; I think I speak for all of us.) Thank you for being honest with us; we are not a very performance-oriented group sometimes. We're generally in favor of things which accelerate message passing between shared-nothing concurrent actors. Hardware mailboxes or transactional memory are nifty. OS-supported message queues are nifty; can those be lowered to hardware in a useful way?
- ori_b 10y agoIt seems to me that Intel TSX and general work on improved atomics is what you want. The high level constructs themselves probably shouldn't be directly in hardware.
- qznc 10y ago> Hardware mailboxes or transactional memory If you do them in hardware, they always come bounded. At most n elements of size m bytes and both numbers usually single digit. If you want to lift that limitation, it usually is just as slow as doing it in software.
- dkersten 10y agoUnbounded queues are arguably not a good idea (although single digit bounds are possibly too low?), at the very least there would probably need to be some concept of back pressure.
- meredydd 10y agoWell, I never thought I'd be plugging my PhD research here, but: "Asynchronous Remote Stores for Inter-Core Communication" http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.592.1178&rep=rep1&type=pdf http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.592... To my knowledge, this is still the only hardware-assisted message passing scheme that is virtualisable (ie compatible with a "real" OS like Linux). Hardware mailboxes are great, but time-sharing OSs can't deal with finite hardware resources that can't be swapped out easily. Software-based queues die a fiery death thanks to cache coherency - reading something that another core just wrote will block you for hundreds of cycles.
- aidenn0 10y agoOne simple and concrete request: Traps on integer overflow: http://blog.regehr.org/archives/1154 http://blog.regehr.org/archives/1154
- deleted 10y ago[deleted]
- dnautics 10y agoYossi, do you think there's a market for special matrix mult machines that use low-precision FP, maybe as a systolic array?
- _yosefk 10y agoThis is OT, but - most of linear algebra dies quickly even in single precision, meaning that your equation system solver produces a solution that doesn't really solve these equations, etc. One exception is neural networks where Google's TPU is just the start and in general GPUs while beating CPUs leave a lot of room for improvement.
- naasking 10y agoA promising alternate architecture that places some previously features in hardware [1]. The execution model still closely matches current architectures. See also different approaches to programming that make space-time tradeoffs more explicit [2], and use natural-law like principles to distribute computing across a simpler but highly connected computing fabric. [1] https://millcomputing.com/ https://millcomputing.com/ [2] http://web.mit.edu/jakebeal/www/Publications/PTRSA2015-Space-Time-Programming-survey-preprint.pdf http://web.mit.edu/jakebeal/www/Publications/PTRSA2015-Space... [3] http://blob.lri.fr/ http://blob.lri.fr/
- jstimpfle 10y agoWhat do you mean by "for loop instruction added"? How would that look from a developer's perspective and what could be done in hardware to improve efficiency?
- _yosefk 10y agoI used that as an example of a bad idea; I don't have details on this bad idea but you could have an instruction looking at an init, bound and increment registers and a constant telling where the loop ends and voila, the processor runs for loops without needing lower-level increment and branch instructions, and it shaves off one instruction (not a cycle, necessarily, but an instruction): FOR counter_reg, init_val_reg, bound_reg, step_reg, END_OF_LOOP ... END_OF_LOOP: ...instead of: MOVE counter_reg, init_val_reg START_OF_LOOP: ... ADD counter_reg, step BRANCH_LESS_THAN counter_reg, bound_reg, START_OF_LOOP I was saying that this obviously not-so-good idea is not much different in spirit from building hardware for quickly creating and applying lambda terms, which is what the Reduceron does. Lowering lambda calculus to simpler operations so that lambda expressions are not represented in a runtime data structure at all much of the time, the way GHC and other compilers approach the problem, is a better idea.
- jstimpfle 10y agoI see - sorry, I'd missed the "not" in "the way not to go"
- deleted 10y ago[deleted]
- Symmetry 10y agoWell, the Reduceron seems to count as an example. I'm not sure I'm convinced about it's performance, though. https://www.cs.york.ac.uk/fp/reduceron/ https://www.cs.york.ac.uk/fp/reduceron/ That's specialized for just one language, though. In general you can always speed things up, sometimes by quite a bit, if you're willing to make your general purpose computer somewhat less general purpose. Some of what the Mill folks are doing with hardware assisted stack operations might fall under the category of higher level instructions but those are for C just as much as any other language. https://millcomputing.com/ https://millcomputing.com/ EDIT: Oh, and Linus likes to wax eloquent about the wonders of rep movs and I think he sort of has a point about having good facilities to call routines specific to the hardware, using instructions specific to the hardware that aren't exposed in the public ISA. But again, that's a high level function in hardware but it isn't specific to a high level language and it's mostly about accelerating C.
- catern 10y ago1. Eliminate cache coherency protocols (replacing it with cache manipulation/inter-CPU communication instructions) 2. Eliminate virtual memory (replacing it with nothing) I'm not a CPU designer, but my understanding is that removing features allows for a denser/faster CPU. Well, these are two features that a suitably high-level language has no need for, because a high-level language doesn't expose "memory" to the programmer. Edit: Though, I 100% agree with what I believe is the core point of the author. We should not implement high-level features in hardware. In fact we should implement as little as possible in hardware, moving as much as possible into software. If Intel would let third parties generate microcode for their CPUs, we could move a lot further in that direction...
- Symmetry 10y agoUsually replacing software with dedicated hardware tends to speed things up and I don't think it's obvious that using explicit communication would actually speed things up. Most memory operations don't require any form of sharing and having the sharing that's required happen automatically seems efficient. Getting rid of virtual memory is potentially a big win, especially for architectures where you can't make the L1D cache virtually indexed but physically tagged. And in general there are a lot of special cases you don't even have to think about if different memory addresses can't alias to the same memory. You do lose out on a lot of software tricks there, though.
- chrisseaton 10y ago> Most memory operations don't require any form of sharing Maybe that was the point? Most of the time it isn't needed, but you are still paying for the logic to detect when it does, and you are paying for false sharing when the automagic gets it wrong.
- dman 10y agoFinally I run into someone who shares my exact pet peeves!
- _yosefk 10y ago#2 saves a little but not much and precludes unsafe low-level code completely. #1 I think impairs many HLLs, certainly every multithreaded imperative shared memory ones (and IMO nothing is close to these in terms of efficiency on multicore); which languages can work well without cache coherence? (I worked in C++ on multicore with no hw coherence btw. Quite the cruel and unusual punishment.)
- rogerbinns 10y agoIntel did try to introduce a high level CPU in 1981: https://en.wikipedia.org/wiki/Intel_iAPX_432 https://en.wikipedia.org/wiki/Intel_iAPX_432 It failed due to very poor performance. There is an excellent paper by Bob Colwell about why the performance turned out the way it did. Prior HN discussion: https://news.ycombinator.com/item?id=9447097 https://news.ycombinator.com/item?id=9447097
- nickpsecurity 10y agoThe i960 was a much better attempt. Baseline version had just enough smarts to improve safety or reliability while still overall a fast RISC. Got some customers in embedded.
- jeffsco 10y agoI loved this (paraphrased) quote: To quote Ken Thompson (from memory) – "Lisp is not a special enough language to warrant a hardware implementation. The PDP-10 is a great Lisp machine. The PDP-11 is a great Lisp machine. I told these guys, you're crazy."
- Animats 10y agoThe history of "higher level" instructions isn't good. The DEC VAX had an assembly language intended to make life easier for assembly programmers, but it slowed the machine down. The Intel IAPX 432 had lots of bells and whistles, but was really slow. The RISC machines with lots of registers turned out not to be all that useful, and too much register saving and restoring was required. RISC is a win until you want to go superscalar and have more than one instruction per clock. Then it's a lose. Stack machines that run some RPN form like Forth or Java code have been built, but don't go superscalar well. A useful near-term feature would be zero-cost hardware exceptions on integer overflow. This is an error in both Java and Rust, and tends to be turned off at compile time because it has a performance penalty. The problem is that people will want to be able to unwind and recover, which means exact exceptions and a lot of compiler support for them. If you could figure out how to do zero-cost subscript checking, that would be a step forward. That check needs additional info about bounds, which usually means a delay. I used to be a fan of schemes for safely calling from one address space to another. i386 almost has this, with call gates, which don't quite do enough to be useful. A few machines have had hardware context switching, but that hasn't been a big performance improvement. All that has to be tightly integrated with the OS or it's a lose. It's an enhancement to Plan 9, not anything anybody uses. The same is true of fancy schemes for inter-CPU communication, but that probably needs more attention. Like it or not, we have to figure out what to do with large numbers of non-shared-memory CPUs. Some way to set up memory-safe message passing between non-shared-memory CPUs without involving the OS after setup would be useful. An IOMMU that allows drivers in user space with minimal performance degradation is a good thing. Those exist.
- i336_ 10y ago> Stack machines that run some RPN form like Forth or Java code have been built, but don't go superscalar well. I've been interested in Forth (and related stack) processors for a while, and my armchair observations over a few months have suggested that the (much-vaunted) performance gains associated with such processor designs are apparently not straightforward to map or relate to or take advantage of. I remember (unfortunately not sure where right now) reading how the GA144 was built at a time when 18-bit memory was the current trending novelty and that it's not really a perfect processor design. I'm still fascinated by it though (sore lack of on-chip memory notwithstanding). What sort of scale are you referring to when you say "superscalar"? 144 processors? 1000? Do stack-based architectures remain a not-especially-practical-or-competitive novelty, or are they worth pursuing outside of CompSci? (FWIW, everything else you've written is equally interesting, but slightly over my current experience level)
- BuuQu9hu 10y agoI would love to know whether alternatives to floating-point, such as unums or quote notation, are worth considering for those languages which have rationals in their numeric towers.
- trsohmers 10y agoWell unums in particular (especially the new type 3 "posits" ones) are very interesting and practical for hardware implementation. John Gustafson, the creator of unums, gave the first talk on type 3 unums at Stanford earlier this month, and there are already multiple hardware implementation efforts (one of which I am indirectly overseeing). While many people have recently hyped up low precision floating point for things like machine learning, Posits provide greater dynamic range AND greater accuracy for very few bits (as low as four, but with good ML starting 8 to 10 bits). On top of all of that, it costs less area and power in silicon than IEEE hardware. Check out the talk: https://www.youtube.com/watch?v=aP0Y1uAA-2Y https://www.youtube.com/watch?v=aP0Y1uAA-2Y
- razorunreal 10y agoWait - Pure functional languages making lots of copies and lacking side effects is supposed to be a bad thing from a hardware perspective? As I understand it, synchronising shared memory is a massive source of complexity and stalls in modern processors. You get better performance in multithreaded code when you churn through new memory rather than mutate what you've already got, and even better in a stream processor where you know that is what is going to happen. Still, I suppose that's not a great argument for custom hardware because code that would benefit could most likely be shoehorned onto a GPU. Or maybe that's an argument that GPUs are already pretty close to being the custom hardware that we need.
- smitherfield 10y ago> You get better performance in multithreaded code when you churn through new memory rather than mutate what you've already got My understanding is this is related to aliasing at the software, language or compiler level, not the hardware level.
- qznc 10y ago> synchronising shared memory is a massive source of complexity and stalls in modern processors Correct. > making lots of copies is supposed to be a bad thing from a hardware perspective Yes, because memory accesses are costly. You want your data to be packed tightly in memory instead of chasing pointers all the time. An array is more efficient than a linked list mostly due to caches. The balance is key. You want one copy per core, but not one copy per data update.
- gravypod 10y agoIf you want a "high-level" computer that can easily run high-level languages you can try building a stack machine. Check out the old LISP machines from the eairly 80s. High level and fast for the time with some of the most interesting compiler design work. Some other possible crazy ideas: * Write instructions for forcing a read/write of memory into L1 cache Allow me to tell the CPU to keep a chunk of memory im L1 for the lifetime of my program. I'll save you the transistors for figuring this out on the chip and it'll be easy to implement in a compiler. This is stuff that was done back in the NES and gameboy days (although by hand) with hi and lowmem. * Put an easily interfaced FPGA on the die. Get rid of the stupid "one-size-fits-all" vectorization hardware we have today. We can just write our own vectorizations ourselves in the compiler level. Just make a big enough FPGA and we'll do all the math as fast as we want and in the parallelization we want. This removes a very specific job oriented bit of transistors and allows it to be used for pretty much any problem that's complex enough to warrent attempting to use vectorization. One use of a combination of these ideas is writing an image filtering algorithm. It's a program that loops through the rows of my image, runs the "L1 cache" instruction from earlier, then passes it to my FPGA code which applies some complex filter, then writes it back to memory and continues. With traditional systems you'd be limited to 128 wide segments but I can presumably configure my FPGA segment to read the entire block of cache I want to load into and write back to it all as fast as possible. This is an extremely hard thing to build but on-die FPGAs would change the face of computing. When they get big enough, who needs GPUs? * Make core-shared fully atomic (large) registers for extremely fast but crude IPC This is how most high level languages operate so I'd imagine this could definetly be made use of. * Make the idea of a Core to Core MPI via interrupts and allow them to be configurally "synchronious" (page the program out until handled) This allows the idea of calling a function on an object running in parallel with "this" object and it will be returned when this function is returned. These are really crazy ideas. If I had inifnite money, time, and resources I probably couldn't pull something this "out there" off. I think the on-die FPGA is possible but impossible to get right. Most features on a modern processor that take up space are a conbination of functions that could all be done on an FPGA rather then wasting transistors on that specific function. In a server do I really need 10 million transistors on my CPU for H264 decoding? Replace that with cache and a live-reprogrammable FPGA and we can do that all in software when and if we need it.
- Const-me 10y agoGreat article. > I have images. I must process those images and find stuff in them. I need to write a program and control its behavior. You know, the usual edit-run-debug-swear cycle. What model do you propose to use? Looks like you need a GPU, not a CPU. Much image processing stuff (and also much neural network stuff) is very suitable for the programming models of modern GPUs. For a first prototype buy an nVidia GPU and use CUDA. That will unlikely work for embedded stuff but if it’ll work OK on your PC with CUDA, there’s almost 100% chance you’ll be able to do that in OpenCL/DirectCompute/whatever.
- piinbinary 10y agoSome assorted ideas: * Expose the data dependency graph directly to the processor, rather than forcing the processor to infer it from the instructions. * Annotate when data is read-only, to reduce communication between cores (for the sake of avoiding latency, not for bandwidth savings). * Add a mechanism for much cheaper (if more limited) parallelism where a single core would work on multiple, related thunks in parallel, even if those units of work would be far too small to be worth coordinating with another thread to offload. This would likely be largely implicit, taking advantage of the first feature. * Instructions for graph traversal. The CPU could (for some uses) order the traversal in a way that improves cache locality, based on how the graph is actually laid out in memory (and prefetch uncached nodes while working on cached ones). * Something like map & reduce, where you can apply a small (pure) function to a list of data. Again, this would likely be done in parallel.
- qznc 10y ago> data dependency graph directly to the processor What format would you propose? How is it different to instructions? > Annotate when data is read-only What overhead would you save? On the cache coherence protocol level? > a single core would work on multiple, related thunks in parallel VLIW https://en.wikipedia.org/wiki/Very_long_instruction_word https://en.wikipedia.org/wiki/Very_long_instruction_word > instructions for graph traversal Your Intel CPU can already do prefetching. Others have special instructions for it. Hyperthreading does the "work on cached stuff first" part. > like map & reduce Use the GPU.
- yvdriess 10y agoAnswering with my own take on the subject. > What format would you propose? How is it different to instructions? It would still be instructions, but either by reworking what their operands are and introducing a destination (instruction address + operand port). Another option that could maybe work is a second and mirroring instruction stream that contains only the dataflow/dependencies. > What overhead would you save? On the cache coherence protocol level? Current strategy is write-invalidate, I believe. In some heavy-contention situations (e.g. lots of spin-loops on one variable), using a 'write-once' instruction would traffic. Thinking more radically: With an overhaul of the entire virtual-memory system, a write-once instruction becomes potentially very interesting (Jack Dennis had an interesting paper on the subject: http://www.cs.ucy.ac.cy/dfmworkshop/wp-content/uploads/2014/08/DFM2014-9-On-the-Feasibility-of-a-Codelet-Based-Multi-core-Operating-System.pdf http://www.cs.ucy.ac.cy/dfmworkshop/wp-content/uploads/2014/...). > VLIW Only if you mean something like Mill (https://en.wikipedia.org/wiki/Mill_architecture https://en.wikipedia.org/wiki/Mill_architecture), instead of Itanium.
- CapacitorSet 10y agoThis is a good time to mention asynchronous architectures: https://en.wikipedia.org/wiki/Asynchronous_circuit#Asynchronous_CPU https://en.wikipedia.org/wiki/Asynchronous_circuit#Asynchron... Intel itself [claimed](http://stackoverflow.com/a/530494 http://stackoverflow.com/a/530494) that the async CPU performs better than the sync one, but they didn't pursue the project further for lack of large-scale profitability.
- x0x0 10y agoI think yosef misinterpreted JWZ. In large applications a bump pointer allocator plus generational gc is really fast (yes, stack allocation is fast too, but you can't always do it). A compacting gc avoids arenas, object reuse, or other awfulness. gc enables lock-free / CASed data structures; otherwise memory ownership is too complex to implement (though there's a Rust guy doing really cool stuff [1]). And gc in a threaded program is wildly easier. Unless you liked eg COM-style explicit ref-counting. As for lisp plus large programs, the large program I worked on did end up with a bespoke (unfortunately) type-free internal language in order to orchestrate itself. Large = low millions of LOC of c++. [1] https://aturon.github.io/blog/2015/08/27/epoch/ https://aturon.github.io/blog/2015/08/27/epoch/
- deleted 10y ago[deleted]
- rjsw 10y agoThe article doesn't seem to consider the SOAR & SPARC family of RISC CPUs, they had instructions to do arithmetic on tagged integers and trap if the tags were incorrect. The feature has been dropped from 64-bit SPARC though.
- kainolophobia 10y agoThe author is stuck on building a better horse. Not to beat this analogy dead, but the reason Alan Kay et al. are so quick to discuss alternative computing methods should be quite obvious to anyone who doesn't limit their worldview to concepts that humans are already using. Right now most processors are ridiculously general. They take a handful (ok, a couple thousand or so) instructions and they do their best to parallelize the instructions both on a single loop (core) and multiple cores. These instructions are of the "add, multiple, load, store" variety, with a few additional instructions for machine learning[1] and whatever HP wants[2]. This is it. This is the state of computing. How do bees work? Why can spiders hunt? When did crows start using tools? What makes us different than bonobos? How are all of these creatures so capable, yet so energy efficient? We are taking a single solution, RISC/CISC architecture, and brute-forcing the hell out of it. Rather than build adaptive or purpose-built hardware, we're stuck on this concept of compile everything to x86/ARM and shrink the transistors (or try and offload parallel number crunching to the GPU). What the author fails to realize is that computers are just fancy looping mechanisms. We use "HLLs" to compile abstract loops into instructions that run on general purpose machines. That's it. The "apparently credible" people see the world in this light. They understand that the solution we've chosen is subpar, but the physics will make it work for some time. A few other commenters have mentioned FPGAs. I'm not here to pitch a future on FPGAs; the die is still flat, the gates can only be reprogrammed so many times and they're generally "expensive." I will say that we need better tools. FPGAs are a good start. Intel knows this[3]. Microsoft knows this[4]. With an FPGA you can dynamically program the exact logic a given operation will need. Whether it's real-time signal analysis, AI-built logic, or memcached, your logic will run exactly as specified. Using purpose-built logic to run functions "natively" will drastically improve the efficiency of computation; both in time and energy. It's really hard to build a horse that will fly to the moon. It's a lot easier to build a spaceship that can carry a horse to the moon. [1]http://lemire.me/blog/2016/10/14/intel-will-add-deep-learning-instructions-to-its-processors/ http://lemire.me/blog/2016/10/14/intel-will-add-deep-learnin... [2]https://en.wikipedia.org/wiki/IA-64 https://en.wikipedia.org/wiki/IA-64 [3]https://www.bloomberg.com/news/articles/2015-06-01/intel-buys-altera-for-16-7-billion-as-chip-deals-accelerate https://www.bloomberg.com/news/articles/2015-06-01/intel-buy... [4]https://www.wired.com/2016/09/microsoft-bets-future-chip-reprogram-fly/ https://www.wired.com/2016/09/microsoft-bets-future-chip-rep...
- chriswarbo 10y agoThe claim that static types (in regards to Haskell) make things low level strikes me as wrong. It's probably a difference in the interpretation of the word "type", which is unfortunately used to describe many different things. From a C perspective, the word "type" pretty much means "memory layout", e.g. the difference between "char" and "int" is that the latter (may) use more memory than the former; the difference between "int" and "float" is that the bits are interpreted in a different way; a struct describes how to lay out a chunk of memory; and so on. Static checks can be layered on top of this, but it mostly boils down to 'not misinterpreting the bits'. I've seen this concept distinguished by the phrase "value types". In Haskell land, types have no particular relationship to memory size/layout; e.g. we don't really care what bit pattern will be used to represent something of type "Functor f => (a -> b) -> (forall c. c -> a) -> f b". Unlike C, there's no underlying assumption that "it's all just bits" which we must be careful to interpret consistently; instead, it's all left abstract, grammatical and polymorphic, leaving it up to the compiler to map such concepts to physical hardware however it likes. I certainly think it's a mistake to think in terms of "high level == dynamic types"; there's the obligatory https://existentialtype.wordpress.com/2011/03/19/dynamic-languages-are-static-languages https://existentialtype.wordpress.com/2011/03/19/dynamic-lan... and I'd also consider Homotopy Type Theory to be a very high-level language. It's also a rather constraining simplification too, as it ignores the many other dimensions of a language. Haskell's garbage collection is an obvious abstraction over memory which has nothing much to do with static/dynamic types (linear/affine/uniqueness types are very related, but are yet another overloading of "type" ;) ). As another example, I would count Prolog as more high-level than (say) Python since it abstracts over "low level" details like control flow; again, nothing to do with their types. Likewise, a message-passing language, operating transparently over a network would be more high-level than a language which communicates by (say) opening sockets on particular ports of particular IP addresses and serialising/deserialising data across the link when instructed to by the programmer; we'd be abstracting over the ideas of physical machines, locations and networks. Calling Haskell low level because of its types ignores such other dimensions; in the context of hardware design, machine code, von Neumann architecture, etc. I'd say that abstracting over control-flow with non-strict evaluation is enough to make Haskell high-level.
- chriswarbo 10y agoI admit I'm not too well versed in hardware tech. One thing that comes to mind is using associative memory to implement objects, namespaces, etc. (e.g. http://www.vpri.org/pdf/tr2011003_abmdb.pdf http://www.vpri.org/pdf/tr2011003_abmdb.pdf ); although that general approach seems to be mentioned in the comments.