9 ms·
The next operating system: building an OS for hundreds, thousands of cores
- ph0rque 16y agoSince starting to learn Erlang, I've been wondering if there's been any effort to build an operating system on top of the Erlang VM to exploit its concurrency... seems like it might be a good fit for this project.
- masklinn 16y agoI'd think the Erlang VM a bit slow as a systems language, but I dream of Erlang as an apps language (à la Objective-C) with a deep integration of the OS and the language runtime (a Lisp Machine in Erlang, and without hardware support, kinda)
- kmod 16y agoThe problem is that parallelism, and performance at this level in general, is not the kind of problem that you can solve by adding additional layers of abstraction. I can only speak to FOS, and not the other projects, but the idea behind it is to optimize the stack all the way down. My guess is that if you tried to run the Erlang VM on a 1000-core cpu, it would run into Linux scalability issues, regardless of how parallel Erlang is. Instead, we need to think about the base primitives of how we interact with the hardware, and potentially redesign them in a way that let us fully utilize the hardware.
- ph0rque 16y agoHmm... is it possible to extend the Erlang turtles further down, as it were, and re-write some of the stuff the VM is based on in Erlang, similar to Ruby/Rubinius?
- gtani 16y agohttp://groups.google.com/group/erlang-programming/msg/75cbeab74644dc0a http://groups.google.com/group/erlang-programming/msg/75cbea... http://www.tilera.com/about_tilera/press-releases/linux-applications-are-scaling-better-ever http://www.tilera.com/about_tilera/press-releases/linux-appl... ( June 2009, Erlang (BEAM) on Tilera 64-core )
- inoop 16y agoAs long as they don't write the thing in C ...
- marshray 16y agoWhat would you write it in?
- stcredzero 16y agoGo. Maybe LuaJIT with some additions to the parser to allow optional Static typing. Smalltalk with an advanced JIT VM and parse/compile time type enforcement ala Strongtalk. Scheme.
- Animus7 16y agoThere's a big difference between an OS and a VM, and they accomplish different things. Go might be feasible, but forcing system-wide GC at random times for the entire system? GC is very hard to make concurrent and a single random-alloc GC'd memory space can't possibly scale to thousands of cores.
- marshray 16y agoI think the problem is deeper than just concurrency and preventing GC pauses. Since a kernel is something that is expected to run forever, it can't afford to leak anything over the long term. For most GCs, collecting that last little bit of garbage (in deterministic time) requires O(committed address space) memory bandwidth. A full-copy style GC may take O(object memory), which could be an improvement. Now that memory and applications are routinely many gigabytes, this is a big deal. It's hard enough for an ecommerce web server to maintain responsiveness, I couldn't imagine trying to respond to hardware IO interrupts in real time while running a collector like that.
- masklinn 16y ago> Go might be feasible, but forcing system-wide GC at random times for the entire system? GC is very hard to make concurrent and a single random-alloc GC'd memory space can't possibly scale to thousands of cores. Erlang (and its way) is a much better fit there, I think it'd be a delightful apps language: the GC runs at the (erlang) process level, each process has its own heap, so even though the GC is a vanilla generational GC by the magic of the Erlang VM it turns into a highly concurrent pauseless GC (only needs to pause a single Erlang process at a time, and you generally have tens of thousands chugging along).
- stcredzero 16y agoan OS for hundreds, thousands of cores That's exactly what Microsoft Azure claims to be in their marketing literature.
- kmod 16y agoThis is an unfortunate case of people adopting technical terms for PR purposes. When we (FOS) say "OS", we mean that when you write a program for FOS, you think about it as if you are writing it for a single computer. It doesn't matter if that computer has multiple processors in it, or even if those processors are in different boxes and are only connected by ethernet.
- stcredzero 16y agoBy reporting what Microsoft is saying in their marketing literature, I'm neither advocating what they say is true, nor am I saying the work referred to in the article is a copy of their efforts. Evidently, you jumped to some conclusion like this. I am merely commenting on the behavior of their marketers. As a long time observer of the tech industry, I always find the behavior of marketers interesting, though not always in pleasant ways. (I think it's wise to pay attention to how such forces change language.)
- runT1ME 16y agoAzul's systems seem to be doing fine with hundreds of cores, and I'm pretty sure it's just modified linux.
- deleted 16y ago[deleted]
- russell 16y agoColor me skeptical. The project will fail, of course, because it is too ambitious. There are too many required new developments for it all to come together: new chip, new OS, new forms of scheduling, a lot more bookkeeping, not to mention new programming paradigms and compiler technology. I would not be surprised that bookkeeping and bandwidth would eat up 90% of the processing cycles. (Obviously, I'm pulling this out of my nether end.) It feels to me that is the next iteration of refinements of technology that is 40 - 50 years old. Caches are what, from 1970? Interconnect issues date from the same time. I think machine cycles are in abundance, the scarce resource is interconnections for data flow. So one of the first things you want to do is organize your data flow so that processing is local. Think simulations like vision processing, weather prediction, rendering, where a processor can work locally and pass on a reduced amount information to its neighbors. The interesting problems arise when the results have to be delivered non-locally. If you store them in main memory for the recipient to pick them up, you run into bandwidth problems. So what I see as needed is gazillions of low level worker bees with modest bandwidth requirements that have semi-permanent connections to the consumers of their output. Think the human brain, Google search, image rendering. Apologies for the rambling, lack of citations, etc. etc, but I am interested in HNer's views on these issues.
- m0th87 16y agoI think you're seeing this project too much from the perspective of whether it will succeed/fail in the marketplace. That is not the point of research. This should be viewed as an exploratory search for new paradigms in multicore computing. Some components might work out to something usable, most won't.
- russell 16y agoI agree with you completely. It is worth doing for what we learn, but what I was trying to say is that the real breakthroughs are going to come from other directions. The people who will get rich are probably the generation after that.
- _delirium 16y agoNot the original poster, but I think it'll fail in the sense of "won't actually invent and tie together all the new components that this press release claims will be invented and tied together", even without including marketplace considerations. However I agree that it'll likely produce interesting research advances and usable components, and some up-front overselling is probably unavoidable... saying something like, "we're initiating a collection of research projects to produce technologies that will be needed for a future generation of operating systems" isn't as good PR as saying, "we're building the next-gen operating system".
- Mongoose 16y agoToo much overview, not enough meat. Are there any whitepapers available for the Angstrom Project? The publication page of their website just says "coming soon." http://projects.csail.mit.edu/angstrom/Publications.html http://projects.csail.mit.edu/angstrom/Publications.html
- kmod 16y agoI think there's some confusion about this project, since the article doesn't go into to much detail about what it actually is, apart from some high-level technical highlights. My understanding (which is potentially outdated) is that Angstrom is a collaboration effort between different groups who are trying to combine their separate research into a larger project. For the criticism that the project will "fail", I think there's some merit. Will this produce the next computation system that the world uses? Probably not. Will this push the boundaries of engineering and improve on the state of the art? Almost certainly. And due to its scale, and the fact that Angstrom was created by combining multiple separate projects, there's also room for individual projects or ideas to succeed even if others don't. My background is that I worked on the FOS project, which is the part that is most related to the title of this post. If you were looking for more technical meat, I encourage you to check out the project page: http://groups.csail.mit.edu/carbon/?page_id=39 http://groups.csail.mit.edu/carbon/?page_id=39