16 ms·
Your computer is a distributed system
- rbanffy 5y agoThis was true for several home computers since the late 70's. Atari 8-bit computers had all peripherals connecting via a serial bus, each one with its own little processor, ROM, RAM and IO (the only exception, IIRC, was the cassete drive). Commodores also had a similar design for their disk drives. A couple months back a 1541 drive was demoed running standalone with custom software and generating a valid NTSC signal.
- Frenchgeek 5y ago( https://youtu.be/zprSxCMlECA https://youtu.be/zprSxCMlECA )
- catern 5y agoWow! Reminds me of https://www.rifters.com/crawl/?p=6116 https://www.rifters.com/crawl/?p=6116 A hydrocephalic demo!
- rbanffy 5y agoI think that plan hits a wall for heat dissipation and nutrient/oxygen consumption - not sure we have lungs large enough to keep a brain doing 10x more computation oxygenated, nor perspiration glands to keep it cool. But I'd be totally in to a 10% increase in IQ in exchange to being able to eat 10% more sugar.
- dkersten 5y agoWow, that is cool!
- YZF 5y agowell, it's been true since a wire has been connecting any two bits. The processor, ROM, RAM are all "distributed" systems internally.
- rbanffy 5y agoThat's not what "distributed system" means.
- YZF 5y agoWhat's your definition of "distributed system" then? Two flipflops interconnected on one wafer. Two flipflops inteconnected on one PCB. Two flipflops interconnected with a cable between two PCBs. These are all "distributed". They're all subject e.g. to the CAP theorem. Sure, the probability of one flipflop failing on the same wafer is quite small. The probability of one flipflop failing on one PCB is slightly larger. But fundamentally all these systems are the same. If you have two computers on a network you can make the probability of failure (e.g. of the network) pretty small.
- rbanffy 5y agoI start counting them as independent computers when they have their own firmware.
- TickleSteve 5y agoIt absolutely is. distribution of signals within smaller systems (microcontrollers, ASICs, FPGAs, etc) are all distributed systems. Ask anyone doing any kind of circuit design about distributing clocks and clock skew, etc.
- rbanffy 5y agoIf you read the article, you’ll understand it’s about our computers being networks of smaller computers. The SSD, GPU, NIC, and BMC has its own CPU, memory, and operating system.
- sesuximo 5y agoI think there’s a big difference which is that your computer is allowed to crash when one component breaks whereas a distributed system is typically more fault tolerant.
- uvdn7 5y agoThis is actually what makes handling the distributed system in a single computer easier – everything crashing together makes it an easier problem. E.g. you have multiple CPU cachelines, caching different values of a main memory location. And there are different cache coherence protocols to keep them sane. But cache coherence protocols never need to worry about the failure mode when one cacheline is temporarily unavailable but the others are. So yes, there's a distributed system in each multi-core computer, but it's a distributed system with an easier failure mode. If you like more analogies between CPU caches and distributed systems, https://blog.the-pans.com/cpp-memory-model-as-a-distributed-system/ https://blog.the-pans.com/cpp-memory-model-as-a-distributed-... :p
- harperlee 5y agoIdeally a peripheral crashing should not crash the whole system.
- catern 5y agoAnd indeed it does not: Modern operating systems like Linux can perfectly well deal with all kinds of devices crashing or disappearing at runtime. Just like in larger distributed systems.
- zerohp 5y agoThat's not entirely true. There's usually some level of fault recovery built in but it doesn't extend to the level of allowing any component to fail at any time.
- throwaway787544 5y agoThe thing we are missing still is the distributed OS. Kubernetes only exists because of the missing abstractions in Linux to be able to do computation, discovery, message passing/IO, instrumentation over multiple nodes. If you could do ps -A and see all processes on all nodes, or run a program and have it automatically execute on a random node, or if (grumble grumble) Systemd unit files would schedule a minimum of X processes on N nodes, most of the K8s ecosystem would become redundant. A lot of other components like unified AuthZ for linux already exist, as well as networking (WireGuard anyone?).
- ohYi55 5y ago
- pjmlp 5y agoKubernetes only exists because people wanted to do Application Servers in any language, and now they are rediscovering them trying to sell us on Kubernetes + WebAssembly, the irony.
- zozbot234 5y agoThe abstractions are there in Linux, largely imported from plan 9. And work is ongoing to support further abstractions, such as easy checkpoint/restore of whole containers. Kubernetes is a very new framework intended to support large-scale orchestration and deployment in a mostly automated way, driven by 'declarative' configuration; at some point, these features will be rewritten in a way that's easier to understand and perhaps extend further.
- MisterTea 5y ago> The abstractions are there in Linux, largely imported from plan 9. Which abstractions are those?
- zozbot234 5y ago> to be able to do computation, discovery, message passing/IO, instrumentation over multiple nodes. Kernel namespaces are the building blocks for this, because an app that accesses all kernel-managed resources via separate namespaces is insulated from the specifics of any single node, and can thus be transparently migrated elsewhere. It enables the kind of location independence that OP is arguing for here.
- taeric 5y agoSo much of programming languages is to hide the distributed nature of what the computer is doing on a regular basis. This is somewhat obvious for thread abstractions where you can get two things happening. It is blatant for CUDA style programming. As this link points out, it gets a bit more difficult with some of the larger machines we have to keep the abstractions useful. That said, it does mostly work. Despite being able to find and harp on the areas that it fails, it is amazing how well so many of the abstractions have held up. Would be neat to see explicit handling of what features are basically completely hiding distributed nature of the computer.
- jayd16 5y agoThe abstractions aren't just for simplicity. In many cases, ensuring that the distributed nature is unknown or unobserved means the system can make different decisions without affecting the program. This leaves room for flexibility in the system design.
- zozbot234 5y agoThe distributed nature can never be unobserved, by definition. What a well-designed distributed system can do is offer facilities to enable useful constraints on its operation, that might then be used as necessary via a programming language.
- deleted 5y ago[deleted]
- Karrot_Kream 5y agoMaybe? Alternatively by bringing the distributed nature up front-and-center you can have more flexible designs. If I could timeout my drawing routine when the screen has already refreshed (or context has been stolen from the OS) then I have a lot more flexibility in how to recover instead of pretending to do my best and ending up with a lot of screen tearing when I miss my frame budget.
- 5y ago
- simne 5y agoUnfortunately, this idea fights vs idea of least responsibility. Because, user level programs are all at one level of abstraction, and this distribution is distributed over many levels of abstraction. So in desktop systems, mean mostly successors of business micro machines, access to other levels of abstraction intentionally hardened for measures of security and reliability. The same thing applied to crowd computing - there also vps's are isolated from hardware and from other vps's. These measures usually avoided in game systems and in embedded systems, but they are not allowed to run multiple programs from independent developers (for security and reliability), and their programming magnitudes more expensive than desktops and even server side (yes, you may surprised, but game consoles software in many case more reliable than military, and usually far surpass business software). To solve this contradiction, need some totally new paradigms and technologies, may be some revolutionary, like usage of GAI to write code.
- andrey_utkin 5y agoYour body is a distributed system. Your brain is a distributed system. A live cell is a distributed system. A molecule is a distributed system. In other news, water is wet.
- pkilgore 5y agoWhat is the kernal and the bus for the cloud?
- simne 5y agoThese all now are virtual state machines, which store some state and convert all kernel/bus behavior to interaction with connected via network devices. At the moment there are lot of such devices - exists for sure many full featured, like Raspberry; but also there are network connected ATA drives, network connected sensors, RAM, ROM (Flash); BTW IEEE 1394 FireWire is serial interface, could been used as networking bus; exists adapters ethernet-usb (and many commodity devices work well with such connection), so virtually anywhere could been considered as connected via network bus. Even exists USB 3.0 to PCIe adapter, to use PCIe device throw USB connection. And in reality exists problem, that FireWire so distributed, that it where possible on Macs with FireWire interface, to read memory via this interface. So hardware and software exists, but need some steps to make it's usage safe.
- Koshkin 5y agoYes, and concurrency is, in fact, an implementation detail. Which is why I think that in most applied scenarios it should be hidden, and taken care of, by the compiler.
- amelius 5y agoYes, and your computer is a ball of interconnected microservices too.
- JL-Akrasia 5y agoYou are also a distributed system.
- __turbobrew__ 5y agoSome more distributed than others
- dredmorbius 5y agoThe people are already here, they're just not very evenly distributed.
- hsn915 5y agoYes but your computer will not gracefully handle CPUs randomly failing or RAM randomly failing. Sure, storage devices can come and go, but that's been the case since forever, and most programs are not written to handle this edge case gracefully. Except for the OS kernel. The links between the components of your computer are solid and cannot fail like actual computer network connections. In terms of "CAP" theorom, the system has no Partition tolerance. If one of the the links connecting CPUs/GPUs/RAM breaks, all hell breaks loose. If a single instruction is not processed correctly, all hell might break loose. So I find the analogy misleading.
- aidenn0 5y agoI think that TFA gets it exactly backwards. It's not that we will be able to treat multi-node systems as non-distributed it's that single-nodes will have to start being treated like distributed systems. > The links between the components of your computer are solid and cannot fail like actual computer network connections. I've personally had this disproven to me on multiple occasions.
- catern 5y ago>I've personally had this disproven to me on multiple occasions. That sounds like interesting stories! Can you elaborate?
- aidenn0 5y agoAccidents on desktop hardware: Multiple bad disk cables (more common in IDE era, but happened once with SATA). Interestingly enough, Windows would reduce the drive speed on certain errors, so I had a drive that booted up in UDMA/133 and the longer it was running the slower it got, eventually settling in at PIO mode 2. Switching the drive cable fixed it. A sound card that wasn't screwed in to the case, so if you pushed the phone connector in too hard it would unseat. I still don't know how that happened; it must have been me (unless someone pranked me) but the sound-card hadn't been changed in like 2 years at that point. A DIMM wasn't fully clipped in, but the system worked fine for weeks until someone bumped into the case. Things that were actually intentional: We expect anything plugged in externally (e.g. USB, ethernet, HDMI) to be plugged and unplugged without needing to restart the system. This sounds banal, but wasn't always the case. I had a network card with 3 interfaces (10BASE5 AUI, 10BASE2 BNC, 10BASE-T modular plug) and you needed to power off the system and toggle a DIP switch to change which was in use. I've seen server and minicomputer hardware with hotpluggable CPUs and RAM Eurocard type systems (e.g. VME, cPCI) could connect all sorts of things, and could run without restarting. This sort of blurs the line as to what a "node" is. If you have multiple CPUs on the same PCI bus, is that one node or many? eGPUs have made hotplugging a GPU something that anyone might do today. If you run this setup, then the majority of the computational power in your system can appear and disappear at will, along with multiple GB of RAM.
- ilaksh 5y agoThis proves that conventional wisdom (such as the idea that abstracting distributed computation is unworkable) is often wrong. What happens is enough people try to do something and can't quite get it to work quite right that it eventually becomes assumed that anyone trying that approach is naive. Then people actively avoid trying because they don't want others to think they don't know "best practices". Remember the post from the other day about magnetic amplifiers? Engineers in the US gave up on them. But for the Russians, mag amps never became "unworkable" and uncool to try, and they eventually solved the hard problems and made them extremely useful. Technology is much more about trends and psychology than people realize. In some ways, so is the whole world. It seems to me that at some level, most humans never _really_ progress beyond middle-school level. The starting point for analyzing most things should probably be from the context of teenage primates.
- it 5y agoThe Erlang VM (BEAM) can be viewed as a distributed operating system, or at least the beginnings of one.
- simne 5y agoAgree, and could add, that ALL Erlang flavors (exists at least 4 independent implementations for different environments and for different targets) are distributed. And Erlang is based on relatively new syntax from Prolog, which also have cool ideas.
- WestCoastJustin 5y agoGreat post called "Achieving 11M IOPS & 66 GB/s IO on a Single ThreadRipper Workstation" [1, 2] that basically walks through step-by-step that your computer is just a bunch of interconnected networks. Highly recommend the post if you're into this and also sort of amazing how far single systems have come. You can basically do "big data" type things on this single box. [1] https://tanelpoder.com/posts/11m-iops-with-10-ssds-on-amd-threadripper-pro-workstation/ https://tanelpoder.com/posts/11m-iops-with-10-ssds-on-amd-th... [2] https://news.ycombinator.com/item?id=25956670 https://news.ycombinator.com/item?id=25956670
- deleted 5y ago[deleted]
- bumblebritches5 5y ago
- AdamH12113 5y ago>[The fact that computers are made of many components separated by communication buses] suggests that it may be possible to abstract away the distributed nature of larger-scale systems. This is a neat line of thought, but I don't think it can go very far. There is a huge difference in reliability and predictability between small-scale and large-scale systems. One way to see this is to look at power supplies. Two ICs on the same board can be running off of the same 3.3V supply, and will almost certainly have a single upstream AC connection to the mains. When thinking about communications between the ICs, you don't have to consider power failure because a power failure will take down both ICs. Compare this to a WiFi network where two devices could be on separate parts of the power grid! Other kinds of failures are rare enough to be ignored completely for most applications. An Ethernet cable can be unplugged. A PCB trace can't. I used to work with a low-level digital communication protocol called I²C. It's designed for communication between two chips on the same board. There is no defined timeout for communication. A single malfunctioning slave device can hang the entire bus. According to the official protocol spec, the recommended way of dealing with this is to reset every device on the bus (which may mean resetting the entire board). If a hardware reset is not available, the recommendation is to power-cycle the system! [1] Now I²C is a particularly sloppy protocol, and higher-level versions (SMBus and PMBus) do fix these problems, so this is a bit of an extreme example. But the fact that I²C is still commonly used today shows how reliable a small-scale electronic system can be. Even at the PC level, low-level hardware faults are rare enough that they're often indicated only by weird behavior ("My system hangs when the GPU gets hot"), and the solution is often for the user to guess which component is broken and replace it. [1] Section 3.1.16 of https://www.nxp.com/docs/en/user-guide/UM10204.pdf https://www.nxp.com/docs/en/user-guide/UM10204.pdf
- charcircuit 5y ago>A PCB trace can't. Sure, but physical damage can disconnect it.
- anyfoo 5y agoBut in that case we accept the device as "broken" and need to replace or repair it. If you are lucky the PCB trace was only relevant to a single feature, say an LED. If you are unlucky, it's part of the memory bus and everything is toast. But no significant engineering went into making one more resilient than the other. Whereas in a distributed system, a single broken communication line, even or especially if it's extremely important, still means that the distributed system has to recover somehow, preferably gracefully.
- benreesman 5y agoEric Brewer thinks this is a good point of view on such things: https://codahale.com/you-cant-sacrifice-partition-tolerance/ https://codahale.com/you-cant-sacrifice-partition-tolerance/ L1-blockchain entrepreneurs and people who got locked into MongoDB aside, I think most agree.
- syngrog66 5y agoonce you learn to bias to thinking in terms of message passing between actors, and, bias to having immutable shared state, then,a lot of problems become easier to decompose and solve elegantly, esp at scale
- jmull 5y ago> This is something unique: an abstraction that hides the distributed nature of a system and actually succeeds. That's not even remotely unique. OP is grappling with "the map is not the territory" vs. maps have many valid uses. Abstractions can be both not accurate in every context and 100% useful in many, many common contexts. Also (before you get too excited), abstractions have quality: there are good abstractions -- which are useful in many common contexts -- and bad abstractions -- which overpromise and turn out to be misleading in some or many common contexts. I'll put it this way: the idea that The Truth exists is a rough (and not particularly useful) abstraction. If you have a problem with that, it just means you have something to learn to engage reality more fruitfully.
- tonymet 5y agoi recommend people model their apps this way. spin up more threads than needed, one each for api , DB , LB, async, pipelines etc. you can model an entire stack in one memory space. It's a great way to prototype your complete data model before scaling to the proper solutions. Lots of design constraints are found this way . everything looks great on paper but then falls apart when integrating layers.
- alexisread 5y agoThere are lots of good resources in this area: The programming language of the transputer https://en.m.wikipedia.org/wiki/Occam_(programming_language) https://en.m.wikipedia.org/wiki/Occam_(programming_language) Bluebottle active objects https://www.research-collection.ethz.ch/bitstream/handle/20.500.11850/147091/eth-26082-02.pdf https://www.research-collection.ethz.ch/bitstream/handle/20.... with some discussion of DMA Composita components http://concurrency.ch/Content/publications/Blaeser_Component_Operating_System_PLOS_2007.pdf http://concurrency.ch/Content/publications/Blaeser_Component... Mobile Maude (only a spec) http://maude.sip.ucm.es/mobilemaude/mobile-maude.maude http://maude.sip.ucm.es/mobilemaude/mobile-maude.maude Kali scheme (atop Scheme48 secure capability OS) https://dl.acm.org/doi/pdf/10.1145/213978.213986 https://dl.acm.org/doi/pdf/10.1145/213978.213986 Kali is probably the closest to a distributed OS, supporting secure thread and process migration across local and remote systems (and makes that explicit), distributed profiling and monitoring tools, etc. It is basically an OS based on the actor model. It doesn't scale massively as routing nodes was out of scope (it connects all nodes on a bus), but that can easily be added. Extremely small (running in 2mb ram), it covers all of R5rs, and the VM has been adapted to bare metal. I feel that there is more to do, but a combination of those is probably the right direction.
- saltcured 5y agoI think that someone newly interested in this should consider the longer history of distributed OS concepts too. What are you trying to do differently? What set of tradeoffs give you a well-defined solution space which is a gap in current approaches? https://en.wikipedia.org/wiki/Sprite_(operating_system) https://en.wikipedia.org/wiki/Sprite_(operating_system) was a project at UC Berkeley that ran building-scale networks of workstations as a distributed OS. It ended in 1992 but produced innovations such as the log-structured filesystem along the way. Various commercial products like Domain/OS or SGI Irix also were distributed operating systems at different scales. It seems like these lost out to the more typical HPC solutions with distributed/parallel OS instances limited to each node. Or on the other extreme, mainframe systems continue to be a kind of distributed OS built for high availability and scaling in a very controlled environment.
- kmod 5y agoI worked on this for my masters thesis! The thesis was for a specific part but the group worked on the problem as a whole, see https://dspace.mit.edu/handle/1721.1/49844 https://dspace.mit.edu/handle/1721.1/49844 IMO there are two things that make the current abstraction of a computer as a unit make sense: - You (mostly) don't have to try to handle partial failures within a computer. Partial failures are what make distributed systems hard. - The difference in communication costs between two cores in a single machine is several orders of magnitude lower than communicating with a separate machine using commodity technologies. So while yes, "it's all just distributed" and you can use a common abstraction, a large enough constant factor difference means that you still will have to look through the abstraction to build a performant system.
- zozbot234 5y agoThe latency involved in communicating with a separate machine is comparable to loading data from disk. So yes, it is slow, but not so slow that you couldn't reuse many of the abstractions involved in programming a single machine.