10 ms·
Managing two million web servers
- akkartik 11y agoThis article got me to go figure out precisely what Erlang processes are. Tl;dr - they aren't OS processes. So it is still conceivable that an error in Erlang can bring down all your web servers. http://stackoverflow.com/questions/2708033/technically-why-are-processes-in-erlang-more-efficient-than-os-threads http://stackoverflow.com/questions/2708033/technically-why-a...
- ghayes 11y agoWell, not if you distribute your erlang processes across several physical nodes.
- jbardnz 11y agoYou mean just like you can with Apache or Nginx?
- felixgallo 11y agoNo. Not at all like those.
- niij 11y agoCan you please expand on why?
- vertex-four 11y agoThe processes can communicate with processes on other servers in exactly the same way that they can communicate with processes on the same server. If you have one traditional application server and want to do some form of cross-user interaction - let's say chat - you can do that trivially, put it in a queue for that user in a global map. Now when you outgrow that server, you need to rewrite all your code to understand the concept of users being on other servers or use an external message queueing system. In Erlang, all of this is built in by default - if you write code for the OTP framework (the standard library for dealing with messaging, process supervision, etc), all you need to do is connect the two servers together and point them at the same shared user->process mapping process (which you have to build whether you're dealing with one server or 20, as there's no global data otherwise). Of course, if you have absolutely no direct interaction between users, it's trivial to scale anything - fire up a new server and direct some portion of your traffic at it. Erlang's trick is to make it that easy even when you do have direct interaction between users. And of course that works for even backend workloads - if you have a backend server that your frontend servers talk to and need to scale it, if you've coded in Erlang and put as much logic as possible in per-connection processes, you're probably a significant chunk of the way there.
- ww520 11y agoWhat does any of these have to do with the equivalence between Apache/Ngix and Erlang when running multiple servers of them to avoid crashing all?
- vertex-four 11y agoYou can't use traditional servers to scale in lots of cases without a lot of additional development work. You can use Erlang servers to do so. Therefore, they're not the same - Erlang covers a broader range of use cases.
- reinhardt1053 11y agoA constructive comment would be appreciated.
- statictype 11y agoYes - it's exactly like that. The Erlang VM itself has to be rock solid just like you expect Apache and Nginx to be rock solid. Fortunately this more or less seems to be the case.
- phamilton 11y agoI would say a better comparison is the Linux kernel. If the kernel crashes, everything on that box goes.
- ww520 11y agoThe Erlang error in code could be replicated to every single physical node and all of them would crash.
- felixgallo 11y agoThat is not what happens.
- rusanu 11y agoPresumably one would catch such a blatant error early in testing. The problem is the odd case, the code that only crashes in a rare state, never seen in test. When the odd state does occur the effect is mild, as only one actor dies as opposed to the entire OS process.
- ww520 11y agoI was responding to the assertion that distributing code to multiple "processes" would avoid crashing all of them, which is simply not true. Even the rare state you mentioned would crash all the processes. The rare state would be an input that testing failed to catch. The code in any of the processes encountering the rare input could crash.
- tvon 11y agoI don't follow, how does Erlang processes not being OS processes relate to what could happen if an Erlang error occurred?
- signa11 11y ago> ... So it is still conceivable that an error in Erlang... are you talking about erlang-runtime here ? it is not very clear what you are implying. can you please elaborate ? if i take your statement to mean 'erlang runtime' then this is no different than saying, if there is a bug in glibc which gets tickled when i do something, then my process will crash, or if there is a bug in jvm, then the application hosted on the jvm will crash etc. etc.
- akkartik 11y agoFrom OP: "In Erlang we create very lightweight processes, one per connection and within that process spin up a web server. So we might end up with a few million web-servers with one user each." I don't see the difference between spawning a thread for each connection in other languages and spawning a process for each connection in Erlang, so I was curious if the word 'process' implied some extra level of insulation. Reading all the comments here it's still not clear to me what benefit is being obtained by having "2 million web servers" rather than 2 million connections.
- phamilton 11y agoShared memory space is one thing. A process has isolated memory. An error in a spawned thread could mess up shared memory, crashing multiple threads. An error in a process can't mess with the memory space of other processes. Since there's no shared memory, performance characteristics change. Garbage collection is done per process and so "stop the world" doesn't happen, meaning each individual "web server" isn't at the mercy of other requests for performance. It's lots of little things like that.
- rdtsc 11y ago> I don't see the difference between spawning a thread for each connection in other languages and spawning a process for each connection in Erlang, That is a huge difference. Otherwise by now, why would Erlang not just run pthread calls and just spawn a thread. Why bother doing all that work? The reason is because Erlang processes do not share heap memory. That is one of the most important features of Erlang runtime. A crashing process can just crash on its own without leaving all the other 1999999 processes in an undetermined state.
- toast0 11y agoAn error in the beam VM (or a nif) could lead to crash of the OS process, and related termination of all of the erlang processes within. However, the beam VM is pretty robust, because of the design, it turns out to be quite small and relatively simple[1], so you don't tend to end up with things were weird edge cases cause terrible crashes that are hard to find. For nifs you write yourself, it's up to you :) Process isolation isn't complete though, one erlang process can use all of many of the limited resources and end up with the whole vm shutting down. A single process can't use all the CPU, but it can use all the memory, all the atoms, all the filedescriptors, all the ets tables, etc; basically everything in the system limits[2] [1] Calling it simple is a compliment -- a normal person can look at the vm code and understand it, if they need to. [2] http://erlang.org/doc/efficiency_guide/advanced.html#id70200 http://erlang.org/doc/efficiency_guide/advanced.html#id70200
- akkartik 11y agoThanks, yours was finally the comment I was hoping for. Could you point me at the right place to read (or read about) the beam sources? I'm looking at https://github.com/erlang/otp/wiki/Routemap-source-tree https://github.com/erlang/otp/wiki/Routemap-source-tree -- is there something better?
- toast0 11y agoYou're in the right place; unfortunately there's not a lot of good documentation on the internals (although do check out erts/emulator/internal_doc ). It's much easier to dive into the source if you have something specific on your mind (like debugging a performance problem, or if you find an beam crash...). erts/emulator/beam has the most basic parts of the virtual machine -- processes, memory allocation, process switching, message passing, interfaces to the operating system (although a lot of that is in erts/emulator/drivers), etc.
- rdtsc 11y ago> So it is still conceivable that an error in Erlang can bring down all your web servers. No it won't. Not if you mean an error in an Erlang process. It almost sounds impossible, but that is the beauty of the BEAM virtual machine. It is really a marvel of engineering. Now if you mean an error in Erlang VM itself, then yeah, if that crashes it will bring down the all the connections. But even that there is a answer -- distributed Erlang. Erlang VM can talk to processes on other Erlang VMs, even if those are running on another server on other side of the planet.
- jondubois 11y agoYou don't need to "crash the server" in response to an error from a single user - It is sufficient to just close the connection and destroy the session. I doubt that erlang spawns millions of OS processes because that would be extremely inefficient due to CPU context switching. So in reality, all erlang is doing behind the scenes is closing the connection and destroying the session... It's not actually crashing and restarting any processes... You can easily implement this behavior with pretty much any modern server engine such as Node.js and tornado.
- gcr 11y agoHow is that different than just throwing an uncaught exception? Most web toolkits will catch that, log it, and 500 the client for you.
- statictype 11y agoYour web service is stateless - for those cases it's fine. However with this model, you are offloading all your state to another layer - likely the database layer. With Erlang (or rather, the Actor model), you can build state-full pieces of code that are also resilient to crashes that cascade across the entire system, by having a standard pattern of restarting individual actors from a known clean state through a clean supervisor hierarchy.
- mason55 11y ago> I doubt that erlang spawns millions of OS processes because that would be extremely inefficient due to CPU context switching. You are correct, Erlang uses something akin to "green processes" where it manages its own threads but they do not share anything in the way that normal threads would.
- felixgallo 11y agoJoe Armstrong is one of the original authors of Erlang. He's using the erlang nomenclature for servers and processes.
- vonkow 11y agoYes, you can (mostly) implement this behaviour in pretty much any modern server framework or language. The point is, you get this behaviour for free as a core part of the language. I write a lot more javascript than erlang these days and I really miss being able to just say, "let it fail".
- stephen_mcd 11y agoI really love the idea of explaining the actor model as tons of tiny little servers compared to a single monolithic server. I tried to make the same comparison recently when I talked about adding distributed transactions to CurioDB (Redis clone built with Scala/Akka): http://blog.jupo.org/2016/01/28/distributed-transactions-in-actor-systems/ http://blog.jupo.org/2016/01/28/distributed-transactions-in-...
- StreamBright 11y agoWhat was the motivation to rebuild it in Scala?
- stephen_mcd 11y agoIf each key/value in the DB is treated as distributed (as per the actor model), a lot of the limitations Redis faces in a distributed environment are solved. I wrote a lot more about it here: http://blog.jupo.org/2015/07/08/curiodb-a-distributed-persistent-redis-clone/#curiodb http://blog.jupo.org/2015/07/08/curiodb-a-distributed-persis...
- smaili 11y agoCould someone explain how in the context of the article, "process" differs from a "thread", in say Java or Python? Or are they one in the same?
- wheresmypasswd 11y agoGenerally, processes do not share memory, but threads may.
- cheald 11y agoIn Erlang, a "process" is analogous to a Java or Python thread in that it's not spinning up a separate OS-level process, but it really functions as its own self-contained process within the Erlang VM. Processes communicate with messages (like OS-level processes) but can't directly share state (like threads would). The result is that in a single Erlang VM, you can have a butt-ton of processes, each effectively runs a full copy of whatever it needs to run, and it can fail, crash, or otherwise explode without affecting the other processes in the VM. The model is what gives Erlang its remarkable resilience and scalability.
- rjayatilleka 11y agoAn Erlang process isn't really equivalent to a Java thread. A Java thread maps to a native thread, whereas an Erlang process is entirely in userspace. Erlang processes are basically green threads without shared state. Java has no equivalent in the standard library.
- cheald 11y agoIt's not equivalent, but it's roughly analogous from the Java programmer's perspective. You would spawn a new process much like you would start a new thread, but the similarities end there.
- phamilton 11y ago"Whenever you would spawn a new thread in Java, you would spawn a process." While that is true, it represents a subset of cases where you would spawn a process. Processes are meant to model the real world. I'd say the truth is closer to "Any time you would instantiate an object in Java, you spawn a process." That's not always going to be true, but more true than you might expect.
- rasengan 11y agoI can't help but think this is madly in-efficient with cache misses and the like.
- phamilton 11y agoWhy?
- mkhpalm 11y agoI think what he's trying to say is that silver bullets don't exist. Only pros and cons.
- phamilton 11y agoI'm specifically asking about caching.
- wang_li 11y agoPresumably he's talking about processor caches. A shared nothing means that every time you "context" switch to another "process" you're going to have to reload all your cache lines in the L1 D-cache.
- phamilton 11y agoIn Erlang you have a "reduction count budget" of 2000 reductions. This is fairly low, less than 1ms of execution, but during that time you have exclusive use of a CPU. At the end of your budget, you might be preempted, or you might get another window. So you take a bit of a hit to cache, but it's not like you are infinitely context switching. In practice it works fairly well.
- yelnatz 11y agoErlang process != OS process.
- waxjar 11y ago
- frik 11y agoWith the same speak, you could say Facebook mangages billions of PHP web servers, though no one speak like that. (PHP has a shared nothing architecture; HHVM works simlar to the Erlang VM, if one can say so)
- jontro 11y agoExactly, also I do not understand the comparison with apache. Apache can be configured to spawn one process per connection. You also could call a thread a "server" in it's own, so it's just a matter of playing with words. Also if an apache process dies it will not crash other requests
- signa11 11y ago> Also if an apache process dies it will not crash other requests can you have 2m Apache processes on your machine ?
- jontro 11y agoWell that would depend how long you run it for. They don't have to be simultaneous. It wouldnt make much sense to recycle the process for each request however. MaxConnectionsPerChild defines how often the process is recycled
- klibertp 11y ago> They don't have to be simultaneous. Now you're just moving a goalpost. We're talking about Erlang processes - you can have millions of them simultaneously. > It wouldnt make much sense to recycle the process for each request however. It does make a great deal of sense from the fault-tolerance perspective. It also helps security. I have an unpleasant feeling that you don't know much about Erlang and its philosophy. The "one process per connection" architecture has many benefits; it's a canonical way of getting concurrency and parallelism on Unix systems, for example. The problem is the way OSes handle and schedule processes, it simply doesn't work with a very large number of them. Erlang implements its own kind of processes, which enforce the same constraints os OS-level processes without taking megs of RAM each and without being pressed after spawning a few thousand processes.
- rodionos 11y agoThe title is somewhat misleading. I clicked expecting to read how someone is managing 2 mln web server instances such nginx or apache. I was curious what kind of company would claim that.
- yelnatz 11y agoThis is Joe's blog, he co-created Erlang 30 years ago. So he's not really affiliated with any company, but rather a voice for Erlang itself.
- rodionos 11y agoIt's not a clickbait of course, it's just that the term 'web server' and 'manage' made it easy to misinterpret the subject. Typically, by managing a web server we mean administrative tasks involved in configuring and running a multi-threaded http daemon. Perhaps "2 million web server processes" would have been better.
- Vendan 11y agoEven that would be confusing, as "Erlang process" != "process" for anyone that's not an Erlang programmer(or who doesn't parse "process" as the Erlang variant by default)
- giancarlostoro 11y agoAs has been said, Joe is known for creating Erlang which is known for running as many processes as the need arises, giving you access to all the cores in a machine as a result.
- krylon 11y agoUtilizing multiple processor cores was something of an afterthought, IIRC. Initially, the Erlang VM only used a single CPU, and in order to utilize multiple CPUs, one had to start one Erlang VM per CPU/core. One of the nice things about Erlang's general approach is that this was a lot less painful than in many other languages. (I have no idea, though, how much work it was to get the Erlang VM itself to use multiple threads.)
- sandra_saltlake 11y agoprocesses do not share memory, but threads may be,
- deleted 11y ago[deleted]
- mpweiher 11y agoBeautiful way of putting it. Also very close to Alan Kay's vision of "object oriented" "In computer terms, Smalltalk is a recursion on the notion of computer itself. Instead of dividing “computer stuff” into things each less strong than the whole – like data structures, procedures, and functions which are the usual paraphernalia of programming languages – each Smalltalk object is a recursion on the entire possibilities of the computer. Thus its semantics are a bit like having thousands and thousands of computer all hooked together by a very fast network." -- The Early History of Smalltalk [1] I also personally like the following: a web package tracker can be seen as a function that returns the status of a package when given the package id as argument. It can also be seen as follows: every package has its own website. I think the latter is vastly simpler/more powerful/scalable. What's interesting is that both of these views can exist simultaneously, both on the implementation and on the interface side. [1] http://gagne.homedns.org/~tgagne/contrib/EarlyHistoryST.html http://gagne.homedns.org/~tgagne/contrib/EarlyHistoryST.html
- krylon 11y ago> Also very close to Alan Kay's vision of "object oriented" Yes! I remember reading a quote by Alan Key about how pretty much every modern OOP language gets it right entirely wrong, because OOP is supposed to be - in Kay's view - about classes or inheritance, but about sending messages to each other. If one thinks of processes as objects and sending messages as "method calls", it is very object-oriented indeed.
- mpweiher 11y agoYep. Except: please don't think of them as "method calls" :-)) "Smalltalk is not only NOT its syntax or the class library, it is not even about classes. I'm sorry that I long ago coined the term "objects" for this topic because it gets many people to focus on the lesser idea. The big idea is "messaging" -- that is what the kernal of Smalltalk/Squeak is all about (and it's something that was never quite completed in our Xerox PARC phase). The Japanese have a small word -- ma -- for "that which is in between" -- perhaps the nearest English equivalent is "interstitial". The key in making great and growable systems is much more to design how its modules communicate rather than what their internal properties and behaviors should be." [1] My contribution towards this is Objective-Smalltalk[2], where I am working on making connectors user definable. So far, it seems to be working. [1] http://lists.squeakfoundation.org/pipermail/squeak-dev/1998-October/017019.html http://lists.squeakfoundation.org/pipermail/squeak-dev/1998-... [2] http://objective.st/ http://objective.st/
- z3t4 11y agoI would like to see the code for the chat or presence server. I have a hunch it will look different depending on the experience of the programmer. I'm especially interested in how they manage state. Because when you do not have to manage state, everything becomes easy and scalable. With state I mean for example a status message for a particular user.
- lucaspiller 11y agoState is one of the big differences between CGI-esque scripting languages and how things are usually done in Erlang. As processes are cheap, one way is to fire up a session process for each user that connects to the system, and keep this around for the lifetime of their session. The process will hold within it any state that is needed between requests, without needing to refer back to the database every time. Whenever a request comes in, you then route the request to that session process (even when it's on another physical machine). This not only provides a clearer mapping between your external and internal APIs, but also allows you to fairly easily ensure that a user can't perform concurrent actions. Edit: Not sure if that's exactly what you mean, so let me know if you have more questions :)
- yelnatz 11y agoThe Phoenix Framework (web framework for Elixir) already solves this for you.[1] It's called Phoenix.Presence. They used a combination of CRDTs with heartbeats to implement it. You are right, it is a hard problem because of the distributed nature of Erlang/Elixir. That's why Chris provided a framework level solution for it. [1] https://youtu.be/XJ9ckqCMiKk?t=921 https://youtu.be/XJ9ckqCMiKk?t=921 (Erlang Factory SF 2016 Keynote Phoenix and Elm – Making the Web Functional)
- siscia 11y agoA little while ago I wrote an extremely short introduction to distributed, highly scalable, fault tolerant system. It is marketing material for my consulting activity anyway some of you can find it interesting. The PDF is here: https://github.com/siscia/intro-to-distributed-system/blob/master/intro_to_distributed.pdf https://github.com/siscia/intro-to-distributed-system/blob/m... The source code is open, so if you find a better way to describe things feel free to open an issue or a pull request...
- kennydude 11y ago> why does the Phoenix Framework outperform Ruby on Rails? Ruby is known to be a slow language. Most things will easily outperform it
- vemv 11y agoLanguage performance rarely impacts web applications performance. Actual bottlenecks are the database, and your webserver/architectural choices.
- andy_ppp 11y agoI get the feeling that people reading this and saying "it's just kind of like a pool of PHP FastCGI instances or Apache worker pools" etc. Do not understand that Phoenix + Elixir can serve minimal dynamic requests about 20% slower than nginx can serve static files. This is very very fast. It also leads to better code due to being functional, lots of amazing syntactic sugar like the |> operator and the OTP can easily allow you to move processes (these basically have very little overhead) to different machines as you wish to scale. Pattern matching and guards are also incredible. I really do not want to write anything else!
- jerf 11y agoAnd I think Erlang programmers sometimes, if not frequently, forget the Erlang is not magic, it runs on the same CPUs as everybody else with the same access to memory protection, same assembler, etc., and consequently, it can't actually do anything that other many languages can't do too. (And the languages that can't do it only can't do it because they've somehow locked themselves out of it.) Yes, it's a great default to be shared-nothing, and it's great to have a VM that supports this, the differences in affordances are important and I think Erlang was a milestone language. Very serious about that. But when it comes down to it, the practical difference between nginx (as in, a finished piece of software that exists right now, not the hypothetical space of future C programs) and an Erlang webserver is not much. So when the Erlang community tries to rewrite the definition of "server" to be "an Erlang process", it is not an unreasonable response to point out that there are plenty of other web servers that have similar levels of isolation that just happen to be written in other languages, and that we don't run around saying "Oh, my nginx has two hundred thousand web servers in it!" This is bad advocacy, and I'd really suggest that Erlangers stop trying to defend this point. It's not a defensible position. There's no reason to try to redefine "server" to be "the number of isolated processes", because even if you do, Erlang will not have some sort of unique claim to being able to run lots of such processes. Any two "processes" that don't write into each other's space and can't crash each other are "isolated", even if some implementation work had to be done to get it that way. And of all things, webservers are the definitive programs that have had that isolation bashed on and banged on to the nth degree; I wouldn't be surprised that nginx's isolation is more tested than Erlang itself's.
- klibertp 11y agoOMG, guys, this is getting really strange. Half of the commenters here read the word "process" and jumped to their own conclusions, possibly true in general, but obviously wrong in the case of Erlang. It bears repeating: Erlang processes are not OS-level processes. Erlang Virtual Machine, BEAM, runs in a single OS-level process. Erlang processes are closer to green-threads or Tasklets as known in Stackless Python. They are extremely lightweight, implicitly scheduled user-space tasks, which share no memory. Erlang schedules its processes on a pool of OS-level threads for optimal utilization of CPU cores, but this is an implementation detail. What's important is that Erlang processes are providing isolation in terms of memory used and error handling, just like OS-level processes. Conceptually both kinds of processes are very similar, but their implementations are nothing alike.
- deleted 11y ago[deleted]
- collinmanderson 11y agoThanks. That wasn't really clear from the article.
- DougWebb 11y agoIn the late 90s I implemented the same concept for a web application written in Perl. (It's still running today.) There were three tiers to it: Tier 1: a very small master program which ran in one process. It's job was to open the listening socket and maintain a pool of connection handler processes. Tier 2: connection handler processes, forked from the master program. When they started they would load up the web application code, then wait for connections on the listening socket or for messages from the master process. They also monitored their own health and would terminate if they thought something went wrong. (ex: this protected them from memory leaks in the socket handling code.) When an http connection came in on the socket, they would fork off a process to handle the request. Tier 3: request handlers. These processes would handle one http request and then terminate. When they started, they had a pristine copy of the web application code (thanks to Copy-On-Write memory sharing of forked processes) so I knew that there was no old data leaked from previous requests. And since they were designed to terminate after a single request, error handling was no problem; those would terminate too. In cases where a process consumed a lot of memory it would get released to the OS when the process ended. We also had a separate watchdog process that would kill any request handler that consumed too much cpu, memory, or was running much longer than our typical response time. This scaled up to handling hundreds of concurrent requests per (circa 2005 Solaris) server, and around six million requests per day across a web farm of 4 servers. That was back in 2010; I don't know how much the traffic has grown since then but I know the company is still running my web app. This was all very robust; before I left I had gotten the error rate down to a handful of crashed processes per year in code that was more than one release old. BTW, while my custom http server code could handle the entire app on its own, and was used that way in development, for production we normally ran it behind an Apache server that handled static files and would reverse-proxy the page requests to the web app server. So those 6 million requests per day were for the dynamic pages, not all of the static files. That also meant that my web app didn't have to handle caching or keep-alive, which simplified the design and makes the one-request-then-die approach more viable.
- yandrypozo 11y agodoes anybody know how to undo an upvote here in HN ?? This reading was a terrible waste of time :(
- tvon 11y agoYou cannot alter your vote.