21 ms·
The C10M problem
- jared314 13y agoPrevious Discussion: https://news.ycombinator.com/item?id=5699552 https://news.ycombinator.com/item?id=5699552 (9 months ago) High Scalability Post: http://highscalability.com/blog/2013/5/13/the-secret-to-10-million-concurrent-connections-the-kernel-i.html http://highscalability.com/blog/2013/5/13/the-secret-to-10-m... Original Shmoocon Presentation: http://www.youtube.com/watch?v=73XNtI0w7jA http://www.youtube.com/watch?v=73XNtI0w7jA
- BadassFractal 13y agoThis article on High Scalability also covers part of the problem: http://highscalability.com/blog/2014/2/5/littles-law-scalability-and-fault-tolerance-the-os-is-your-b.html http://highscalability.com/blog/2014/2/5/littles-law-scalabi...
- ksec 13y agoI think OSv or something similar would be part of that solution. Single User / Purpose OS designed to do one / few things and those only. I could only hope OSv development would move faster.
- ehsanu1 13y agoAn implementation of the idea: http://www.openmirage.org/ http://www.openmirage.org/ A good talk about it by one of the developers/researchers: http://vimeo.com/16189862 http://vimeo.com/16189862
- nwmcsween 13y agoSo an exokernel?
- wmf 13y agoPeople don't use actual exokernels; they just use Linux like an exokernel. Aka "1975 programming".
- voltagex_ 13y ago>Content Blocked (content_filter_denied) >Content Category: "Piracy/Copyright Concerns" I'm starting to use these blocks at my workplace as a measure of site quality (this will be a high quality article). Can someone dump the text for me?
- toomuchtodo 13y agoDoes it work through IA? http://web.archive.org/web/20140217043940/http://c10m.robertgraham.com/p/manifesto.html http://web.archive.org/web/20140217043940/http://c10m.robert...
- voltagex_ 13y agoAlso blocked, unfortunately. At least they didn't block all of archive.org
- sb057 13y agoThat page is essentially a glorified intro to his series of blog entries, and they are on another domain, so perhaps they are not blocked: http://blog.erratasec.com/search/label/C10M http://blog.erratasec.com/search/label/C10M
- voltagex_ 13y agoAh, erratasec. I'm surprised that isn't blocked here, too. Thanks for the link.
- JanneVee 13y agoWhy would that page be a "Copyright Concern"?
- rdtsc 13y agoHere is how C2M<x<C3M connections problem was solved in 2011 using Erlang and FreeBSD: http://www.erlang-factory.com/upload/presentations/558/efsf2012-whatsapp-scaling.pdf http://www.erlang-factory.com/upload/presentations/558/efsf2... It shows good practical tricks and pitfalls. It was 3 years ago so I can only assume it got better, but who knows. Here is the thing though, do you need to solve C*M problem on a single machine? Sometimes you do but sometimes you don't. But if you don't and you distribute your system you have to fight against sequential points in your system. So you put a load balancer and spread your requests across 100 servers each 100K connections. Feels like a win, except if all those connections have to live at the same time and then access a common ACID DB back-end. So now you have to think about your storage backend, can that scale? If your existing db can't handle, now you have to think about your data model. And then if you redesign your data model, now you might have to redesign your application's behavior and so on.
- chongli 13y agoIf you could do 10M connections on one machine, then why not 1B on 100? Does it even make sense to have a billion simultaneous connections?
- erichocean 13y agoIf by "connection", you mean TCP, probably not. But that's not the only way to maintain connections, and it's certainly not the only reliable network protocol. My latest project keeps every "connection" open at all times, but it's a custom UDP based reliable messaging protocol, not TCP. At Facebook's scale, we'd have the equivalent of one billion connections "open". It's easy to keep them open, despite changing IP addresses, because every packet is public-key authenticated and encrypted, so you don't have to rely on IP addresses to know who you're talking to... It also means you only pay for a connection setup time once. For mobile devices, the improvement in latency is palpable.
- mh- 13y agoCan I ask some questions about that? Extremely interested.. I've considered doing something similar for our messaging/signalling protocol (currently standard TCP sockets established to several million mobile devices.) I had concerns about what the deliverability of UDP would be on mobile networks; many carriers are going towards NAT'ing everything, (forced) transparency proxies, etc. Are you only using UDP from device->infrastructure? If not, do you rely on the devices providing an ip:port over a heartbeat of sorts (to keep up with IP changes?) Do you have any issues with deliverability, in either direction? (not due to UDP's inherent properties, but because of carrier network behavior) thanks very much for anything you're able to answer.
- ubikation 13y agoI think cheetah OS, the MIT exo kernel project proved this and halvm by Galois does pretty well for network speed that xen provides, but I forget by how much. The netmap freebsd/linux interface is awesome! I'm looking forward to seeing more examples of its use.
- oscargrouch 13y agoi would just love to see that exo kernel from MIT in practice some day in some OS.. i think the research is from the nineties, isnt? Also, netmap from freebds was the first thing that come to my head, as a relief from the IO bottleneck from moderns systems.. As in the original C10k, freebsd to the rescue here.. since it was the first OS with the kqueue interface.. and now is netmap.. the numbers from the speedup in the original paper are astounding
- wpietri 13y agoOn the one hand, I love this. There's an old-school, down-to-the-metal, efficiency-is-everything angle that resonates deeply with me. On the other hand, I worry that just means I'm old. There are a lot of perfectly competent developers out there that have very little idea about the concerns that motivate thinking like this C10M manifesto. I sometimes wonder if my urge toward efficiency something like my grandmother's Depression-era tendency to save string? Is this kind of efficiency effectively obsolete for general-purpose programming? I hope not, but I'm definitely not confident.
- logicchains 13y agoI believe it's what used to be called 'craftsmanship'. Taking pride in creating things that are efficient and not wasteful for no other reason than the desire to make the best product possible.
- PhasmaFelis 13y agoA beautifully carved chair may display excellent craftsmanship, but if you spend three months carving the world's most beautiful chair when the client asked for a dozen basic ladderbacks for a dinner party next Friday, you're not a very wise craftsman. Similarly, if your software runs in 1% of CPU on a typical customer's machine, spending 10x the time and resources to make it run in 0.1% is not laudable. Context is everything.
- logicchains 13y ago>Similarly, if your software runs in 1% of CPU on a typical customer's machine, spending 10x the time and resources to make it run in 0.1% is not laudable. No, but spending 1.2x the time and resources might be. That's not an unreasonable proposition; something written in Go, C# or Scala can run 10 to 100 times faster than the equivalent code in Ruby, but it certainly doesn't take 10 to 100 times longer to write the code. It also conserves resources; you're reducing the amount of the world's non-renewable energy that your app consumes by up to 90%. You're also freeing up your users' machines to do more things simultaneously while running your program, potentially increasing the user's enjoyment of the system in the case of a desktop app.
- erichocean 13y agoWhat's significant to me is that you can do this stuff today on stock Linux. No need to run weird single-purpose kernels, strange hypervisors, etc. You can SSH into your box. You can debug with gdb. Valgrind. Everything is normal...except the performance, which is just insane. Given how easy it is, there isn't really a good excuse anymore to not write data plane applications the "right" way, instead of jamming everything through the kernel like we've been doing. Especially with Intel's latest E5 processors, the performance is just phenomenal. If you want a fun, accessible project to play around with these concepts, Snabb Switch[0] makes it easy to write these kinds of apps with LuaJIT, which also has a super easy way to bind to C libraries. It's fast too: 40 million packets a second using a scripting language(!). I wrote a little bit about a recent project I completed that used these principles here: https://news.ycombinator.com/item?id=7231407 https://news.ycombinator.com/item?id=7231407 [0] https://github.com/SnabbCo/snabbswitch https://github.com/SnabbCo/snabbswitch
- stefan_kendall 13y agoNot really. Last I checked, stock Ubuntu was choking around 60k concurrent connections for no reason, and Fedora could handle a lot more. This was a couple years back, but I'd demand numbers before assuming the situation has changed.
- erichocean 13y agoThe entire purpose is to NOT route stuff through the kernel. TFA explains how to do it, and I've done it myself. You can set a flag on the Linux kernel when it boots limiting it to the first N cores. I usually use 2. The remaining cores are completely idle—Linux will not schedule any threads on those cores. Then you build an app that works more-or-less like Snabb Switch, which talks directly to the Ethernet adaptor, bi-passing the kernel (Ubuntu, Fedora, etc. isn't relevant in the least). So, you launch your app as a normal userland app. For each of your app's threads, schedule them on the remaining CPU cores however you want (I schedule one thread per core). Linux will not schedule its own threads or threads from any other process on those cores, so you own them completely—it'll never context switch to another thread. That means when you SSH in, it's running on a thread on cores 1 or 2 only. Same with every other Linux process but your own. Other than sucking up available memory bandwidth and potentially trashing your L2/L3 cache, these other processes don't impact your own app at all. Thus, even though you're running stock Linux, and SSH and gdb works, and you've got a normal userland app, your app is the ONLY app running on the remaining cores, and you're talking directly to the hardware. It's just as fast as doing everything without a kernel, except it cost you 2 cores. IMO, it's more than worth it for the convenience. This approach is so easy that there's really no reason not to do it. There are so many situations in the past where I wanted the performance of those single-app kernels, but it just wasn't worth the dev effort. That's no longer true.
- EdwardDiego 13y agoAt the risk of sounding dumb, aren't we still limited to 65,534 ports on an interface?
- dxhdr 13y agoPort numbers must only be unique for ip:port pairs. A TCP connection is identified by the "quadruple" source_ip:source_port, dest_ip:dest_port. You can have as many connections as you want on the same source_ip on port 80 as long as there aren't 65,535 to the same dest_ip (ie as long as the quadruple is unique).
- EdwardDiego 13y agoCheers for the info! I guess that's where I got the wrong idea from - attempting to stress test one machine from another machine, I'd always hit that limit, but now I understand why.
- turbojerry 13y agoTry Tsung for stress testing, it can use multiple IPs and therefore you can open up as many connections as you like - http://tsung.erlang-projects.org/ http://tsung.erlang-projects.org/
- perlgeek 13y agoAlso with IPv6 you can easily route a whole /64 net (264 IPs!) onto a single machine.
- sp332 13y ago2^64 IPs. :)
- deleted 13y ago[deleted]
- leoh 13y agoProjects such as the Erlang VM running right on top of xen seem like promising initiatives to get the kind of performance mentioned (http://erlangonxen.org/ http://erlangonxen.org/).
- rdtsc 13y agoI wish they open sourced it and let others look at the code and experiment with it. For a lot of developers if it isn't open = it doesn't exist. Now it is their code and they do whatever they want but that is my view of the project.
- zerop 13y agoOne more problem is cloud. We host on cloud. cloud service providers might be using old hardware. Newest hardware or specific OS might be winner but no options on cloud. How do you tackle that ?
- jon-wood 13y agoIf this sort of thing matters that much to you then you bite the bullet and rent a rack to fill with physical servers somewhere.
- ganessh 13y ago"There is no way for the primary service (such as a web server) to get priority on the system, leaving everything else (like the SSH console) as a secondary priority" - Can't we use the nice command (nice +n command) when these process are started to change its priority? I am sorry if it is so naive question
- slashnull 13y agoHe probably meant that from a TCP point of view, as in there is no way to give a higher priority to incoming TCP connections going into the server than to those going into SSH, even if you could use nice to assign more cpu time to your server. Or perhaps he meant that the infrastructure used to do multitasking still have to interrupt both his server and SSH, but then described how the kernel can be set to leave some cores free of work then to set the server to use only those and then run absolutely uninterrupted. Not the only bizarre and confusing statement he wrote, anyways.
- axman6 13y agoIt seems we've already passed this problem: "We also show that with Mio, McNettle (an SDN controller written in Haskell) can scale effectively to 40+ cores, reach a throughput of over 20 million new requests per second on a single machine, and hence become the fastest of all existing SDN controllers."[1] (reddit discussion at [2]) This new IO manager was added to GHC 7.8 which is due for final release very soon (currently in RC stage). That said, I'm not sure if it can be said if all (or even most) of the criteria have been met. But hey, at least they're already doing 20M connections per second. [1] http://haskell.cs.yale.edu/wp-content/uploads/2013/08/hask035-voellmy.pdf http://haskell.cs.yale.edu/wp-content/uploads/2013/08/hask03... [2] http://www.reddit.com/r/haskell/comments/1k6fsl/mio_a_highperformance_multicore_io_manager_for/ http://www.reddit.com/r/haskell/comments/1k6fsl/mio_a_highpe...
- porlw 13y agoIsn't this more-or-less how mainframes work?
- memracom 13y agoJust what are these resources that we are using more efficiently? CPU? RAM? Are they that important? Should we not be trying to use electricity more efficiently since that is a real world consumable resource. How many connections can you handle per kilowatt hour?
- logicchains 13y agoGenerally electricity use is roughly proportional to CPU and RAM usage, as they're powered by electricity. If you have two otherwise identical programs, one of which uses 50% of the cpu and the other of which use 10%, chances are the latter will use less electricity.
- joosters 13y agoIf you are going to write a big article on a 'problem', then it would be a good idea to spend some time explaining the problem, perhaps with some scenarios (real world or otherwise) to solve. Instead, this article just leaps ahead with a blind-faith 'we must do this!' attitude. That's great if you are just toying with this sort of thing for fun, but perhaps worthless if you are advocating a style of server design for others. Also, the decade-ago 10k problem could draw some interesting parallels. First of all, are machines today 1000 times faster? If they are, then even if you hit the 10M magic number, you will still only be able to do the same amount of work per-connection that you could have done 10 years ago. I am guessing that many internet services are much more complicated than a decade ago... And if you can achieve 10M connections per server, you really should be asking yourself whether you actually want to. Why not split it down to 1M each over 10 servers? No need for insane high-end machines, and the failover when a single machine dies is much less painful. You'll likely get a much improved latency per-connection as well.
- slashnull 13y agoThe two bottom-most articles (protocol parsing and commodity x86) are seriously pure dump, but fortunately the ones about multi-core scaling are pretty damn interesting.
- dschiptsov 13y agoSo, he is trying to suggest that pthread-mutex based approach won't scale (what a news!) and, consequently JVM is crap after all?) The next step would be to admit that the very idea to "parallelize" sequential code which imperatively processes sequential data by merely wrapping it into threads is, a nonsense too?) Where this world is heading to?
- swah 13y agoThose two articles, http://blog.erratasec.com/2013/02/multi-core-scaling-its-not-multi.html http://blog.erratasec.com/2013/02/multi-core-scaling-its-not... (from Robert Graham) and http://paultyma.blogspot.com.br/2008/03/writing-java-multithreaded-servers.html http://paultyma.blogspot.com.br/2008/03/writing-java-multith..., seem to say opposing things about how threads should be used. Having no experience with writing Java servers, I wonder if any you guys have an opinion on this.
- cjbprime 13y ago> There is no way for the primary service (such as a web server) to get priority on the system, leaving everything else (like the SSH console) as a secondary priority. Just for the record -- the SSH console is the primary priority. If the web server always beats the SSH console and the web server is currently chewing 100% CPU due to a coding bug..
- eranation 13y agoWhat about academic operating system research that was done years ago? Exokernel, SPIN, all aim to solve the "os is the problem" issue. Why don't we see more in that direction?
- alberth 13y agoWhatsApp is achieving ~3M concurrent connections on a single node. [1][2] The architecture is FreeBSD and Erlang. It does make me wonder, and I've asked this question before [3], why can WhatsApp handle so much load per node when Twitter struggled for so many years (e.g. Fail Whale)? [1] http://blog.whatsapp.com/index.php/2012/01/1-million-is-so-2011/ http://blog.whatsapp.com/index.php/2012/01/1-million-is-so-2... [2, slide 16] http://www.erlang-factory.com/upload/presentations/558/efsf2012-whatsapp-scaling.pdf http://www.erlang-factory.com/upload/presentations/558/efsf2... [3] https://news.ycombinator.com/item?id=7171613 https://news.ycombinator.com/item?id=7171613
- SkyMarshal 13y agoHaven't you answered your own question? FreeBSD + Erlang vs Rails. Not hating on Rails but it wasn't remotely designed for this use case, Erlang was and there's arguably nothing better in the world at it. Twitter got better after they moved to the JVM, another battle-hardened platform designed for scale.
- barrkel 13y agoThe problem of 1:1 messaging is slightly different to Twitter, which is more m:n. 1:1 messaging can be handled reasonably easily with a mailbox per user, and there is no shared state. Messaging with m:n has different optimal patterns depending on the relative ratios of m and n. Twitter has many users with millions of followers; if Twitter used a 1:1 mailbox approach like a chat app, these users would be whole countries worth of load on their own. That's not to say that Twitter's scaling issues where wholly forgivable. They weren't fatal to the service, but I don't think they were necessary with good design from the start. High popularity is a good problem to have though.
- Aloisius 13y agoWhat's the current state of internet switches? Back when I used to run the Napster backend, one of our biggest problems was that switches, regardless of whether or not they claimed "line-speed" networking, would blow up once you pumped too many pps at them. We went through every single piece of equipment Cisco sold (all the way to having two fully loaded 12K BFRs) and still had issues. Mind you, this was partially because of the specifics of our system - a couple million logged in users with tens of thousands of users logging in every second pushing large file lists, a widely used chat system which meant lots of tiny packets, a very large number of searches (small packets coming in, small to large going out) and a huge number of users that were on dialup fragmenting packets to heck (tiny MTUs!). I imagine a lot of the kind of systems you'd want 10M simultaneous connections for would hit similar situations (games and chat for instance) though I'm not sure I'd want to (I can't imagine power knocking out the machine or an upgrade and having all 10 million users auto-reconnect at once).
- wmf 13y ago10 Gbps switches are pretty good and are generally line rate (as long as you avoid ten-year-old chassis).