3 ms·
> Stack Overflow serves all the traffic with only 9 on-premise web servers, and it’s on monolith! It has its own servers and does not run on the cloud. This is
by JSavageOne 3y ago
> Stack Overflow serves all the traffic with only 9 on-premise web servers, and it’s on monolith! It has its own servers and does not run on the cloud. This is contrary to all our popular beliefs these days.
Love this part. I wonder how much of the complexity of microservices and various cloud services are actually adding value on a net basis vs. resume-driven development of some bored backend engineer.
- artimaeis 3y agoI also love this anecdote -- but the SO engineering team that made that possible is _not_ your average LOB engineering team in my experience. Those guys were like C#/SQL technomancers. SO is a great example of just how much is possible with a .NET monolith, but very few teams are going to be capable of driving the machines that hard.
- wutwutwat 3y agoAs I understand it, SO sites are mostly read traffic and they keep something like 2/3 of all of their data in cache at any given time. When your entire dataset is sitting in ram, you can handle insane amounts of traffic with few machines using any programming language, because most of the traffic is never even touching any code you wrote. Still impressive though! Edit: I'd be curious how they hold up to a redis `flush all` or memcached `flush_all`. I wonder if they've ever game day'd such a scenario.
- mixmastamyk 3y ago> how they hold up to a `flush all` Populate the cache as part of the startup sequence.
- wutwutwat 3y agoI was talking more about the app servers would be up and running and their cache layer would just go offline and the app servers would get 100% of the traffic. The `flush all` would be to simulate the cache going away. Populating caches on app boot is a bad idea, it'll slow the boot times, and if the cache is down, now your apps can't boot either to at least run in a degraded state. Cache should never be in the critical path like that, it's a cache, and it going away is not unexpected. Just my 2 cents. Some cool info about SO From 2019 > Layers of Cache at Stack Overflow > We have our own “L1”/”L2” caches here at Stack Overflow, but I’ll refrain from referring to them that way to avoid confusion with the CPU caches mentioned above. What we have is several types of cache. Let’s first quickly cover local and memory caches here for terminology before a deep dive into the common bits used by them: > “Global Cache”: In-memory cache (global, per web server, and backed by Redis on miss) > Usually things like a user’s top bar counts, shared across the network > This hits local memory (shared keyspace), and then Redis (shared keyspace, using Redis database 0) > “Site Cache”: In-memory cache (per site, per web server, and backed by Redis on miss) > Usually things like question lists or user lists that are per-site > This hits local memory (per-site keyspace, using prefixing), and then Redis (per-site keyspace, using Redis databases) > “Local Cache”: In-memory cache (per site, per web server, backed by nothing) > Usually things that are cheap to fetch, but huge to stream and the Redis hop isn’t worth it > This hits local memory only (per-site keyspace, using prefixing) and > For the curious, some quick stats from last Tuesday (2019-07-30) This is across all instances on the primary boxes (because we split them up for organization, not performance…one instance could handle everything we do quite easily): > Our Redis physical servers have 256GB of memory, but less than 96GB used. - 1,586,553,473 commands processed per day (3,726,580,897 commands and 86,982 per second peak across all instances – due to replicas) - Average of 2.01% CPU utilization (3.04% peak) for the entire server (< 1% even for the most active instance) - 124,415,398 active keys (422,818,481 including replicas) - Those numbers are across 308,065,226 HTTP hits (64,717,337 of which were question pages) https://nickcraver.com/blog/2019/08/06/stack-overflow-how-we-do-app-caching/ https://nickcraver.com/blog/2019/08/06/stack-overflow-how-we...
- mixmastamyk 3y agoYou have it work the way you chose, therefore not a bad idea.
- harshalizee 3y agoStackoverflow is a much simpler mostly read only, easily cacheable site compared to a workflow heavy SaaS. They're not even remotely comparable.
- aprdm 3y agoI think you're underestimating how much write happens in SO, and overestimating how much writes happens on 99% of the SaaS that have very few paying customers..
- harshalizee 3y agoCurrently work for a very large household name SaaS, previously on Salesforce. I'm willing to bet good money that SO does not remotely come close to the amount of writes in either of these systems.
- quickthrower2 3y agoResume AND sales-driven. It is hard to make a high-growth startup selling co-located servers whose docs page is a link to Linux documentation and RFCs.