4 ms·
> English Wikipedia gets about 250M page views per day, which is only about 3,000 per second. If they used an efficient language instead of PHP, that's well wit
by afc 5y ago
> English Wikipedia gets about 250M page views per day, which is only about 3,000 per second. If they used an efficient language instead of PHP, that's well within the capability of a single modern server!
This type of comment betrays a complete lack of knowledge about how large scale internet systems operate (which is something I've worked on for the last 10+ years) and comes across as incredibly naive. I can immediately point out huge flaws in your thinking:
* You're assuming that traffic is constant through the day (that you can divide daily traffic rate to get a sense of daily peak; or, in other words, you provision for mean usage, not peak)
* You're ignoring the problems of unpredictable hotspots (single articles suddenly becoming orders of magnitude more popular than the median in relatively unpredictable and spikey ways; think World Cup final or major earthquake)
* You're not reserving any safety margins for unexpected traffic growth, so you'd likely run into cascading failures
* You're assuming that the entirety of Wikipedia can be held in ram. Including all the history and such. Or, if not, that a single server has enough drives that spindle capacity won't be a problem.
* You're not even considering networking/bandwidth costs. How many network cards would your server need?
* You're ignoring any problems with replication/redundancy (e.g., no backups?) so your site would be a nightmare in reliability terms. Of course, once you do that, you'll need to reason about consistency problems.
I think your logical fallacy is that you're thinking that because your can't understand something (that running Wikipedia could cost so much), it must mean that the thing you can't understand isn't true. Instead, I'd suggest that you'd get further by focusing on discovering the limits of your understanding and genuinely asking yourself: "How could it be that it costs so much? What may I be missing?"
- SpelingBeeChamp 5y agoThanks for the insight.
- jiggawatts 5y agoOh don't get me wrong, I'm not saying I would host a site like Wikipedia on a single server, I'm saying that it's more than possible. Also, don't assume I'm speaking from a position of ignorance here, I know more than a little about how Wikipedia works. I've perused the source, played with the database dumps, etc... I find comments like yours amusing, because I hear similar things in the industry all the time! People like you are not wrong, it's just that correct knowledge is very rapidly outdated in this industry. It's shocking what exponential growth really means! As a random example, I once had an engineer going on and on about their "powerful" email server that cost them six figures. I pointed out that it was slower than the Blackberry phone it was synchronising the email to. (No, really!) You really underestimate the spec of a modern, 2021 era server! You really are. You can buy, right now, for "normal" money, a server with two 64-core AMD EPYC CPUs. That's 128 cores that are the fastest individually of pretty much any server CPU. This platform provides 256 hardware threads. Right there, you're looking at something like 12 requests per second per thread. As long as you can serve them in under 85 milliseconds each on average, that's plenty! I've written a Wikipedia clone for an in-house CMS. The content generation was so fast that for laughs I compiled it to javascript and had it run per keystroke for a live preview in the browser. This worked fine for content up to about 64 KB. This was in 2008, by the way. Things have moved on in terms of performance. Try this with Rust and some judicious use of high-performance AVX-enhanced parsers, and you could generate a typical Wikipedia page in under 10 ms, no sweat. Then there's output caching to get it even lower... That server can have something like 4 TB of memory in it. The entire Wikipedia database is just 5.6 TB uncompressed. Yes, you really can fit most of Wikipedia in memory! Bandwidth? Mellanox makes 200 Gbps Ethernet cards. They're dual port, so that's 400 Gbps. That's enough to handle the bandwidth of a decent sized regional telecommunications company through a single box. I should know, I've seen the network diagram of mine, and they have only 100 Gbps inter-city links! Spindles!? Are you kidding me? There are gaming PCs being built right now with 2 TB NVMe drives in them that can individually put out 7 GB/s and 1M IOPS! Who the hell puts servers on spinning rust in this day and age? That lone, single NVMe drive is putting out 56 Gbps by itself. Throw in a smidge of caching, and it could saturate those aforementioned 2x200 Gbps ports. So no, I would not put the entirety of Wikipedia on just one server, that would be silly. I would put it on two. Okay, maybe three, for redundancy. Just in case.
- kaba0 5y agoCome on.. you are being ridiculous. Also, how is your 500ms ping?
- jiggawatts 5y agoWhich part is ridiculous to you?
- BeefWellington 5y ago> Oh don't get me wrong, I'm not saying I would host a site like Wikipedia on a single server, I'm saying that it's more than possible. This is incorrect. Many of the numbers you throw out ignore that in order to serve its content a server must do some work, which requires both CPU time and memory. Assuming you were able to eliminate hard drives and fit everything into RAM, you still have the problem that to serve requests and do that work you need RAM. >That server can have something like 4 TB of memory in it. The entire Wikipedia database is just 5.6 TB uncompressed. >Yes, you really can fit most of Wikipedia in memory! This is incorrect. According to Wikimedia[2], the entire uncompressed size of Wikipedia is 19TB as of 2019, and the site has grown since. Talking specifically about hard drives: > Spindles!? Are you kidding me? There are gaming PCs being built right now with 2 TB NVMe drives in them that can individually put out 7 GB/s and 1M IOPS! Who the hell puts servers on spinning rust in this day and age? That lone, single NVMe drive is putting out 56 Gbps by itself. Throw in a smidge of caching, and it could saturate those aforementioned 2x200 Gbps ports. Firstly, gaming PCs are almost always ahead of the curve when it comes to performance on everything. Secondly, to your question about who puts servers on spinning rust -- Pretty much everyone. Backblaze even releases HDD reliability ratings[1] based on what they observe in terms of the spinning platter drives. Also, your 2TB NVMe drive is going to cap out at around a quarter of what the large-spindle raid controller can, and you can have a few of those in each system if you want. Sure, a PCIe 4.0 NVMe drive is going to get you 7.88GB/s in your motherboard slot but a PCIe 4.0 raid controller can have four times that bandwidth available to it. But for argument's sake let's say you decided to swap out your raid controller for a PCIe 4.0 card that does NVMe and has onboard raid (these don't exist for purchase yet that I can find at my usual suppliers but I imagine they're coming). Now you have a size problem. The largest current best available SSDs are Sabrent's 8TB PCIe 3.0 drives. This will cut the performance considerably but is probably the sweet spot for size & speed. You need 4 of them just to host current Wikipedia, limiting any growth, and certainly without factoring in any other parts of the system that may take away from the operation. For comparison's sake to the 2TB NVMe drive in your example, my old homelab can hit 5 GB/s on five spindles that I bought in ~ 2011-12 and they're all old and cheap, not enterprise-grade gear. > So no, I would not put the entirety of Wikipedia on just one server, that would be silly. > I would put it on two. > Okay, maybe three, for redundancy. Just in case. Then you have issues of bandwidth by having so few connections, as well as the poor user experience of trying to visit a site that exists and is served in only three places (hopefully around the globe). It's also bad for reliability because if any one site loses its connection suddenly your 200GB/s traffic is pushed onto the other servers. If you want a sense for how reliability engineering works you should look at how Netflix does testing and why they've managed to withstand regional AWS outages despite being hosted on AWS[3]. Wikimedia lays out their system set up pretty well on their own page[4]. Could you technically host wikipedia on your own systems? Sure, lots of people do, go visit Reddit's r/datahoarders and I'm sure you can find more than a few who have their own local copies. Could you actually replace production wikipedia on the Internet with the setup you describe? No, you cannot. I also suspect you haven't costed out the server setup you describe, as it seems incongruous to say that they spend too much while advocating replacing their entire server stack with cutting edge hardware (and will they need to do so annually in order to keep up?). [1]: https://www.backblaze.com/blog/backblaze-hard-drive-stats-for-2020/ https://www.backblaze.com/blog/backblaze-hard-drive-stats-fo... [2]: https://meta.wikimedia.org/wiki/Data_dumps/Dumps_sizes_and_growth https://meta.wikimedia.org/wiki/Data_dumps/Dumps_sizes_and_g... [3]: https://netflixtechblog.com/lessons-netflix-learned-from-the-aws-outage-deefe5fd0c04?gi=d2cab0be29ff https://netflixtechblog.com/lessons-netflix-learned-from-the... [4]: https://meta.wikimedia.org/wiki/Wikimedia_servers https://meta.wikimedia.org/wiki/Wikimedia_servers