5 ms·
I would be very surprised if their spend was just $10M/year on their tech infra. That's very low. I expect it would be much higher. PS: I work for cloud comput
by thoughty 5y ago
I would be very surprised if their spend was just $10M/year on their tech infra. That's very low. I expect it would be much higher.
PS: I work for cloud computing company and have insight on how much costs increase with scale.
edit: fixed typos.
- jiggawatts 5y agoWikipedia is hosted on ordinary rackmount servers at a handful of colocation facilities: https://meta.wikimedia.org/wiki/Wikimedia_servers https://meta.wikimedia.org/wiki/Wikimedia_servers It does not significantly use cloud services, and as you can see by the picture of their racks, they use standard 1 RU pizzaboxes. Their workload is highly scalable horizontally, and doesn't even need much in the way of expensive "enterprise" hardware such a SAN arrays: https://upload.wikimedia.org/wikipedia/commons/thumb/c/ca/Wikimedia_Foundation_Servers_2015-90.jpg/1920px-Wikimedia_Foundation_Servers_2015-90.jpg https://upload.wikimedia.org/wikipedia/commons/thumb/c/ca/Wi... If you buy beige box rackmount hardware with basic compute and CPU, it's pretty easy to spend less than $250K per rack, and the hardware can be kept for 5+ years. That's about 50K per rack per year. Double that for opex, and assume twenty racks distributed across a a handful of sites for capacity and availability. That's just $2M per annum. Keep in mind they're a high-profile charity, so they get discounts and tax breaks. (I'm not saying that's all there is to Wikipedia! There's software development, management, finance, etc... I'm just saying that the infrastructure alone is probably less than you think.) PS: English Wikipedia gets about 250M page views per day, which is only about 3,000 per second. If they used an efficient language instead of PHP, that's well within the capability of a single modern server! There have been comments along these lines made in other forums, and the response was that it would cost more to rewrite the software than the maintain the hardware required for PHP. I don't remember the numbers, but I vaguely remember tens of millions for a rewrite, a high risk of breaking issues, and a significantly lower hardware cost.
- zozbot234 5y ago> It does not significantly use cloud services They do host their own private "cloud", mostly used for research and project development. > English Wikipedia gets about 250M page views per day, which is only about 3,000 per second. If they used an efficient language instead of PHP, that's well within the capability of a single modern server! Nitpicking, page views from outside users are not supposed to hit PHP at all - those requests should be served by their caching layer. But yes, you're correct that they're relying on quite a bit of legacy tech from their early days, not just PHP but MariaDB as well even though many enterprises would now rely on Postgres wrt. these use cases. I suppose that they could rewrite the stuff piecemeal, and recent developments in PHP itself would make it easier. (Some things have been rewritten already, notably the wikitext parser. That's clearly one thing where you would want to avoid breakage at all costs, but they did get it done.)
- vosper 5y ago> But yes, you're correct that they're relying on quite a bit of legacy tech from their early days, not just PHP but MariaDB as well even though many enterprises would now rely on Postgres wrt. Postgres is great and all, but there’s really nothing wrong with relying on MariaDB or MySQL for what Wikipedia is doing.
- deleted 5y ago[deleted]
- bawolff 5y ago> Some things have been rewritten already, notably the wikitext parser The new parser is being rewritten in php in order to be integrated with the rest of the php code.
- yoz-y 5y ago> If they used an efficient language instead of PHP, that's well within the capability of a single modern server! Hm, since Wikipedia dumps are public, has anybody tried? Moving wikipedia itself might be a large endeavor, but seeing some submissions on HN I think a very dedicated dev or two could possibly manage a rewrite to have a proof of concept.
- mschuster91 5y agoMediawiki markup language is not efficiently parse-able and extremely complex with lots of bolted-on features, which is why it took ages for the visual editor to be implemented as a prototype and to iron out the bugs. For any attempt at porting over MW, you will need to re-implement the whole parse/include stuff... and that is a lot of work, even for a proper team.
- bawolff 5y agoSo just steal the parser part
- lozenge 5y agoMy information might be ten years out of date, but the vast majority of page views never hit PHP. If you are viewing a even slightly popular article and aren't logged in, your request is served by a Varnish cache.
- tored 5y agoWhy do you think that MediaWiki is not optimized to handle thousands request per second just because it is written in PHP? MediaWiki repository size is about 1.4 GB and according to phploc it has closer to a million lines of PHP code. Sounds like fools errand to rewrite it.
- jiggawatts 5y agoI looked into the topic in some detail year ago, because I copied Wikipedia's approach for an in-house CMS editor. I looked through the code, did loads of performance experiments, read through related forums posts, etc... The wiki staff themselves admitted that the current parser is inefficient, partly because of PHP, partly because the underlying grammar was not designed to be efficient, and partly because the parser itself was built up over time and wasn't easy to optimise. My approach was to write the parser and a matching "markdown inspired" format at the same time, optimised for speed. Just a handful of small tweaks to the syntax were all that was required to largely eliminate backtracking and achieve a nearly linear parsing time in most cases. If I remember correctly, I had it down to about 1-5ms for a typical 64 KB page, and then HTML generation was another 5-10ms depending on various factors. What a lot of people are missing here is that Wikis are not at all like typical "ERP" applications. The latter sometimes requires dozens of API calls and thousands of database queries to generate just one kilobyte of output HTML. Wiki is very linear, with a single 1-100KB blob of text as input, a matching 1-100KB blob of HTML as output. It all boils down to the parsing and HTML generation efficiency, nothing else matters!
- tored 5y agoYou don't need to parse the grammar for every request and as they admitted it was an old implementation. What I understand they have a new one now. You can parse grammar on save and make an optimized compiled format where static content is already resolved. Next step is to substitute semi-static content (like author name) to a runtime format, this last format can handle dynamic data substitution like current date & time, but probably going to be rare that any substitution is needed for the runtime format (how much truly dynamic data does a wiki have?), most of the time just print it with readfile or similar. Only difficult part here is cache invalidation for the runtime format (I know, it is one of the three difficult things you can do in programming). 3000 request/second in PHP is not hard, especially if you plan for it.
- afc 5y ago> English Wikipedia gets about 250M page views per day, which is only about 3,000 per second. If they used an efficient language instead of PHP, that's well within the capability of a single modern server! This type of comment betrays a complete lack of knowledge about how large scale internet systems operate (which is something I've worked on for the last 10+ years) and comes across as incredibly naive. I can immediately point out huge flaws in your thinking: * You're assuming that traffic is constant through the day (that you can divide daily traffic rate to get a sense of daily peak; or, in other words, you provision for mean usage, not peak) * You're ignoring the problems of unpredictable hotspots (single articles suddenly becoming orders of magnitude more popular than the median in relatively unpredictable and spikey ways; think World Cup final or major earthquake) * You're not reserving any safety margins for unexpected traffic growth, so you'd likely run into cascading failures * You're assuming that the entirety of Wikipedia can be held in ram. Including all the history and such. Or, if not, that a single server has enough drives that spindle capacity won't be a problem. * You're not even considering networking/bandwidth costs. How many network cards would your server need? * You're ignoring any problems with replication/redundancy (e.g., no backups?) so your site would be a nightmare in reliability terms. Of course, once you do that, you'll need to reason about consistency problems. I think your logical fallacy is that you're thinking that because your can't understand something (that running Wikipedia could cost so much), it must mean that the thing you can't understand isn't true. Instead, I'd suggest that you'd get further by focusing on discovering the limits of your understanding and genuinely asking yourself: "How could it be that it costs so much? What may I be missing?"
- SpelingBeeChamp 5y agoThanks for the insight.
- jiggawatts 5y agoOh don't get me wrong, I'm not saying I would host a site like Wikipedia on a single server, I'm saying that it's more than possible. Also, don't assume I'm speaking from a position of ignorance here, I know more than a little about how Wikipedia works. I've perused the source, played with the database dumps, etc... I find comments like yours amusing, because I hear similar things in the industry all the time! People like you are not wrong, it's just that correct knowledge is very rapidly outdated in this industry. It's shocking what exponential growth really means! As a random example, I once had an engineer going on and on about their "powerful" email server that cost them six figures. I pointed out that it was slower than the Blackberry phone it was synchronising the email to. (No, really!) You really underestimate the spec of a modern, 2021 era server! You really are. You can buy, right now, for "normal" money, a server with two 64-core AMD EPYC CPUs. That's 128 cores that are the fastest individually of pretty much any server CPU. This platform provides 256 hardware threads. Right there, you're looking at something like 12 requests per second per thread. As long as you can serve them in under 85 milliseconds each on average, that's plenty! I've written a Wikipedia clone for an in-house CMS. The content generation was so fast that for laughs I compiled it to javascript and had it run per keystroke for a live preview in the browser. This worked fine for content up to about 64 KB. This was in 2008, by the way. Things have moved on in terms of performance. Try this with Rust and some judicious use of high-performance AVX-enhanced parsers, and you could generate a typical Wikipedia page in under 10 ms, no sweat. Then there's output caching to get it even lower... That server can have something like 4 TB of memory in it. The entire Wikipedia database is just 5.6 TB uncompressed. Yes, you really can fit most of Wikipedia in memory! Bandwidth? Mellanox makes 200 Gbps Ethernet cards. They're dual port, so that's 400 Gbps. That's enough to handle the bandwidth of a decent sized regional telecommunications company through a single box. I should know, I've seen the network diagram of mine, and they have only 100 Gbps inter-city links! Spindles!? Are you kidding me? There are gaming PCs being built right now with 2 TB NVMe drives in them that can individually put out 7 GB/s and 1M IOPS! Who the hell puts servers on spinning rust in this day and age? That lone, single NVMe drive is putting out 56 Gbps by itself. Throw in a smidge of caching, and it could saturate those aforementioned 2x200 Gbps ports. So no, I would not put the entirety of Wikipedia on just one server, that would be silly. I would put it on two. Okay, maybe three, for redundancy. Just in case.