10 ms·
How we made editing Wikipedia twice as fast
- programminggeek 12y agoI feel like 3 seconds to load a page is still slow. I assume that is time to build a page that doesn't hit cache, and that Wikipedia is using something like varnish to cache pages most of the time. Still, 3 seconds to load a page feels like a slow page and should be a lot faster.
- sp332 12y agoWikipedia only caches the latest version of each page. Older versions have to be rebuilt on the fly every time.
- sarciszewski 12y agoHmm, this sounds remarkably like a possible DoS vector.
- Zikes 12y agoI would be surprised if the servers/instances responsible for rebuilding old pages had any other responsibilities. If that's the case, DoS-ing by flooding old page versions would likely only bring down that particular site feature.
- matthewmacleod 12y ago3 seconds is page saving time for editors. Since wikipedia is a read-heavy site, I would expect that to take longer than loading a page, as writes will always be more expensive and less optimised. By comparison, uncached page load time appears to be ~800ms
- nightpool 12y agoDid you even read the article, or just skip to the comments when you saw the graph? He spends a whole section talking about Wikipedia's aggressive caching strategy, and how 96-98% of the pages are served from Squid cache servers. Also you didn't even scroll down to the next graph, where he shows the average page load time (for logged in, not anonymous, users) as 800 ms. (which is obviously larger then the average page load time served from cache). The first graph was the average amount of time it takes to save a wikipedia edit.
- sparkzilla 12y ago>The CPU load on our app servers has dropped drastically, from about 50% to 10%. Our TechOps team member Giuseppe Lavagetto reports that we have already been able to slash our planned purchases for new MediaWiki application servers substantially Which means that the hosting costs for the site will drop even lower. So why is Wikipedia still begging for money? http://newslines.org/blog/stop-giving-wikipedia-money/ http://newslines.org/blog/stop-giving-wikipedia-money/
- ahelwer 12y agoFrom your article: >What is the point of a site saying they don’t want to show ads, then covering up 50% of the screen with a request for money? This is an article you seriously recommend reading?
- sparkzilla 12y agoIs there a factual problem? See the enclosed screenshot where the begging banner covered 50% of my 23" screen.
- sarciszewski 12y agoThere is a substantial difference between a "Please give us money" served from the same servers that you are accessing, and a third-party advertisement loaded with tracking beacons, malicious cookies (and often evercookies) from the evil advertising industry. The former is annoying. The latter is annoying and invasive.
- Karunamon 12y agoSeeing the content of the page shifted down by 50% so you can beg for money leads to me not donating, but tweaking my filter rules. ABP users can kill that banner by toggling on the "Fanboy's Annoyances" list.
- sp332 12y agoYes, I've seen the banners myself. Asking users directly for money means you're not worried about what advertisers, or other investors, will think about the content on your site. https://pbs.twimg.com/media/B5L1CR0CEAEN78C.jpg https://pbs.twimg.com/media/B5L1CR0CEAEN78C.jpg <- Sorry for the text-in-a-picture but the original thread is dead.
- leeoniya 12y agoworth mentioning: phpng (PHP7) has cut cpu time in half over the past year [1] (scroll down), don't know what the mem situation is. i don't know if HHVM has additional advantages over plain PHP, but certainly the list of major benefits will be smaller by the next major version. [1] https://wiki.php.net/phpng https://wiki.php.net/phpng
- ck2 12y agoThere won't be a stable php7 release until late 2015
- ecaron 12y agonode.js doesn't have v1.0 release yet, but I still use it... I think we're reaching a more acceptable point in technologies where our proper test coverage ensures that using an unblessed release in production is acceptable.
- jaredmcateer 12y agoThat analogy is a bit flawed, both NodeJs and PHP advertise their stable production ready versions on the homepage. Node doesn't advertise v0.11+ anywhere that could be misconstrued as production ready, neither does PHP advertise NG as such. While you may be ready to take such risks, I can assure you most companies are not.
- jaredmcateer 12y agoThe main advantage is that HHVM is available now, PHPNG won't be stable until later this year at best. For my code base HHVM isn't compatible so we're eagerly awaiting the stable release of NG but damned if we didn't try to upgrade to HHVM.
- psykovsky 12y agoThe advantage of PHPNG is that it has a dedicated development team who will not go away when Facebook goes the way of Myspace. And, yes, I know it is open source and development can continue. My point still stands.
- ExpiredLink 12y agoBTW, what's the relationship between Facebook and Zend? Is PHP now 'owned' by Facebook?
- sarciszewski 12y agohttp://www.infoworld.com/article/2628093/open-source-software/facebook-s-php-doesn-t-really-compete-with-zend.html http://www.infoworld.com/article/2628093/open-source-softwar...
- smsm42 12y agoNo, PHP is not owned by Facebook, and wasn't owned by Zend either. Zend contributed a lot into PHP development (phpng effort is largely sponsored by Zend) but it doesn't mean it "owns" PHP. Nobody really owns PHP. Facebook owns HHVM and Hack (which I guess was one of the reasons they created it - with organization this large, it makes sense to have a platform that is predictable for them - i.e. owned by them). Some people from Facebook also contribute to PHP, including those working for HHVM/Hack team.
- TazeTSchnitzel 12y agoAlso note that the Zend Engine, confusingly, isn't made by Zend, it's just a component of PHP. Zend Technologies, Inc. came after the Zend Engine, but both were created by the same duo (Zeev Suraski and Andi Gutmans).
- Cherian 12y agoAre the Wikipedia server and network performance graphs public?
- bhauer 12y agoGreat write up. This will be a good point of reference for future debates concerning the value of selecting high-performance platforms for web-applications. A common refrain among advocates of slower platforms is that computational performance does not matter because applications are invariably busy waiting on external systems such as databases. While that may be true in some cases, database query performance is often only a small piece of an overall performance puzzle. Blaming external systems is a too-convenient umbrella to avoid profiling an application to discover where it's actually spending time. What to do once you've made external systems fast and your application is still squandering a hundred milliseconds in the database driver or ORM, fifty milliseconds in a low-performance request router, and another 250 milliseconds in a slow templater or JSON serializer? Yes, three seconds for a page render is still uncomfortably slow, but it's substantially faster than the original implementation, and unsurprisingly frees up CPU for additional concurrent requests. It's a shame Wikimedia didn't have this platform available to them earlier. Today web developers have many high-performance platform options that offer moderate to good developer efficiency. Those who use low-performance platforms may do their future selves a service by evaluating (comfortable) alternatives when embarking on new projects.
- tracker1 12y agoI've been very happy with node as a platform for 3+ years now... Though it tends to take the mindset of designing to scale instead of as a monolith... breaking workers up from services, and spreading the data out. I'd be more interested to see what, if any changes they have made to their other server components. Are they using NGINX vs Apache? What kinds of caching systems are they using? What is their database layout? Is their data sharded? These kinds of things I think are much more interesting, and at their scale at least as impactful.
- jonesnc 12y agoI would also like to hear more about their infrastructure. I do know they use Varnish as their cache server. https://wikitech.wikimedia.org/wiki/Varnish https://wikitech.wikimedia.org/wiki/Varnish
- 12y ago
- drzaiusapelord 12y agoCurious why they're using squid and not varnish for caching. Weird how they're progressive with PHP but still sticking with the antiquated squid.
- neilk 12y agoI think Gabriel was referring to a historical moment -- Wikimedia does use Varnish now. "We currently use Varnish for serving bits.wikimedia.org, upload.wikimedia.org, text of pages retrieved from WMF projects, and various miscellaneous web services. Nothing uses Squid." -- https://wikitech.wikimedia.org/wiki/Varnish https://wikitech.wikimedia.org/wiki/Varnish
- deleted 12y ago[deleted]
- renaudg 12y agoThey said they were using Squid back in 2004, not necessarily now still. And "progressive" with PHP but not Squid ? They're both technologies from the same era, Facebook just made PHP more usable.
- RobAley 12y agoModern PHP isn't the PHP of 2004, HHVM or not.
- deleted 12y ago[deleted]
- Shish2k 12y ago> Between 2-4% of requests can’t be served via our caches, and there are users who always need to be served by our main (uncached) application servers. This includes anyone who logs into an account, as they see a customized version I run a similar site (95% read-only), and have been pondering whether it would make sense to use something like Varnish's Edge Side Includes (Like SSI, combining cached static page parts and generated dynamic page parts) -- I wonder if they've considered that and what the results would be like?
- dugmartin 12y agoI've played with this configuration but never put it into production: Varnish (with ESI enabled) -> Nginx (with memcached module enabled) -> PHP-FPM. Most of the page is served with Varnish with the ESI directives hitting the Nginx server which serves the fragments from memcached if present, otherwise the PHP-FPM server is hit which then returns the results and sets the fragment in memcached with a low ttl.
- aonic 12y agoI've run into a similar case. In my case using ESI defeated the whole purpose of the Varnish implementation because the dynamically generated portion of the page had to rebuild the user session and bootstrap the framework, which, while faster than doing it for the entire page, was still not as fast as I would have liked to deliver the content to the browser. In the end I ended up placing special HTML tags where the ESI tags would have been placed, and then making AJAX calls (on dom ready) to swap out the static content with the dynamic user-specific content. It worked well for our needs.
- MatmaRex 12y agoThere is an old proposal to use it, but it seems that not much effort was put into it. https://phabricator.wikimedia.org/T34618 https://phabricator.wikimedia.org/T34618 https://www.mediawiki.org/wiki/Requests_for_comment/Partial_page_caching https://www.mediawiki.org/wiki/Requests_for_comment/Partial_... Personally I doubt that it would help much – there are hardly any non-dynamic parts of the page in MediaWiki. (Everything is customizable using server-side configuration, modifying magical pages, or both; or dependent on user permissions or preferences; or both.)
- jedberg 12y ago> and there are users who always need to be served by our main (uncached) application servers. This includes anyone who logs into an account, as they see a customized version of Wikipedia pages that can’t be served as a static cached copy I keep hearing this, but it isn't true anymore. For something like wikipedia, even when I'm logged in, 95% of the content is the same for everyone (the article body). You can still cache that on an edge server, and then use javascript to fill in the customizations afterwards. This will get you two wins: 1) The thing the person is most likely interest in will load quickly (the article) and 2) your servers will have a drastically reduced load because most of the content still comes from the cache. The tradeoff of course is complexity. Testing a split cache setup is definitely harder and more time consuming as is developing towards it. But given the page views of Wikipedia, would be totally worth it.
- qopp 12y agoYou don't necessarily need a split cache. Just serve some extra javascript to all users that only performs the extra work if there is presence of a cookie. (Also make sure the javascript sets a cookie so that wikipedia can fall back to non-cached if javascript isn't enabled.)
- gaius 12y agoYeah LiveJournal was doing that with memcached like 10 years ago.
- twic 12y agoYou don't even need to use JavaScript - you can do it in the cache with edge-side includes: https://www.varnish-cache.org/docs/3.0/tutorial/esi.html https://www.varnish-cache.org/docs/3.0/tutorial/esi.html
- natrius 12y agoAfter using ESIs with Varnish for a project, I'd never do it again. Cached static pages with Javascript that pulls in the dynamic parts is an easier solution to maintain and has better failure modes.
- tokenadult 12y agoSeveral of the previous comments here have quoted a key, interesting fact from the submitted article: "Between 2-4% of requests can’t be served via our caches, and there are users who always need to be served by our main (uncached) application servers. This includes anyone who logs into an account, as they see a customized version." That's my experience when I view Wikipedia. I am a Wikipedian who has been editing fairly actively this year, and I almost always view Wikipedia as a logged-in Wikipedian. I see the Wikimedia Foundation tracks the relevant statistics very closely and has devoted a lot of thought to improving the experience of people editing Wikipedia pages. I can't say that I've noticed any particular improvement in speediness from where I edit, and I have definitely seen some EXTREMELY long lags in edits being committed just in the past month, but maybe things would have been much worse if the technical changes this year described in this interesting article had not been made. From where I sit at my keyboard, I still think the most important things to do to change the user experience for Wikipedia editors is to change the editing culture a lot more to emphasize collaboration in using reliable sources over edit-warring around fine points of Wikipedia tradition from the first decade of Wikipedia. But maybe I feel that way because I have worked as an editor in governmental, commercial, and academic editorial offices, so I've seen how grown-ups do editing. I think the Wikimedia Foundation is working on the issue of editing culture on Wikipedia too, but fixing that will be harder than fixing the technological problems of editing a huge wiki at scale. Human behavior is usually a tougher problem to solve than the scalability of software. By the way, the article illustrates the role for-profit business corporations like Facebook have in raising technical standards for everybody through direct assistance to nonprofit organizations running large websites like the Wikimedia Foundation. That's a win-win for all of us users.
- kanamekun 12y ago<< From where I sit at my keyboard, I still think the most important things to do to change the user experience for Wikipedia editors is to change the editing culture a lot more to emphasize collaboration in using reliable sources over edit-warring around fine points of Wikipedia tradition from the first decade of Wikipedia. >> Completely agree with you. Also agree that fixing culture is often harder than fixing technological scaling problems! All that said, kudos to Wikimedia Foundation for addressing the speed issues for uncached pages. Great work!
- nosage 12y agohttps://ganglia.wikimedia.org/latest/ https://ganglia.wikimedia.org/latest/ ooo pretty!
- yuvipanda 12y agoAnd http://gdash.wikimedia.org/ http://gdash.wikimedia.org/ :)
- vladmk 12y agoWard Cunningham, now there's a guy who should have played on an nfl team.
- PanMan 12y agoHonest questions: What kind of benchmark do others uses for a 'reasonable' response time? Of course it fully depends on the use-case (rendering a video can be hard in 500ms), but for user facing stuff? In my previous startup we tried to stay within 500ms. Not saying this isn't a great improvement, but to me 3s still sounds quite long? (not saying it's easy to do quicker!)
- sparkman55 12y agoThere's a difference between requests that the user "expects" to take a long time, and those that can never be fast enough. For example, POSTs, credit card transactions, and things like Wikipedia edits generally have lengthy forms prior to the request, and the user can tolerate a correspondingly-lengthy response time. I prefer 2s as a target for anything like that, and rely on a queue for asynchronously processing anything that takes longer. For GET requests, particularly those reached by clicking a link from elsewhere on the site, faster is better... Luckily, many of these types of requests can leverage a cache.
- fitshipit 12y agoIt depends completely on the application. It's also usually best to focus on the tail end -- I've always found alerts on 95th or 99th percentile latency useful and easy to decide on thresholds for (ask yourself -- how slow can I tolerate this being for 1% or 5% of users?)
- laurencerowe 12y agoIt's great to see such a significant improvement, but it goes to show just how limiting the CGI era architecture really is. A modern persistent web apps running in Python/Java/Ruby/etc is able to perform preparatory work at startup in order to optimize for runtime efficiency. A CGI or PHP app has to recreate the world at the beginning of every request. (Solutions exist to cache byte code compilation for PHP, but the model is still essentially that of CGI.) Once your framework becomes moderately complex the slowdown is painful.
- eru 12y agoIf you write it right, you should even be able to cache the state of the whole execution right until your code sees the first byte that depends on the user request or the environment. Sort of like a copy-on-write, but it's copy-on-read.
- putlake 12y agoHere's a video of the author's presentation at Scale Conf on migrating Wikipedia to HHVM: http://www.dev-metal.com/migrating-wikipedia-hhvm-scale-conference-2014/ http://www.dev-metal.com/migrating-wikipedia-hhvm-scale-conf...
- dynjo 12y agoStill slower than https://slimwiki.com https://slimwiki.com
- B-Con 12y ago> Between 2-4% of requests can’t be served via our caches, and there are users who always need to be served by our main (uncached) application servers. This includes anyone who logs into an account, as they see a customized version of Wikipedia pages that can’t be served as a static cached copy, Is this why we get logged out every 30 days, to boost cache hits for users who rarely need to be logged in? (It seems like every time I want to make an edit I have to login again.)
- virtuabhi 12y agoFacebook is doing a lot of awesome open-source work. They already have Open Compute, many projects on Github - https://github.com/facebook https://github.com/facebook, and sent a developer to MediaWiki to help with the migration to HHVM. I hope Facebook keeps open-sourcing their internal projects (in addition to contributing in existing ones)!
- illumen 12y ago= How to make it take 0.0 seconds to 1.0 seconds. Save early, before the Editor actually presses save. Only commit the change if the Editor actually presses save. This will improve the Editor experience by making save faster at the expense of CPU time. Predict well enough, or do enough processing client side, then you won't need extra server side CPU used. What is more precious to you? The human Editors or some dumb pieces of silicon?
- RobAley 12y agoWhen you don't have much funding, both are extremely important, and when you can't buy more silicon, you either need to optimise the code or let the human wait.
- illumen 12y agoSure. However, the method I mentioned can be done with zero extra server load if done correctly. I've done this before with other editors, and it definitely did not take a team 12 months to complete. It is a full stack optimisation however, from UI all the way through the DB and the rest of the stack. Now that they're using only 10% load, there's plenty of room to breath. Saving a draft of what someone may have spend hours typing into a crash prone browser is good UI practice regardless.
- RobAley 12y agoAgain, "full stack optimisations" cost money, if there's not the money to throw more machines against it (which is often, though not always, cheaper than dev time) then there's not likely to be the money for extensive dev time either. Saving a draft is indeed good UI practice, but best practice is never "regardless" in the real world, particularly in cash/manpower strapped non-profits.
- illumen 12y agoThere was over a year of dev time effort put into this optimisation by a team of people. I think the effort of my optimisation would cost less time and effort, and give better results. It would also waste less Editor time. Mediawiki already allows drafts to be saved.