7 ms·
Continuous Deployment at IMVU: Doing the impossible fifty times a day.
- patio11 18y agoOne of the greatest lines I have ever read on a blog: It may be hard to imagine writing rock solid one-in-a-million-or-better tests that drive Internet Explorer to click ajax frontend buttons executing backend apache, php, memcache, mysql, java and solr. I am writing this blog post to tell you that not only is it possible, it’s just one part of my day job.
- ntoshev 18y agoThese are automated functional tests. It would be interesting to know if the author uses unit test too and how he writes tests in general. Also some code:test ratio metrics would be indicative of how much weight does all this testing add.
- jmathes 18y agoI work at IMVU as well, know the author personally, and can tell you that yes, he unit tests too. In fact, like all of us, writes all code test-first. For test strategy in general, you need a whole seminar. There's a good blog on testing over at http://googletesting.blogspot.com/ http://googletesting.blogspot.com/ <- go there! As for code:test ratio, I'm not sure what metric you want, but using just php files, there are ~1k active test files and ~7k php files total. So if you take into account that 1k of that 7k is the test, and some is test-setup and obsolete files we no longer use, it's maybe about 5:1 code:test. However, a lot of the code is third party OS software packages. In-house code is about 1.5:1 test:code, in my experience. These, by the way, are very rough estimates.
- ntoshev 18y agoThanks a lot, this is pretty much what I wanted to know. Looks like there is not too much overhead and the results are amazing.
- plinkplonk 18y ago"writing rock solid one-in-a-million-or-better tests that drive Internet Explorer" I find this unparseable. (English is not my native language). As far as I know "one in a million" means something like "very rare". Help?
- briansmith 18y ago"rock solid" is "very reliable" "one-in-a-million-or-better tests" is "tests which fail less than one in a million times". Our Internet-Explorer-based tests are very reliable; they fail less than once per million executions.
- sbraford 18y agoWhile the "one in a million" better is a cool blurb, what does it really mean? Let's say your team makes 25 commits per day. 25 * ~300 working days = 7,500 commits per year That would take 133+ years to reach 1 in a million. The more interesting metric to me is how often the build gets broken.
- nuclear_eclipse 18y agoThe part you're missing is the 15000 tests, multiplied by a new commit every 9 minutes, which in 8 working hours, is roughly 50 commit-test cycles, so 750,000 tests run in a day's timespan... Edit: of course that assumes a peak commit rate matching or exceeding the commit-test cycle period. The point being that even a considerably low rate of failure in the testing mechanism could manifest itself as a blocked commit-test-deploy cycle at least once a day, hence the importance placed on rock-solid testing systems that should only ever fail when the tested code itself fails.
- TimothyFitz 18y agoWe empirically have on average 70 builds a day. The number is higher than your calculation because we don't all work 9-5, we're commiting frequently from around 8am to 9pm. We also run builds repeatedly overnight to flush out any intermittently failing tests we may have recently introduced. We'll run the builds as fast as they can go from 2am-4am.
- edw519 18y agonot only is it possible, it’s just one part of my day job A necessary, but not sufficient, requirement for a nimble start-up.
- mechanical_fish 18y agoDon't be too disappointed if a single submission gets a lukewarm or confused response on HN. The upmods and comments on here are a lot less consistent than what you're used to. ;) Just keep writing. It's really valuable. Also, it's clear to me why your daily routine might sound like science fiction to the median HN reader: A lot of programmers have never seen a system like this. As those of us who were online during a specific half-hour period a couple weeks ago can attest, even Google doesn't have a system that's remotely as reliable as this: It appears to be possible to break all of Google search, worldwide, in ten minutes by misplacing a single character in a text file.
- jeremyw 18y agoHmm, on your Google point, we know that they use partial-cluster deployments extensively, and several presentations point to sophisticated testing of these momentary guinea pig users. I wouldn't hold a one-time lack of a sanity check against their total uptime history. Tests ain't perfect.
- mechanical_fish 18y agoI agree that we shouldn't extrapolate too much from this one incident. But it's not like Google's super-secrecy policy gives us much choice. If anyone from Google wants to tell us about their deployment infrastructure and explain why this one incident really was a nigh-impossible black-swan one-in-one-billion-hour freak of nature -- or why Google has sensibly traded away a certain amount of uptime in exchange for a more flexible architecture (or, perhaps, more cash to spend on tasty gourmet pizzas) -- I'm sure we'll all listen with rapt attention. Until then, we get to tease them mercilessly. ;) Meanwhile, I'm sure that the original submitter would agree that tests ain't perfect. If you read the link at the top of this blog post: http://timothyfitz.wordpress.com/2009/02/08/continuous-deployment/ http://timothyfitz.wordpress.com/2009/02/08/continuous-deplo... ...you'll find that this isn't merely an article about automated testing. Automated testing is just a part of the mighty continuous-deployment ecosystem being described here. It isn't even the real heart of that system: The heart is a planned, well-designed, semi-automated routine for rolling back changes in production. They roll out a change to a subset of their servers, monitor for statistical anomalies in the usage patterns of real, live users, and only continue the rollout if there are no anomalies. If they run into trouble, back they go.
- DannoHung 18y agoI tell people that we should aim for this sort of automation and they pat me on the head and say, "No, no, that will never do." I think there's an idea that if something goes wrong because you let an automated system do it, it's somehow much worse than if something goes wrong because there was human error. I don't really understand the reasoning.
- TimothyFitz 18y agoExactly. Drew Perttula put it better than I'll be able to: "IMHO, manual testing has only two advantages: it’s the easiest thing to [try to] do; and it has a lovely accountability chain. You can always blame the developer, and non-technical people will easily accept that this is the “inevitable cost of software engineering”."
- mst 18y agoThe thing I'd be really interested in is how you deal with UI changes - I've never found a satisfactory way to test "is this ugly/confusing" other than letting a few users bang on it on a staging server.
- TimothyFitz 18y agoUI is an interesting problem. The ultimate solution is to have business metrics drive your UI changes, usually in the form of an A/B test. Then you have a clear winner. This A/B would be run separate from the roll out structure (and indeed, we do LOTS of A/B tests). Sometimes that's not possible, for a new feature or for content without a clear business metric to evaluate for. Either way we often have someone manually test new UI, so that we're not exposing users to something fundamentally broken. We usually do this by using the existing deploy system, but turning the frontend on only for QA users. In the end, you do what works and is cheap, and that's usually something slightly different for every project.
- akronim 18y agoThere's a difference between automation and how often your customers see something going wrong. I'm all for the automation. But lets say the error rate is so low that just 0.1% of automated releases go wrong. Rolling out 50 times a day means you'll expose an error every 20 days. Compare that to a monthly, weekly or even daily cycle and you can see you're exposing yourself and your customers to problems without much corresponding gain.
- jacquesm 18y agoAny chance of more detail than you are giving in your posting ? This is extremely interesting stuff, I'd really like to know a lot more about what goes in to achieving this.
- eries 18y agowhat do you want to know?
- jacquesm 18y agoEverything :) No, seriously, I'd be much obliged if you could tell what tools go into your setup, how much of it is created in house - and thus unavailable - and how much of it is off the shelf, preferably open source. I'd very much like to spend time on recreating what you've done there.
- eries 18y agoI've written in light detail about this in a few places; I'd be glad to share more. Here's an assortment off the top of my head. Feel free to ask anything else you'd like to know. http://startuplessonslearned.blogspot.com/2009/02/continuous-deployment-and-continuous.html http://startuplessonslearned.blogspot.com/2009/02/continuous... http://startuplessonslearned.blogspot.com/2008/11/five-whys.html http://startuplessonslearned.blogspot.com/2008/11/five-whys.... http://startuplessonslearned.blogspot.com/2008/09/new-version-of-joel-test-draft.html http://startuplessonslearned.blogspot.com/2008/09/new-versio... http://startuplessonslearned.blogspot.com/2008/12/continuous-integration-step-by-step.html http://startuplessonslearned.blogspot.com/2008/12/continuous...
- jacquesm 18y agoThank you, I'll be reading all of that later today.
- lgriffith 18y agoInteresting idea but.... Looks like meeting that goal would constrain you to write code to be used by a robot and not by a human. There may be many cases where this is both doable and acceptable to the end user. So no problem with that. I am greatly challenged to see how this could be done for a highly interactive, visually oriented, subtle pattern generating response to user input, type application. Computers are still not as bright as earth worms when it comes to generalized pattern recognition. Which means we programmers are about as bright as earth worms when it comes to writing such code. How then could computers automatically test all the software reactions to the wonderful and totally unpredictable behavior of mere humans as they interact with your software? The test cases would expand to consume all the resources available for development. All you would get done is writing all but impossible test cases. At least you wouldn't ship bugs. This does not consider the explosion of combination and permutations of inputs that prohibits exhaustive testing that no matter how many systems you run tests on. It would be much easier and cheaper to go out of business. Your certainty of being free of shipped bugs would be much better than one in a million.
- mechanical_fish 18y agoAs I noted elsewhere in this thread (http://news.ycombinator.com/item?id=475391 http://news.ycombinator.com/item?id=475391), this article is not merely about automated tests. The author says that his company is using continuous deployment because it lets live, human end users bang on the code, as quickly as possible, in bite-sized chunks that can more easily be rolled back and fixed.
- lgriffith 18y agoThen why do such an exhaustive automated test? Why not have your local tests, automated or not, cover the common cases and error conditions to catch programmer stupidities? Then let the actual humans do the strange corner cases. If your design is even close to correct, testing repeatedly tested code is pointless. If your design is corrupt and your implementation is sloppy, no amount of testing is going to save your ass. I do very rapid turns and I am a one man team. I can turn my system in less that 30 minutes and have the user testing it in a live situation on the other coast. If I want 10 turns a day, I can easily do it. Low coupling, high cohesion, clean correct design, and disciplined implementation makes it possible. I agree that doing things in small chunks is a great way to do it but doing the equivalent of a weeks worth of global automated testing for each small change seems like a silly exercise. That is except for the server hardware salesmen and system admin people. The sales commissions and payroll look rather good. The production of real value is questionable. Bang for the buck is as important for testing as it is in any other part of product development.
- amix 18y agoFrom what I have read Facebook use a similar method: commit and deploy often and rollback if something messes up. We also use this method on Plurk.com and have done so for about a year. Thought, IMVU's case is pretty extreme :) The major problem is rolling back client side changes (that are located in scripts or CSS). This is pretty costly to rollback, because of browser cache - we solve this by having real versioning of the static files so we can force a refresh of browser cache (real versioning = script_{timestamp}.js and not script.js?v={timestamp}).
- delano 18y agoWhat do you mean by real versioning? The major advantage to using "script.js?v={timestamp}" is that it maintains a consistent URI for the resource. Whereas with "script_{timestamp}.js", everything that points to it needs to be updated every time it changes. You could create a symbolic link or rewrite rule that directs requests for "script.js" to the latest "script_{timestamp}.js" but it's more convenient to use a URI parameter.
- amix 18y agoThe problem with script.js?v={timestamp} is that it's ignored by some browsers while script_{timestamp}.js isn't. And with script.js?v={timestamp} you can't set good cache headers. Also, if you ever move to a CDN, then you are forced to use real versioning (at least with Amazon Cloudfront). The versioning scheme we use is `md5 hash of name + file contents + file extension` (and not timestamp).
- delano 18y agoI'm not aware of a browser that ignores URI parameters. Moving to a CDN does not force you to put versioning in the path or filename. The URI parameter merely tricks the browser into thinking there is a new file. The parameter itself is otherwise ignored.
- amix 18y agoAmazon Cloudfront forces you to put versioning if you want to expire objects manually, check out this page: http://docs.amazonwebservices.com/AmazonCloudFront/2008-06-30/GettingStartedGuide/index.html?NextSteps.html#Expiration http://docs.amazonwebservices.com/AmazonCloudFront/2008-06-3... (under `Object expiration`). Unless you specify "Cache-control: no-cache" header you aren't really sure how the browser caches your static files (especially if the user is behind a proxy - and even "Cache-control: no-cache" can easily be ignored).
- gfodor 18y agoThis is great stuff, thanks.
- pj 18y agoContinuous deployment is good, but the comments are valid. There is a certain non-zero probability for errors to occur during deployment. Binaries have to be reloaded, database connections have to be reconnected, sessions have to be restored, etc, so the more you deploy, the larger the coefficient before this probability in the "will something go wrong" equation. So, what we do is break up our system into deployment groups where some handful of users gets updated a few times an hour sometimes. We test the deployment on this small set of users, usually they know the change is coming and are ready to test the change in real time. Sometimes we repeat this process using different deployment groups. Test in this one, then test in that one, until we get a final small errorless deployment and then we roll out to the masses. If it is successful, we roll it out to the masses. Your site doesn't have to be /all/ beta or /all/ production. You can have batches of users in different groups.
- shiranaihito 18y agoWhy do all changes have to end up in production immediately?
- teej 18y agoBecause having thousands of real users running your code gives you insight that automate tests simply cannot match.
- shiranaihito 18y agoYes, but they might give you a hard time too if you put out something silly before thinking things through.
- BeefingJection 18y agoThe obvious solution, of course, is to think things through.
- shiranaihito 18y agoRight, but doing that 50 times per day is more challenging than once, for example.
- jmathes 18y agoWhether you commit your code once in one batch at the end of the day or 50 times in 50 smaller chunks, you have the same amount of complexity about which to be careful. In fact it's more complex in the former case, because in the latter, for each push, you know that all the previous pushes are working.
- kragen 18y agoWhen something breaks in production, it's easier to figure out what it was and fix it if you only changed a few things since the last time you updated production.
- pskomoroch 18y agoI was thinking of doing a basic Django + Selenium + Hudson continuous integration how-to blog post, but this blows me out of the water :)
- TimothyFitz 18y agoI would love to see that post. The first question most people ask me is "How do I get there?" and I don't have a great place to point and say "start here" A well written concise introduction to continuous integration / constant testing would be a boon to this community.
- forkqueue 18y agoI'd still like to see that post, this blog post gives the overview, yours should give the details.
- inerte 18y agoI think the original article misguided some people. It all looked very simple, update the code and put it in production. That _is a horrible idea_, as some have noted. What's not horrible is having thousands of tests, on dozen of machines, 9 minutes to-live, with selective updating of users, and rollbacks, as this article has explained. The original post was too light on details, I guess. Its intention was not to be comprehensive anyway, the focus was why recently changed code should be put in production ASAP. But it looked like the author was simply FTPing after commit. And the whole "SOMEONE IS WRONG ON THE INTERNET" thing kicked in.
- TimothyFitz 18y agoHonestly I think it's a gradient. I'm also one of the developers on a hobby project called http://TIGdb.com http://TIGdb.com (Jeff Lindsay is the other, and has written the majority of the website) We don't have a big Continuous Deploy infrastructure, but we also don't have the users and business requirements of IMVU. We started with the usual, completely manual deploys and hard-to-setup sandboxes, and have been iterating towards a fully automated setup ever since. The entire time we've been doing this, we've been committing and deploying often. Our users are patient, because we're giving them something they can't get elsewhere and we're giving it to them for free. As we do introduce regressions, we'll post-mortem them (probably using the 5 why's technique) and we'll slowly evolve a system to prevent regressions. If the site is a success, we'll have evolved a world class deploy system. If the site never makes it that big then we won't have wasted time on infrastructure. It's classic lean startup thinking (even though TIGdb is really just a hobby project).
- sbraford 18y agoJust curious - who maintains the Selenium tets, and how big is the development / "QA" team? I've never worked in a team big enough that it could devote resources to maintaining all of the following kinds of tests: * unit * functional * AND acceptance * plus writing the actual code IMHO, a neutral third-party group like QA should be responsible for writing & maintaining acceptance tests.
- 18y ago
- benn 18y agoHah, sucks to be hacker news. I wish there was clustering on hackers news, so you could specify your friends and you'd see their posts - not the posts of the angry fanboys who seethe on hn all day.
- s3graham 18y agoWhoa, I love the sound of this as far as development process... But what really blew me away is: 3D Chat makes $1M/month? Really? Or did I find the wrong IMVU?
- TimothyFitz 18y agoYep, that IMVU. Here are a few more staggering statistics that we've published: http://www.vator.tv/news/show/2009-01-22-recession-not-affecting-imvus-virtual-world http://www.vator.tv/news/show/2009-01-22-recession-not-affec...
- mhartl 18y agoI don't recall the last time an article linked from Hacker News so quickly and dramatically expanded my notion of what is possible in software development. Bravo!