8 ms·
Specific technology choices aside, this was an incredible write-up of their migration process: thorough, organized, readable prose about a technical topic. It i
by jack6e 8y ago
Specific technology choices aside, this was an incredible write-up of their migration process: thorough, organized, readable prose about a technical topic. It is helpful to read how other teams handle these types of processes in real production systems. Perhaps most refreshing is the description of choices made for the various infrastructure pieces, because it is reasonable and real-world. Blog posts so often describe building a system from scratch where all latest-and-greatest software can be used at each layer. However, here is a more realistic mix of, on the one hand, swapping out DBs for an entirely new (and better) one, but on the other hand finding new tools within their existing primary language to extend the API and proxy.
Great read. Well done.
- harel 8y agoWriting is their Core business after all :) I agree though, I read it like a fascinating breaking news story
- chronotis 8y agoThis was what was particularly interesting to me - that they went to the effort of writing a purely technical article on the particulars of how parts of their environment operate, and to publish it on their platform even when that's not the sort of content they're known for.
- viraptor 8y agoIt's not their main platform though: > Digital Blog > A blog by the Guardian's internal Digital team. We build the Guardian website, mobile apps, Editorial tools, revenue products, support our infrastructure and manage all things tech around the Guardian
- pdpi 8y agoIsh. From what I can tell, the “Digital Blog” seems to be set up as just another column on the platform.
- vanderZwan 8y agoGiven their employers, one would hope that they get good editorial support though!
- jwdunne 8y agoThis is pretty much how all Guardian articles are formatted. Some of their regular pieces could be called "blog posts" - Felicity Cloake's cooking series comes to mind. Guess it makes sense to reuse the platform that already has the templates than use another platform and reimplement the design.
- wftglf 8y agoHi! Thanks for your comments. I'm one of the authors of this post. It is the same platform at the moment (just not tagged with editorial tags so it stays away from the fronts), though sometimes the team that approves non-editorial posts to the site can be concerned about us writing about outages and things as it might carry a 'reputational risk', so we may end up migrating to a different platform in the future so we can publish more quickly, we'll see!
- AlexCoventry 8y agoDoes SecureDrop run on your AWS infrastructure?
- dmix 8y agoIt’s probably better if they didn’t answer this question...
- AlexCoventry 8y agoIt's a bit disturbing to me that they seem to be using AWS for confidential editorial work. > Due to editorial requirements, we needed to run the database cluster and OpsManager on our own infrastructure in AWS rather than using Mongo’s managed database offering.
- nexuist 8y ago>Since all our other services are running in AWS, the obvious choice was DynamoDB – Amazon’s NoSQL database offering. Unfortunately at the time Dynamo didn’t support encryption at rest. After waiting around nine months for this feature to be added, we ended up giving up and looking for something else, ultimately choosing to use Postgres on AWS RDS.
- deleted 8y ago[deleted]
- 8y ago
- InGodsName 8y agoOld versions of mongo were very bad. We accured lots of downtime due to mongo. But later versions were rock solid and I've matainer mongo installations at many startup and SMEs once you setup alertd for disk/memory usage, off you go. Works like charm 99% of the times.
- kbenson 8y agoMogoDB is proof that with the right strategy, marketing and luck, you really can fake it until you make it. Not that that's really a surprise or was unknown, it's just fairly new to see in the open source ecosystem instead of the enterprise one.
- badloginagain 8y agoI'd posit its more a matter of maturing a new paradigm. There's a lot more edge cases you have to cover as NoSQL became more popular for production-at-scale. SQL has decades of production maturation, and has wider domain knowledge.
- kbenson 8y agoI'm sure there's some of that. But a lot of the early problems were a bit more weighted towards poor engineering in general, IIRC. For example, I seem to recall an early problem was truncating large amounts of data on crash occasionally.
- bigiain 8y agoThat's probably true. Unfortunately the number of my customers who would sign off on just "two nines" is approximately zero...
- hideo 8y agoAgreed! I don't think enough engineering orgs appreciate the value of a great narrative on any technical topic. Part of my duties at work require me to deal with "large" issues. While a solution to them is usually necessary and quick and high quality, I've seen the analyses that come after them vary in quality. Good writeups tend to stick around in people's memories and become company culture, and drive everyone to do better. Bad writeups are forgotten, and thus the lessons learned from them are forgotten as well. This particular article stands out for me. English is not my first language, and I've spent most of my life dealing with very fundamental technical details, so most of my writeups aren't the best. I'm going to bookmark this one and come back to it to learn how to write accessible technical narratives.
- jdietrich 8y agoThe BBC's Online and R&D departments have very interesting blogs, if you like this sort of thing. http://www.bbc.co.uk/blogs/internet http://www.bbc.co.uk/blogs/internet https://www.bbc.co.uk/rd https://www.bbc.co.uk/rd
- revel 8y agoTotally agreed. This is pretty much the definitive guide on how to perform a high stakes migration where downtime is absolutely unacceptable. It's extremely tempting, particularly for startups, to simply have a big-bang migration where an old system gets replaced by something else in one shot. I have never, ever seen that approach work out well. The Guardian approach is certainly conservative but it's hard to read that article and conclude anything other than that they did the right thing at every step along the way. Well done and congratulations to everyone on the team.
- konschubert 8y agoI agree, but it looked them a year if I am reading the article right. In most early stage startups, that would be an unacceptable loss of time. So I don't judge them for doing a one-shot migration even if it causes an hour of downtime. It all depends on the business.
- wftglf 8y agoYeah it did take a long time! Part of this though was due to people moving on/off the project a fair bit as other more pressing business needs took priority. We sort of justified the cost due to the expected cost savings from not paying for OpsManager/Mongo support (as in the RDS world support became 'free' as we were already paying for AWS support) - which took the pressure off a bit. Another team at the guardian did a similar migration but went for a 'bit by bit' approach - so migrating a few bits of the API at a time - which worked out faster, in part because stuff was tested in production more quickly, rather than our approach with the proxy which, whilst imitating production traffic, didn't actually serve Postgres data to the users until 'the big switch' - so not really a continuous delivery migration!
- eecc 8y agoThe article mentions several corner cases that weren’t well covered by testing and caused issues later. What sort of test tooling did you use, Scalacheck?
- brown9-2 8y agoI imagine this is a nice benefit of working at an organization that is primarily about writing.
- luord 8y agoThis is the first thing I noticed too. This was an excellent read.
- ranman 8y agoI was actually a little confused by the article - it seems to go up and down in terms of technical depth. It feels like it was written by several people. The hyperlink to “a screen session” was odd as well ... ammonite hyperlink I get but... screen is a pretty ancient tool... people either know it or can find out about it. Like you link to screen but not “ELK” stack? I like the article but it was a bit hard for me to consume with multiple voices in different parts.
- thanatropism 8y agoscreen is hard to google.