7 ms·
Twilio incident and Redis
- mbillie1 13y agoI'm curious if you're using anything other than redis-cli to set the master/slave relationships, and if you have any failover mechanism. I've used corosync/pacemaker for a high-availability redis cluster, but without an awful lot of confidence (we likely misconfigured it, to be fair). Just "slaveof <masterip>" and other redis-cli commands? Or are you using any automated process? Or has anyone else got a great redis failover/HA solution that they'd care to share? (I apologize for this having nothing to do with Twilio; I'm just curious)
- hijinks 13y agoThe best thing out there is redis sentinel. It's in 2.6 and the issue I ran into 6 months ago is not a lot of drivers supported it yet. http://redis.io/topics/sentinel http://redis.io/topics/sentinel
- eblume 13y agoIt's good to see Twilio post this! That being said - yeah, I really am concerned that Twilio is using an ephemeral database to store such important data. Why not simply use Postgres? Is Twilio really making so many transactions per second that Postgres won't scale?
- michaelmior 13y agoThis wasn't posted by Twilio, but the creator of Redis.
- RobSpectre 13y agoTotally agree. Need to clear up a developing misconception - Redis does not serve as the primary store for the account balance of Twilio's customers. The billing system uses a double bookkeeping model common to many high volume designs with Redis as the in-flight data store (e.g. when a call or SMS message is created) with the transaction also stored independently to an RDBMS post-flight (e.g. when a call or SMS message is completed). Clearly however our implementation failed dangerously and did not recover in a manner that meets our customers' expectations. Totally get how such a misconception would occur from a cursory read of the incident report - just need to be clear.
- justincormack 13y agoBut it does appear (if I understand correctly) that sometimes you used Redis as if it was the system of record rather than checking the actual one, hence the rebilling?
- PommeDeTerre 13y agoIs there actually a legitimate performance or scalability need to incorporate a NoSQL database in this case? Ever since NoSQL databases started receiving a lot of hype a few years back, I've witnessed a number of teams use them without any real justification. They'll build unnecessarily complex systems using one or more NoSQL database systems, all while a relational database would be more than sufficient for their needs. In several of these cases, some of the developers have been quite adamant that these NoSQL databases are essential. Then we rip them out, usually because they've been causing problems like described in this case. It quickly becomes obvious that they were never needed in the first place, and won't be needed even in the face of significantly increased load.
- AYBABTME 13y agoSome problems are better solved with key-value stores than with relational databases. It depends on the data model you need to map. Lot of people use relational databases as a big key value store, when really they should use a key value store. It's not necessarily about scalability or performance. =)
- aidos 13y agoYes. And Redis provides an interesting model with very specific performance characteristics depending on how you choose to store your data in each case. If you haven't worked with Redis it might be hard to appreciate how it's a bit more than just a key-value store and how it's a different tool from a relational db.
- threeseed 13y agoNo offence but attitudes like this are the worst. If we all took your advice we would all be still using punch cards or writing everything in assembler. Sometimes you don't need a clear justification to use newer technologies. Perhaps developers just want to experience the significant developer productivity that comes with using many of the NoSQL databases. Also might be worth dropping the whole "SQL is better" insinuation. We have seen some pretty major data loss bugs in PostgreSQL and MySQL recently.
- deleted 13y ago[deleted]
- StavrosK 13y agoWell, it's not "ephemeral". I don't consider it as durable as postgres either, but I am hard-pressed to find a reason why I think that, which is usually an indication of a faith-based, rather than fact-based, opinion.
- lifeisstillgood 13y agoReasons: longevity Postgres has had it's failures under weird confluences of circumstances - and learnt from them a decade or more ago. I would like to see Twilio standby their stack choices, invest time and effort and sponsorship of redis so that in a decade or two people will be taking it on faith that redis is bullet proof (well the combined postgredis datastore:-)
- StavrosK 13y agoThe append-only file just... well... appends every command to a file. That's pretty solid, especially on journaled filesystems, so I can't find any durability-related reason why it shouldn't be used.
- nknighthb 13y agoThe only durability "problems" I've ever been able to come up with for Redis involve human error as a key component, like what Twilio experienced here with the AOF/RDB mixup. The new CONFIG REWRITE command makes it easier to avoid, but it would be nice if it were harder to get this wrong in the first place. What it definitely isn't is fully ACIDic, though for some workloads you can get close if you use it very, very carefully.
- dkulchenko 13y agoRedis is not an ephemeral database when used in Twilio's configuration. When configured with an AOF with fsync set to 'always', Redis will be as durable if not more durable than postgres and the like.
- VladRussian2 13y ago>Redis will be as durable if not more durable than postgres and the like guessing by your username where you may be from, you may be aware that freshness... err... durability comes in only one grade - first grade :)
- recuter 13y agoI understand how the AOF can be considered similar to a postgres WAL file but what makes you say its even more durable?
- pjscott 13y agoTo expand on that a bit: when redis is in AOF mode, all write commands are appended to a journal file. When fsync is set to 'always', this journal file is fsynced on every write. Redis has a reputation as an "ephemeral database", because it's usually used as one, but there's nothing inherently ephemeral about it. This isn't magic; it's just file I/O.
- banachtarski 13y agoNot twilio, not ephemeral, postgres also has durability issues in the many configurations, your durability is still only as good your kernel's disk controller, tps has very little to do with the decision to use redis or postgres.
- aidos 13y agoAs others have said, Redis is resilient in this configuration and by no mean ephemeral. It's worth nothing from the original Twilio post "This cluster is configured with a single master and multiple slaves distributed across data-centers for resiliency in the event of a host or data-center failure". Ironically it may well be that not having the slaves at all could have prevented the outage. I've seen other scenarios in the past where having a (mysql) master-slave arrangement has caused an outage. Root cause was bad configuration and a bug in the app code but had there been no slaves, there wouldn't have been an outage. Just goes to show that it always pays to think carefully about the complexity you introduce into systems.
- bigiain 13y ago"Is Twilio really making so many transactions per second that Postgres won't scale?" Telco billing systems are (or at least used to be) one of the busiest uses of databases. When you're billing in increments of only a few cents per transaction (typical for outgoing SMS), the volume of transactions for any reasonable amount of money can be _huge_. I don't know how soon Twilio are going to get within an order of magnitude or two of AT&T or Verizon, but I'd hazard a guess that their transactional billing database is _very_ busy by just about _anybodies_ standards.
- pbreit 13y agoBut this issue seemed to be with processing payments vs. metering usage, quite a difference in volume.
- bigiain 13y agoI haven't re-read the article, but if I recall correctly the Redis in-memory data was being used to keep a working copy of the current account balance (not the canonical accounting/bookkeeping version) and it was the results of calculations based on the Redis stored data that was triggering the payment processing. So I think it _was_ a problem with a "metering usage" scale system, which cascaded into the payment processing system working as intended, but with erroneous data.
- aphyr 13y agoTwilio's write volume is on the order of tens of thousands of writes per second. https://twitter.com/dN0t/status/360119871318659074?p=p https://twitter.com/dN0t/status/360119871318659074?p=p
- vertis 13y agoIt has now been noted that Redis is not the primary store of this information.
- mountaineer 13y agoHere's the Twilio post-mortem thread on HN: https://news.ycombinator.com/item?id=6093954 https://news.ycombinator.com/item?id=6093954
- VladRussian2 13y agodog pile - reminds about FB outage couple years ago when their in-memory cache machines got simultaneously flushed by software update and as result piled up upon MySQL databases for the refresh. Twilio's prohibition of master restart seems like a solution to a consequence only.
- encoderer 13y agoThat's a truly common problem. Experienced the same thing when I was working at Formspring. We relied heavily on Redis and SimpleDB for caching and when a large portion of the cache was lost the site was pretty instantly DOSd. Not fun at all.
- mountaineer 13y agoTwilio definitely uses ec2, it's been an oft-highlighted choice in many presentations and posts over the years. - http://www.slideshare.net/twilio/twilio-voice-applications-with-amazon-aws-s3-and-ec2-presentation http://www.slideshare.net/twilio/twilio-voice-applications-w... - http://www.twilio.com/engineering/2011/04/22/why-twilio-wasnt-affected-by-todays-aws-issues/ http://www.twilio.com/engineering/2011/04/22/why-twilio-wasn...
- aidos 13y agoVery clear and thoughtful post from antirez, as ever. It's worth reading his post on how persistence works in Redis (and other dbs). It's very interesting and gives great insights as to what goes on down in dbs to try to keep our data safe - particularly for those of us how don't ever interact with that layer directly. http://oldblog.antirez.com/post/redis-persistence-demystified.html http://oldblog.antirez.com/post/redis-persistence-demystifie...
- zackmorris 13y agoWhat caught my attention was where Twilio said the redis-slaves were timing out to the redis-master: http://www.twilio.com/blog/2013/07/billing-incident-post-mortem.html http://www.twilio.com/blog/2013/07/billing-incident-post-mor... I think timeouts should be abolished for the vast majority of software today. The usual reasoning goes something like this: for a TCP connection, if you don't hear from the server for some period of time, you can assume that something is "wrong" and drop the connection. The fallacy is, the TCP connection is not really important to the shared state of two devices. From the very beginning (I'm talking 1970s!), devices should have been using tokens to identify one another, regardless of communication status. The tokens could be saved in nonvolatile memory on servers so that jobs could always continue where they left off. Instead we have a whole slew of nondeterministic pathological cases -exactly- like the one that hit Twilio. If you take on the burden of timeouts, you end up with dozens of places in your code (even more, potentially) where you just don't know what to do if you lose communication. If you don't take on the burden of timeouts, then you can just track each connection and all it costs you is storage space, which is practically free today and getting cheaper every year. With credentials from the client, you don't even have to worry about duplicate connections. You can write your client-server code deterministically and stick to the logic, and easily stress test failure modes.
- donpdonp 13y agoThis resonated with me because I've been building services with zeromq lately. The zeromq bind and connect calls isolate the caller from managing disconnects. Zeromq will reestablish a dropped connection and the client code is none the wiser. Now I have some extra reasoning as to why this is a good idea. Thanks!
- dpe82 13y agoIt's a good idea until the local send buffer fills up and starts silently dropping data. There's no free lunch.
- tlrobinson 13y agoIn theory don't programs already need to handle that case for slow/congested connections?
- MichaelGG 13y agoI do not understand why, when updating a balance from a CC transaction, you wouldn't be using transactions. Start Transaction Update Balances Call CC Processor Commit That would eliminate "the billing system charged customer credit cards to increase account balances without being able to update the balances themselves" -- you don't go call a non-transactional CC processor until you've actually been able to process the update in your own system (which you can easily rollback). If you're worried about Commits failing (due to not using pessimistic locking, for instance), then separate it into two transactions. That way when you go to process the CC the next time, you have a record stating there's already a transaction in-flight. For financial records, I'd expect a bit more care. Sounds like they had proper records, but only as a backup/logging. (Even for telecom, in which I work. There are fully ACID databases that have no problems handling millions of transactions/sec. In-flight balance information is trivial to handle.)
- superuser2 13y agoWhat databases are those?
- PommeDeTerre 13y agoGiven appropriate hardware, it's possible to get astounding performance, safety and reliability out of relational database systems like NonStop SQL, DB2, and Oracle.
- aphyr 13y agoThese ones, for starters: http://www.tpc.org/tpcc/results/tpcc_perf_results.asp?resulttype=undefined&version=5%&currencyID=1 http://www.tpc.org/tpcc/results/tpcc_perf_results.asp?result... On good server hardware, Postgres will happily push 100-200K small transactions per second. Naturally the definition of "transaction" varies, and you'll see vastly different performance depending on contention, locality, indices, etc. I'd expect logging independent events in a single table to be a good deal faster than multi-table transactions, especially those involving contended rows. On EC2, the story is (naturally) a good deal less heartening. I think you'd be hard-pressed to make it to 10K on TPC-B, given reports like http://www.palominodb.com/blog/2013/05/08/benchmarking-postgres-aws-4000-piops-ebs-instances http://www.palominodb.com/blog/2013/05/08/benchmarking-postg... Granted, single-node performance may not be the question to ask here, because this problem is readily shardable by customer ID, and operations can also be buffered in local memory to some extent. Remember: if it fits in Redis, the problem requires, at maximum, a single node's memory and a single core--and can tolerate network latencies.
- Vitaly 13y agoJust like I commented on the original incident report post, I think systems like Redis are not suitable to work as a db for payment processing and transaction storage. Reading through the report I can't imagine something like this happening with a payment system built around Postgres. Not unless you are doing something incredibly stupid. And stupid those guys are not. They are obviously bright guys meaning well, and yet they've designed and implemented payment system with such a bad failure mode. I do understand that they have a LOT of billing events, and have to update customer billable amounts for each of them. But instead of holding the customer balances in Redis and doing payment processing on top of that, my paranoia would most probably lead me to only store 'amount to charge' in Redis and update it as frequently as needed, and store customer balances and transactions in an RDBMS. And only change during actual charge event. This way, if Redis data were to be lost, I'd under-charge my customers and not over-double-tripple charge them. The failure mode becomes less disastrous.
- ra 13y agoIf you're running with AOF then redis is perfectly fine for storing eg: call logs where the transaction is naturally atomical to a single command. In my experience problems have occur because AOF isn't the default persistence setting (Snapshotting is the default in Ubuntu apt at least). So if Redis get's upgraded in an apt-get upgrade then the sysadmin needs to take care not to override the AOF configuration. This is not unlike postgres upgrading from 9.1 to 9.2, for example. Not catastrophic, but boy it'll make your heart pump! I think the best solution at the moment it to not use apt to manage Redis updates so that you have full control over the configuration.