10 ms·
It's good to see Twilio post this! That being said - yeah, I really am concerned that Twilio is using an ephemeral database to store such important data. Why no
by eblume 13y ago
It's good to see Twilio post this! That being said - yeah, I really am concerned that Twilio is using an ephemeral database to store such important data. Why not simply use Postgres? Is Twilio really making so many transactions per second that Postgres won't scale?
- michaelmior 13y agoThis wasn't posted by Twilio, but the creator of Redis.
- RobSpectre 13y agoTotally agree. Need to clear up a developing misconception - Redis does not serve as the primary store for the account balance of Twilio's customers. The billing system uses a double bookkeeping model common to many high volume designs with Redis as the in-flight data store (e.g. when a call or SMS message is created) with the transaction also stored independently to an RDBMS post-flight (e.g. when a call or SMS message is completed). Clearly however our implementation failed dangerously and did not recover in a manner that meets our customers' expectations. Totally get how such a misconception would occur from a cursory read of the incident report - just need to be clear.
- justincormack 13y agoBut it does appear (if I understand correctly) that sometimes you used Redis as if it was the system of record rather than checking the actual one, hence the rebilling?
- PommeDeTerre 13y agoIs there actually a legitimate performance or scalability need to incorporate a NoSQL database in this case? Ever since NoSQL databases started receiving a lot of hype a few years back, I've witnessed a number of teams use them without any real justification. They'll build unnecessarily complex systems using one or more NoSQL database systems, all while a relational database would be more than sufficient for their needs. In several of these cases, some of the developers have been quite adamant that these NoSQL databases are essential. Then we rip them out, usually because they've been causing problems like described in this case. It quickly becomes obvious that they were never needed in the first place, and won't be needed even in the face of significantly increased load.
- AYBABTME 13y agoSome problems are better solved with key-value stores than with relational databases. It depends on the data model you need to map. Lot of people use relational databases as a big key value store, when really they should use a key value store. It's not necessarily about scalability or performance. =)
- aidos 13y agoYes. And Redis provides an interesting model with very specific performance characteristics depending on how you choose to store your data in each case. If you haven't worked with Redis it might be hard to appreciate how it's a bit more than just a key-value store and how it's a different tool from a relational db.
- threeseed 13y agoNo offence but attitudes like this are the worst. If we all took your advice we would all be still using punch cards or writing everything in assembler. Sometimes you don't need a clear justification to use newer technologies. Perhaps developers just want to experience the significant developer productivity that comes with using many of the NoSQL databases. Also might be worth dropping the whole "SQL is better" insinuation. We have seen some pretty major data loss bugs in PostgreSQL and MySQL recently.
- badclient 13y agoHow are NoSql databases more productive than SQL?
- PommeDeTerre 13y agoBilling systems are not to be taken lightly, namely because money is inherently involved. When developing such systems, it is irresponsible to use new, unproven technologies without justification. When developing such systems, it is irresponsible to trade off the reliability and safety of the system for some "developer productivity". Such irresponsibility is just not acceptable. Failures due to such irresponsibility should not be tolerated, either.
- newman314 13y agoSo what happens when there is a failed transaction mid-flight? It's already in redis but not in RDBMS. How do you rollback then?
- marshray 13y agoRegardlesss, something about the redis query returning a zero balance was triggering bad behavior in the billing system. Also, reading between the lines in the incident report made it sound as if there might have been multiple teams involved in the troubleshooting and not communicating perfectly. For example, were the redis admins informed that customers had been getting billed repeatedly when the decision was made to restart the billing system? Did they have access to the billing system logs which might have contained errors related to redis being read-only? All in all, big props to Twilio for starting to get customer accounts credited back within 11 hours of the first trouble and even more for their wonderful open disclosure.
- deleted 13y ago[deleted]
- StavrosK 13y agoWell, it's not "ephemeral". I don't consider it as durable as postgres either, but I am hard-pressed to find a reason why I think that, which is usually an indication of a faith-based, rather than fact-based, opinion.
- lifeisstillgood 13y agoReasons: longevity Postgres has had it's failures under weird confluences of circumstances - and learnt from them a decade or more ago. I would like to see Twilio standby their stack choices, invest time and effort and sponsorship of redis so that in a decade or two people will be taking it on faith that redis is bullet proof (well the combined postgredis datastore:-)
- StavrosK 13y agoThe append-only file just... well... appends every command to a file. That's pretty solid, especially on journaled filesystems, so I can't find any durability-related reason why it shouldn't be used.
- nknighthb 13y agoThe only durability "problems" I've ever been able to come up with for Redis involve human error as a key component, like what Twilio experienced here with the AOF/RDB mixup. The new CONFIG REWRITE command makes it easier to avoid, but it would be nice if it were harder to get this wrong in the first place. What it definitely isn't is fully ACIDic, though for some workloads you can get close if you use it very, very carefully.
- dkulchenko 13y agoRedis is not an ephemeral database when used in Twilio's configuration. When configured with an AOF with fsync set to 'always', Redis will be as durable if not more durable than postgres and the like.
- VladRussian2 13y ago>Redis will be as durable if not more durable than postgres and the like guessing by your username where you may be from, you may be aware that freshness... err... durability comes in only one grade - first grade :)
- recuter 13y agoI understand how the AOF can be considered similar to a postgres WAL file but what makes you say its even more durable?
- pjscott 13y agoTo expand on that a bit: when redis is in AOF mode, all write commands are appended to a journal file. When fsync is set to 'always', this journal file is fsynced on every write. Redis has a reputation as an "ephemeral database", because it's usually used as one, but there's nothing inherently ephemeral about it. This isn't magic; it's just file I/O.
- banachtarski 13y agoNot twilio, not ephemeral, postgres also has durability issues in the many configurations, your durability is still only as good your kernel's disk controller, tps has very little to do with the decision to use redis or postgres.
- aidos 13y agoAs others have said, Redis is resilient in this configuration and by no mean ephemeral. It's worth nothing from the original Twilio post "This cluster is configured with a single master and multiple slaves distributed across data-centers for resiliency in the event of a host or data-center failure". Ironically it may well be that not having the slaves at all could have prevented the outage. I've seen other scenarios in the past where having a (mysql) master-slave arrangement has caused an outage. Root cause was bad configuration and a bug in the app code but had there been no slaves, there wouldn't have been an outage. Just goes to show that it always pays to think carefully about the complexity you introduce into systems.
- bigiain 13y ago"Is Twilio really making so many transactions per second that Postgres won't scale?" Telco billing systems are (or at least used to be) one of the busiest uses of databases. When you're billing in increments of only a few cents per transaction (typical for outgoing SMS), the volume of transactions for any reasonable amount of money can be _huge_. I don't know how soon Twilio are going to get within an order of magnitude or two of AT&T or Verizon, but I'd hazard a guess that their transactional billing database is _very_ busy by just about _anybodies_ standards.
- pbreit 13y agoBut this issue seemed to be with processing payments vs. metering usage, quite a difference in volume.
- bigiain 13y agoI haven't re-read the article, but if I recall correctly the Redis in-memory data was being used to keep a working copy of the current account balance (not the canonical accounting/bookkeeping version) and it was the results of calculations based on the Redis stored data that was triggering the payment processing. So I think it _was_ a problem with a "metering usage" scale system, which cascaded into the payment processing system working as intended, but with erroneous data.
- aphyr 13y agoTwilio's write volume is on the order of tens of thousands of writes per second. https://twitter.com/dN0t/status/360119871318659074?p=p https://twitter.com/dN0t/status/360119871318659074?p=p
- vertis 13y agoIt has now been noted that Redis is not the primary store of this information.