4 ms·
Being taken down by slow IO in a single cloud zone seems to implicated their overall architecture.
by seriesf 7y ago
Being taken down by slow IO in a single cloud zone seems to implicated their overall architecture.
- jrockway 7y agoLooking at their job postings, they use Cassandra and Elasticsearch, so I am somewhat surprised that they didn't survive this outage. (I have not ever run Cassandra at scale, but their website says in the first paragraph that it's designed to handle regional failures.) If I were on GCP and had paying customers, I would use Cloud Spanner. It's expensive, but it's a good piece of technology.
- sieabahlpark 7y agoMost startup companies worry about building out features than building out reliable infrastructure. Almost always an after thought
- seriesf 7y agoThe gift of Spanner is it’s always slow, so if you start there you will naturally gain experience in either hiding latency or relaxing consistency requirements, which are both good skills.
- vl 7y agoSpanner is pretty fast, but as with any database you have to use it correctly to get the speed. Since it’s a non-standard database, it requires non-standard tricks and skills.
- jrockway 7y agoNo, that's about right. You get one write per second per entity or something like that and that's as fast as it goes. The key is to choose your entity wiseley. I wrote an app inside Google that did on the order of 10,000 writes per second and it was no problem; each entity group only showed up once every minute or so.
- vl 7y agoAFAIR, entity group is a Megastore concept. Megastore is ill-conceived predecessor of Spanner, are you confusing them? Spanner is much faster than 1 transaction per tablet.
- psanford 7y agoI have run Cassandra in production. A replication factor 3 cluster over 3 AZs should be able to survive loss of a single AZ. However, it has been my personal experience that Cassandra handles instances going down much better than it handles instances getting very slow but staying online. In that case Cassandra will work extra hard to just spin its wheels. If you know that is what is happening you can be better off manually bringing those nodes offline.