9 ms·
Distributed Databases Should Work More Like CDNs
- cagenut 9y agothis is true, and why fastly is kindof like a globally distributed nearly-cache-coherent key value store that people use as a CDN. there's a great talk on how its done with a custom gossip protocol implementation: https://www.youtube.com/watch?v=HfO_6bKsy_g https://www.youtube.com/watch?v=HfO_6bKsy_g
- loiselleatwork 9y agoThis is super interesting; thanks for the pointer.
- ddorian43 9y agoCDN post with no performance talk (beside keyword). Never mention lower performance (even on single-node). Add to that aws-vps with pseudo-cores and spectre-upgrade and good luck with your tps-reports.
- vog 9y agoWould you mind to elaborate? Your criticism is so condensed that I'm unable to make a lot of sense of it.
- ddorian43 9y agoTheir performance is 0.1-0.03 of postgres. And spectre makes your aws-vps ~0.7x compared to previously. Meaning you can't use it for performance-sensitive stuff (the whole point of fancy sharding and (no/new)sql).
- vog 9y agoThanks! That's something I can make sense of. And I agree that this is perhaps one of the many situations where people throw away the "C" of "ACID" for no reason beyond it is modern to do so. (At least most strive for "Eventual Consistency", but that's another can of worms.) This is even less understandable once you notice that PostgreSQL offers a lot more features than most other databases (SQL or NoSQL) and is extremely flexible and extensible - even if you use it just as a fancy JSON or XML store.
- timkpaine 9y agoImagine a globally replicated and version controlled object store. You could use this for data, assets, source code, anything you like. Would be incredibly useful for both the development side of things, as well as production.
- ddorian43 9y agoWhy not store the data only where you live ? Worst case meteor strikes and you die with your data.
- saghm 9y agoSometimes people travel, and it would be nice to have the data if a meteor strikes while you're gone.
- prepend 9y agoWe could call it the World Wide Web.
- TeMPOraL 9y agoIt would fall short of all of these goals, as it would get turned into advertisement, colorful magazine pseudocontent and malware distribution platform.
- cookiecaper 9y ago
- appdrag 9y agoVery interesting! Have you used it at big scale for production or not yet?
- yuz 9y agoadvertisement
- jjevanoorschot 9y agoI used CockroachDB for a university project, and while I think it looks very promising, I found the tooling and documentation to be a bit lacking. I wouldn't use it in production yet. However, when CockroachDB matures a bit I can really see it take off.
- hexspeaker 9y agoIf this interests you, you'll probably enjoy reading Google's paper on Spanner. Cockroachdb was heavily influenced by it. https://research.google.com/archive/spanner.html https://research.google.com/archive/spanner.html
- dstroot 9y agoCDN: no trade offs. Faster everywhere. More reliable overall Cockroach DB: trade some performance for geographic redundancy. The trade off may work in your favor - e.g. read heavy workloads (or not). I plugged in CDB I place of Postgres for some testing this week, was surprised it worked so well.
- icebraining 9y agono trade offs This is almost never the case, and CDNs are no exception. A CDN like Cloudflare that reuses your domain(s) means that they become just a useless point that your dynamic requests have to travel to and from the main server. A CDN that uses its own domains requires extra DNS queries, extra TCP & SSL connection setup, etc, plus it only starts loading when the browser has started processing the HTML.
- dstroot 9y agoExcellent points. I definitely oversimplified.
- ec109685 9y agoCDN’s can terminate a user’s ssl connection close to the user (and keep a persistent one open to the origin), so it is more than a useless hop. There are also http headers that can instruct the browser to fetch cdn resource before the html is delivered.
- dasil003 9y agoAt my previous company we built our own CDN because we were streaming video to a long tail of global users for a limited library. The problem with off-the-shelf CDNs was they couldn't keep our content cached for the long tail, but for the same price we could replicate our limited library to bare metal. Traditional caching could then go on top of that. This was a huge win for QoS on cache misses, which were a significant portion of our traffic. There are tons of tradeoffs to make this happen which is why we couldn't find an off-the-shelf solution to deliver the same results.
- zzzcpan 9y ago
- astine 9y agoSooo, being able to distribute data globally is good for performance? Who knew? The thing about distributed systems, including distributed databases is that they need to navigate around the the CAP theorem(Consistency, Availability, Partition Tolerance, pick two, essentially,) and every solution is ultimately a trade off. This article would be a lot more interesting if it showed how CockroachDB made a better trade off than the other solutions listed.
- AYBABTME 9y agoThe article they link at the end explains this: https://www.cockroachlabs.com/docs/stable/transactions.html https://www.cockroachlabs.com/docs/stable/transactions.html
- astine 9y agoThanks, I got the answer to questions I was asking from CDB FAQ, actually. I just wish this article in particular was less fluffy.
- marknadal 9y agoAlso a distributed database engineer, and while I disagree with CockroachDB's CAP Theorem tradeoff decisions, they are definitely well reasoned and principled. RethinkDB was Master-Slave (strongly consistent) with an amazing developer community and actually survived Aphyr's tests better than most systems. CockroachDB is also in the Master-Slave (strongly consistent) camp, but more enterprise focused, and therefore will probably not fail, however their Aphyr report worked... but was unfortunately very slow. But hey, being strongly consistent is hard and correctness is a necessary tradeoff from performance. Other systems, like Cassandra, us (https://github.com/amark/gun https://github.com/amark/gun), Couch/Pouch, etc. are all in the Master-Master camp, and thus AP not CP. Our argument is that while P holds, realtime sync with Strong Eventual Consistency is good enough yet has all the performance benefits, for everything except for banking. Sure, use Rethink/Cockroach for banking (heck, better yet, Postgres!), but for fun-and-games you can do banking on top of AP systems if you use a CRDT or blockchain (although that kills performance, CRDTs don't) on top. So yeah, I agree with you about CAP Theorem and stuff, disagree with Cockroach's particular choice - but they do have some pretty great detailed explainers/documentation on their view, and therefore should be treated seriously and not written off.
- zzzcpan 9y ago> When partitions heal, you might have to make ugly decisions: which version of your customer’s data to you choose to discard? If two partitions received updates, it’s a lose-lose situation. When partitions heal you simply merge all versions through conflict-free replicated data types. No ugly decisions, no sacrificing neither latency nor consistency. We call it strong eventual consistency [1] nowadays. And it's exactly like CDNs, except more reliable. I'm wondering, since CockroachDB keeps lying about and attacking eventual consistency in these PR posts, the whole "consistency is worth sacrificing latency" mantra might not work in practice after all. People just don't buy it, they want low latency, they want something like CDNs, something fast and reliable, something that just works. Something that CockroachDB can never deliver. [1] https://en.wikipedia.org/wiki/Eventual_consistency#Strong_eventual_consistency https://en.wikipedia.org/wiki/Eventual_consistency#Strong_ev...
- imtringued 9y agoMy application is just plain old CRUD. What if two users want to change e.g. the telephone number of an existing record during a network partition. There just is no obvious way to merge a telephone number. One of them is correct, the other is incorrect. Can CRDTs solve my simple problem?
- merb 9y agoI think the issue is solved (while still complex) with an event sourcing system. The chances are really high that the users did not push / updated the number at the exact same time. (so they are at least 1ns apart, even when not it's not that bad). So you can actually restore to a sane state by reapplying the log from scratch. i.e. this is solved with CQRS and Event Sourcing and it probably works in like all databases. It's quite complex but pretty reliable, I'm pretty sure that everybody already built at least a extremly simple append only event log.
- elvinyung 9y agoDoes this imply that there's a single place that log entries get ingested at? Doesn't that make it a single point of failure?
- russell_h 9y agoI'm curious, what sort of read latency is achievable with CockroachDB? Does it support some notion of tunable read consistency in order to achieve lower read latency at the expense of consistency?
- susegado 9y agoReads from CockroachDB go through a lease holder for the piece of data being read without needing confirmation from any replicas about consistency, so there is no overhead from replication. But read latency can be affected by writes on the data (conflicts), because it is fundamentally a consistent system (serializable). This is not tunable.
- russell_h 9y agoThanks, that makes sense (and provides good context for the article). So for a cluster spanning multiple regions, one region can support low-latency reads for a given range, but reads in any other region will have to go cross-region to the leaseholder. Being able to move the leaseholder around to optimize read latency makes a lot of sense. It would also be useful to me, in some cases, to be able to perform a read-only query in an "inconsistent" mode to avoid that cross-region latency, at the expense of potentially receiving stale data.
- HumanDrivenDev 9y agoThe only distributed database I'm familiar with is CouchDB. Can anyone give me a birds eye view on how Cockroach is different?
- tyingq 9y agoUnfortunately, it's not on the list for the table of feature comparisons, but here's a start: https://www.cockroachlabs.com/docs/stable/cockroachdb-in-comparison.html https://www.cockroachlabs.com/docs/stable/cockroachdb-in-com... Might be helpful if you already know the couchdb answer for each feature.
- pdeva1 9y agoAm I wrong, or does seems the article not really tell how CDB deals with the latency issue, especially with regards to writes? If the write has to be consistent and available across multiple regions, it will need to synchronously replicate that write to all the regions, thus incurring the same performance penalty as RDS or any other consistent database.
- irfansharif 9y agoI doubt this particular article was intended to address how CRDB handles writes cross-region, synchronous replication, etc. You'll find articles touching on varying aspects of what you're looking for if you dig through some of the earlier posts on their blog[1] or their FAQ[2]. [1]: https://www.cockroachlabs.com/blog https://www.cockroachlabs.com/blog [2]: https://www.cockroachlabs.com/docs/stable/frequently-asked-questions.html https://www.cockroachlabs.com/docs/stable/frequently-asked-q...
- pdeva1 9y agowell the comparison section to RDS specifically claims that RDS is inferiors because "this forces all writes to travel to the primary copy of your data". So it doesn't explain how CDB is superior to RDS, since writes will incur the same penalty in CDB too.
- irfansharif 9y agoAt the key-value level, CockroachDB starts off with a single, empty range (a set of sorted, contiguous data from your cluster). As you put data in, this single range eventually reaches a threshold size (64MB by default). When that happens, the data splits into two ranges, each again covering a contiguous segment of the entire key-value space. This process continues indefinitely; as new data flows in, existing ranges continue to split into new ranges, aiming to keep a relatively small and consistent range size. Each range is replicated 3-way (by default) as well, and is backed by a single Raft instance. When your cluster spans multiple nodes (physical machines, virtual machines, or containers), newly split ranges (or more specifically, replicas of these ranges) are automatically rebalanced to nodes with more capacity. Writes addressed to a range are handled by the Raft leader for that range (which can hop around its various replicas as needed). Writes to different ranges (non-overlapping key spaces by definition) are processed independently, and very well may be processed across multiple machines. Source: https://www.cockroachlabs.com/docs/stable/frequently-asked-questions.html https://www.cockroachlabs.com/docs/stable/frequently-asked-q...
- philbogle 9y agoAs long as we're advertising, consider: https://cloud.google.com/spanner/ https://cloud.google.com/spanner/
- dantiberian 9y agoI love CockroachDB, but feel like this article was a bit misleading in it's claims. I started out writing a comment response, but it got so long that it turned into a blog post: https://danielcompton.net/2018/02/08/add-context-cockroachdb-databases-acting-like-cdns https://danielcompton.net/2018/02/08/add-context-cockroachdb....
- anon1253 9y agoso immutable, like datomic? https://www.infoq.com/articles/Datomic-Information-Model https://www.infoq.com/articles/Datomic-Information-Model Or more like federated SPARQL on RDF? https://www.w3.org/TR/sparql11-federated-query/ https://www.w3.org/TR/sparql11-federated-query/
- daxfohl 9y agoDoes GDPR really require all Europeans' data to stay on EU soil without explicit consent?
- itcmcgrath 9y agoSeems to be a common misconception. Look at EU Model Contract Clauses: e.g. https://cloud.google.com/terms/eu-model-contract-clause https://cloud.google.com/terms/eu-model-contract-clause
- deleted 9y ago[deleted]
- dantiberian 9y agoI don't think so, you need consent to process personal data anywhere in the world, and you can only transfer it outside of the EU if there are appropriate safeguards - https://gdpr-info.eu/art-46-gdpr/ https://gdpr-info.eu/art-46-gdpr/
- supergirl 9y agowhat a dumb article about nothing. hey, a database replicates stuff but so does a CDN! what a brilliant insight. let's make an article with this title but fill it with marketing text.