10 ms·
TiDB – a global scale distributed DB
- BlackjackCF 9y agoCongrats to PingCAP on the 1.0 launch. I've used TiDB for some side projects and I was really impressed. Excited to see how things continue to develop!
- api 9y agoPaging Jepsen. Paging Jepsen. Edit: some work's been done: https://pingcap.github.io/blog/2017/09/01/tidbmeetsjepsen/ https://pingcap.github.io/blog/2017/09/01/tidbmeetsjepsen/
- anishathalye 9y agoThey're also working on using a custom-built system written in Golang to test TiDB (though this is in the early stages): https://medium.com/@siddontang/use-chaos-to-test-the-distributed-system-linearizability-4e0e778dfc7d https://medium.com/@siddontang/use-chaos-to-test-the-distrib...
- siddontang 9y agoHi @anishathalye, thanks for your https://github.com/anishathalye/porcupine https://github.com/anishathalye/porcupine. It’s a great project! Would you like to help me to check whether our model check is right or not with Porcupine? That would be highly appreciated!
- anishathalye 9y agoSure! Sent you an email.
- foohak42 9y agoSo, how does it compare to cockroachdb? So far i've seen: - Apache v2 license - Aims at compatibility with Mysql vs postgres for cockroach
- shenli3514 9y agoThere are some key differences between TiDB and Cockroach. 1. User interface and eco-system Despite that TiDB and CockroachDB both support SQL, TiDB is compatible with MySQL protocol while Cockroach chooses PostgreSQL. You can directly connect to TiDB server with any MySQL client. 2. Architecture The whole TiDB project is logically divided into two parts: the stateless SQL layer (TiDB) and distributed storage layer (TiKV). As TiDB is built on top of TiKV, developers have the freedom to choose to use TiDB or TiKV, depending on their own business. If you only want a distributed Key-Value database, you can just use TiKV alone for higher performance and lower latency. In a word, our system is highly-layered and modularized while CockroachDB is a P2P system. The design of our system results in the fact that we use two programming languages: Go for TiDB and Rust for TiKV to improve the storage performance. And benefit by the highly-layered architecture, we build another project[1] to run Apache Spark to on top of TiDB/TiKV to answer the complex OLAP queries. It takes advantages of both the Spark platform and the distributed TiKV cluster. 3. Transaction model Even though CockroachDB and TiDB both support ACID transaction, TiDB uses a model introduced by Google’s Percolator. The key feature of this model is that it needs an independent timestamp allocator. Like Spanner, each transaction in TiDB will have a timestamp to isolate different transactions. The model that CockroachDB uses is similar to the TrueTime API that Google described in its paper. However, unlike Google, CockroachDB didn’t build the atomic clocks and GPS receivers to keep the time consistent across different data centers. Instead, it uses NTP for clock synchronization, which leads to the problem of uncertain errors. To solve this problem, CockroachDB adapts the Hybrid Logical Clocks (HLC) algorithm. 4. Programming Language TiDB uses Go for the SQL layer and Rust for the storage engine layer. As Go has a Garbage Collector (GC) and runtime, we think it will cost us days to tune the performance. Therefore, we use Rust, a static language, for TiKV. Its performance is much better. CockroachDB only uses Go. [1] Spark on TiKV: https://github.com/pingcap/tispark https://github.com/pingcap/tispark
- foohak42 9y agoThanks for the details! I'm pretty sure cockroach use RocksDB for the underlying storage so it's written in C++.
- eloff 9y agoFrom their blog: "TiKV is a distributed key-value database. It is the core component of the TiDB project and is the open source implementation of Google Spanner." I think the Cockroach DB guys would take offence at that. So how do they use a Spanner-like design without specialized hardware? I had to really dig but the answer is "We are using the Timestamp Allocator introduced in Percolator, a paper published by Google in 2006. The pros of using the Timestamp Allocator are its easy implementation and no dependency on any hardware. The disadvantage lies in that if there are multiple datacenters, especially if these DCs are geologically distributed, the latency is really high." Using Spanner's design but without the real hardware that makes it practical seems like a step backward. A Calvin[1] based approach like FaunaDB would probably have been better - but they can't do that and still have MySQL compatibility (you can't do interactive client-server transactions with Calvin.) I'm skeptical of a database built over a generic KV-store. There was a much-hyped database a few years ago that failed to live up to the hype because of exactly that architecture. I can't even remember the name now when I was trying to find the post-mortem analysis for it in Google. I'm also skeptical of a database claiming to be good at both OLAP and OLTP. One requires a column store, the other a row store. You can be half-decent as an OLAP store and good as a OLTP store. There are no column-stores that are also good at OLTP. But OLAP is the real big-data problems of today, and using half-decent for that is likely to get you in trouble. There's no reason a database can't do both by using two separate storage engines under the hood, but that doesn't seem to be the case here. [1] https://fauna.com/blog/distributed-consistency-at-scale-spanner-vs-calvin https://fauna.com/blog/distributed-consistency-at-scale-span...
- zzzcpan 9y agoKV-store is probably the reason why they can claim OLAP, since distributing batch load on top of it is easier.
- nemothekid 9y ago>There was a much-hyped database a few years ago that failed to live up to the hype because of exactly that architecture. The only database I can remember that did something like this was FoundationDB. They had the core layer as a K/V, and a layer above that was FSQL or similar. They got acquired by Apple, and most of their material was taken off line. I don't know if they were a failure (IIRC, some people were pretty impressed with the core NoSQL DB), but an exit to Apple may not have been all bad - although I would bet that Apple is not using the technology internally. IME, my beef with the new NewSQL solutions is that they tend to be worse at both NoSQL and SQL workloads (specifically none are "the best" at OLAP and OLTP) except at one specific use case - global-replication. Getting global replication requires a cost (either monetary like Spanner, or performance wise like everything else), and global-replication just doesn't seem like a killer feature to me. Most times I need data in two datacenters globally, I wouldn't replicate the data anyways (don't want EU data in the US), and other times my SLAs just aren't tight enough to warrant the cost.
- olegkikin 9y agoBenchmark from half a year ago: https://pingcap.github.io/blog/2017/05/23/perconalive17/#sysbench https://pingcap.github.io/blog/2017/05/23/perconalive17/#sys... Previous discussions: https://news.ycombinator.com/item?id=13298664 https://news.ycombinator.com/item?id=13298664 https://news.ycombinator.com/item?id=10180503 https://news.ycombinator.com/item?id=10180503
- shenli3514 9y agoWe will release a new benchmark result soon.
- itaifrenkel 9y agoHow would you run spark on top of the production db without affecting its performance?
- c4pt0r 9y agoTiDB doesn't want to solve all the problems, no silver bullet, right? And I think it's all about workload, some analytical workload requires CPU/Network resource instead of I/O, and OLAP workload isn't that frequently. Storage layer and computing layer are separated in TiDB stack, I think it's possible for some workload.
- itaifrenkel 9y agoThe problem is that OLAP is run by analysts/business/data.science people and they may mess up with their queries/workloads. without isolation the production performance could impact the user expirience
- shenli3514 9y agoWe use different isolation levels and priorities for OLAP and OLTP workload.
- galkk 9y agoTheir references to use cases from an article sounds funny "Migration from MySQL to TiDB to handle tens of millions of rows of data per day". Come on, even Excel can handle tens of millions of rows of data per day.
- buryat 9y agoExcel can not handle more than 1048576 rows https://support.office.com/en-us/article/Excel-specifications-and-limits-1672b34d-7043-467e-8e27-269d656771c3 https://support.office.com/en-us/article/Excel-specification...
- aneutron 9y agoNot to mention that file locking would make it literally impossible to concurrently edit the file.
- otterley 9y agoContrary to the assertion in the press release, there's no evidence Abraham Lincoln ever said, "the best way to predict the future is to create it." That would be Alan Kay. https://quoteinvestigator.com/2012/09/27/invent-the-future/ https://quoteinvestigator.com/2012/09/27/invent-the-future/
- deleted 9y ago[deleted]
- jinqueeny 9y agoThat’s interesting! I googled it too, but found out it could be Abraham Lincoln or Peter Drucker...
- qq66 9y agoAnyone who has read any of Lincoln's papers knows that simply isn't the way he would talk, or really anyone from his time period. The language that we use today around innovation, envisioning futures, boldly creating the world we want to live in, etc. is post-WWII.
- scroot 9y agoThis is totally bizarre. It must be a joke, right?