5 ms·
I wonder if the 2020s column store would outperform kdb, which was written in the 1990s with a UI from the 1950s.
by bboreham 9y ago
I wonder if the 2020s column store would outperform kdb, which was written in the 1990s with a UI from the 1950s.
- throwaway7645 9y agoI talked to someone recently who ran kdb using an SSD. Is that the standard approach?
- bboreham 9y agoWhatever works, as much and as fast as you can afford. The kdb disk game is more about serial transfer rates and quantity.
- geocar 9y agoUnlikely. Two reasons are at the top of my mind: 1. The current best efforts in benchmarking are focusing on queries that "look" similar, and yet kdb is still 400x faster than Hadoop for those queries. For example: select avg size by sym,time.hh from trade where date=d,sym in S SELECT sym,HOUR(time),AVG(size) FROM trade NATURAL JOIN s WHERE date=d GROUP BY sym,HOUR(time); To answer this question, the database has to read two or three columns across ten billion rows -- it's hard to be much faster than kdb: 10 billion rows completes in 70msec on kdb, but Hadoop takes something like 30 seconds. The 2020 column store has to do a lot of work to even match kdb, but assuming it does that, and even ekes out a few extra percent of performance on these queries, there's another issue: 2. Most kdb programmers don't write this way. Sure some write their application in Java and send these queries over to the kdb "server" get the results, and do stuff with the results, etc., just like the application programmers that use Hadoop, but most kdb programmers don't. They just write their application in kdb. That means that there isn't an extra second or two delay while this chunky result set is sent over to another process. UDF/Stored Procedures/Foreign procedures are the rest of the world's solution for this problem, and they are massively under-utilised: Tooling like version control and testing of stored procedures just doesn't work as well, and I don't see any suggestion that's going to change in the next decade or so.
- beaumayns 9y agokdb also allows you to do much more than SQL, since a select is more or less just syntactic sugar for a certain set of operations on columns as arrays. The full language is available in the context of the query, or you could just treat your table as a bunch of arrays in the context of a larger program. It's a really elegant way of dealing with large amounts of data, although the downside is that you've typically got to build a lot of the nice-to-have dbms type infrastructure yourself. It would be interesting if BigQuery or Redshift ever figure out that they could have a much more powerful system if they stuck an array language on the front of their storage engines.
- dustingetz 9y agoYes! "code/data locality" is key for real apps. Everyone has seen "numbers every programmer should know" < https://gist.github.com/jboner/2841832 https://gist.github.com/jboner/2841832 > If you're going to do complex data analysis, e.g. machine learning, you want your data access latency to be on the short end of this chart :) When you end up on the long end (like in RDBMS) this is known as N+1 problem. But modern size data doesn't fit into memory. Distributed systems necessarily add latency, and to fix that we add caching, which hurts consistency. I blew up this thread further down about how Datomic's core idea is to provide consistent data to your code; which is the opposite of how most DBMS (including kbd) make you bring the code into the database.
- dozzie 9y ago> it's hard to be much faster than kdb: 10 billion rows completes in 70msec on kdb, but Hadoop takes something like 30 seconds. Indeed, it's hard for exhaustive search beat index lookup.
- geocar 9y ago> it's hard for exhaustive search beat index lookup. kdb isn't using any indexes for this query.