11 ms·
1.1B Taxi Rides on Kdb+/q and 4 Xeon Phi CPUs
- Smca 10y agoAstonishing speed given the scale...
- nnx 10y agoI absolutely love this blog series. Can't wait to read what's next :) First time I noticed (mention of) recap at http://tech.marksblogg.com/benchmarks.html http://tech.marksblogg.com/benchmarks.html
- qume 10y agoIf he made the layout a bit uglier, and made the language more esoteric and generally difficult to understand, this would make a fantastic academic paper. But seriously, what a wonderful world it would be if all papers were this well written.
- chiph 10y agoUnder a second to do an avg across 1.1 billion rows spread over four machines. That's pretty amazing.
- jxy 10y agoFor a columnar database, that's a continuous chunk of memory. Assuming 32bit q defaults to 32bit int, 1.1 billion integers across four machines means each 64-core (with 4 threads/core) KNL chip is averaging over 275M elements of int array, or 1.1M 32bit int operations per thread. Now think again whether that's amazing or not.
- shaklee3 10y agoYou're not accounting for the memory bandwidth at all. Yes, that's still amazing. Try doing that in opentsdb.
- throwawayish 10y ago~4.4 GB in 150 ms are just about 30 GB/s.
- obl 10y agois it ? those things are trivial enough to be entirely bandwidth limited. total_amount is 4 byte, passenger_count is 1 and those are tightly packed in a column layout. streaming through that in 150ms is almost within the reach of a single normal chip with dual channel DDR3 ram. Of course, not quite, and that's discounting the (small) sync overhead but still, no need to shell out 4 big servers, overpriced phi chips and fancy wide bus memory.
- wyldfire 10y agoHow does Phi's MCDRAM compare to GDDR5 (wrt throughput)?
- loeg 10y agoAccording to wikipedia, the fastest GDDR5 can do 256 Gbit per chip[0]. I don't know how many chips are typically used. MCDRAM in the article does 400 GB/s, or 3,200 Gb/s. That would require 12.5 of those GDDR5 chips, assuming they scale linearly. [0]: https://en.wikipedia.org/wiki/GDDR5_SDRAM#Commercial_implementation https://en.wikipedia.org/wiki/GDDR5_SDRAM#Commercial_impleme...
- shaklee3 10y agoGDDR5 can typically do 240GB/s access time on a typical GPU, and there are multiple chips on many cards (Tesla K80). The newer cards use HBM2 and can do 732GB/s (http://www.nvidia.com/object/tesla-p100.html http://www.nvidia.com/object/tesla-p100.html).
- AlphaSite 10y agoHBM2 is 1024GB/s (256 per stack).
- astrodust 10y ago1TB/s is pretty nuts by today's standards but I bet it'll elicit a yawn in ten years time. Amazing indeed.
- shaklee3 10y agoKind of. Nvidia lowered the voltage on their P100 so it does not hit those rates. Theoretically it can go that high, but the power draw was too large. Next gen we'll likely see that.
- throwawayish 10y agoSpeed and size are somewhat independent. Speed is set by the bus frequency and how many chips you use - every chip has a fixed data bus width, typically 32 bits, while overall size is then how many chips you have times how big each chip is. Unlike typical computer memory architectures, where the memory bus connects multiple chips or modules to one controller, GDDR doesn't do that; every slice of the memory controller only speaks with a single chip, strictly point-to-point. (Reducing bus load and layout issues and thus allowing higher clock rates). That's why, with GPUs, it's usually sufficient to say how wide the bus is (often 64 - 128 - 256 - 384 - 512 bits) to get a rough idea of it's performance, since memory clock frequencies occupy a rather narrow range. (However, narrow-bus, lower-end GPUs often don't use the same technology as higher-end GPUs, eg. DDR3 instead of GDDR5)
- smulh76 10y agoUnbelievably fast! Interesting blog tbh.
- gbrown_ 10y agoNice to see some KNL usage outside the traditional large HPC centers :D
- nextos 10y agoYes, I remained a bit skeptical, but it seems to be taking off after latest iteration.
- deleted 10y ago[deleted]
- svan99 10y agoNice writeup. If you are interested in learning KDB/Q, please take a look at this book: http://code.kx.com/mkdocs/qformortals3/ http://code.kx.com/mkdocs/qformortals3/
- picodoc 10y agoor for a really quick high level overview you can use this: https://learnxinyminutes.com/docs/kdb+/ https://learnxinyminutes.com/docs/kdb+/ it's a really beautiful little language once you get into it :-)
- WhitneyLand 10y agoI don't see how these results provide much useful information in terms of being able to say x is faster than y. The hardware doesn't seem consistent across different benchmarks. He says it's fast for a "cpu system", but for practical purposes Phi competes more with GPGPUs. Would this be just as fast with one redis system with 512GB ram? I don't know too many apples to oranges here.
- lorenzhs 10y agoA single machine doesn't have the required memory bandwidth to do this in the same time.
- WhitneyLand 10y agoWhy not? Sounds like the initial data load and indexing are done up front. Once you get past that to run the benchmark it's not clear that a quad channel ddr4 system would be saturated.
- lorenzhs 10y agoA current-generation Xeon E7 has a memory bandwidth of up to 102 GB/s. The article says 90, which is probably a realistic estimate of achievable throughput. But the data is 500GB. So the problem is not with loading it from disk (that can be done beforehand), but getting it from CPU to RAM fast enough.
- throwawayish 10y agoInaccurate; the very point of columnar stores is that if I'm only interested in a column I only expend the memory bandwidth required for that very column and nothing else. Hence typical queries to a columnar store would never stream the whole 500 GB for processing.
- lorenzhs 10y agoRight, a columnar store would touch only a couple (1.1bn * three numerical columns) of those 500 GB. But the question was about Redis :) The author has an overview of the benchmarks with various systems at http://tech.marksblogg.com/benchmarks.html http://tech.marksblogg.com/benchmarks.html – but one can't just compare rows as many factors vary (esp. hardware).
- buckie 10y agoIn my experience, kdb+'s k and q (which is broadly speaking a legibility wrapper around k, which again broadly speaking is APL without unicode) are phenomenally fast for dense time series datasets. Though they can struggle (relatively, still pretty fast) with sparse data that's not really what they are built for. They were built for high-performance trading systems, and trading data is dense. If you like writing dense, clever regexs (which I do) then you'll love k & q. The amount that you can get done with just a few characters is unparalleled. Which leads to, IMHO, their main drawback: k/q (like clever regexes) are often write-only code. Picking up another's codebase or even your own after some time has passed can be very hard/impossible because of how mindbendingly dense with logic the code it. Even if they were the best choice for a given domain, I'd try to steer clear of using them for anything other then exploratory work that doesn't need to be maintained.
- de_Selby 10y agoThat's more on you than it is the language though. You can write obtuse write-only code in any language. I will concede that there is a culture of trying to be a bit too clever on the k4 mailing list but it's perfectly possible to write maintainable code in kdb+
- Cyph0n 10y agoBut from what I recall, the syntax is terse by design - this is not inherently bad, though. In other mainstream languages, you have to go out of your way to write obtuse code (e.g, code golf). I'm guessing that best way to address this issue is through liberal use of explanatory comments.
- de_Selby 10y agoTerse syntax isn't an issue, you just need to get used to reading it. Using 1 character variable names is another matter though. There is no reason not to use camel cased variable names and indent functions, if/else blocks etc, and when written this way the code can be perfectly legible even to non q programmers. Something else that leads people into the write-only trap is that the usual way of working with the language is to use the REPL loop while working, where you can tend to be doing multiple things on 1 line. It's just laziness not to reformat and clean up the code afterwards though.
- mmcclellan 10y agoThis was a good idea for a test. I'll definitely check out the author's other stuff. Commenting briefly on cost: while the article mentions the free 32 bit version early on, the actual benchmarks were done using the commercial version. I've had the impression the comercial version was cost prohibitive for us poor folks. For those interested in experimenting with Xeon Phi though, it looks like you can get started for ~$5k: http://dap.xeonphi.com/ http://dap.xeonphi.com/
- Twirrim 10y agoIf you just want to meddle with a Xeon Phi, you can get some as cheap as $300: https://www.amazon.com/Intel-BC31S1P-Xeon-31S1P-Coprocessor/dp/B00OMCB4JI/ref=sr_1_1?ie=UTF8&qid=1485379661&sr=8-1&keywords=xeon+phi https://www.amazon.com/Intel-BC31S1P-Xeon-31S1P-Coprocessor/... , though that's from the 3100 rather than 7200 family, and so won't perform as fast. That kind of money puts it more into the hobby territory, though.
- protomyth 10y agoIs there a difference in how you program the different Xeon Phi families?
- mtanski 10y agoWith the new generation there is now a difference. This is because this generation has Phis in both addin card (PCIe) and a bootable CPU package (when you run Linux or windows or whatever on the Phi). Generally with the PCIe one you're running something like OpenCL and with system CPU package you run threads and processes like you normally would. Technically you could run software directly on the old addin cards since they boot to Linux but you had handle the distribution, running and communication of your software with the host. (you could run any x86_64 binary)
- VodkaHaze 10y agoDo you still need the costly intel compiler suite to run some C++ code on it? Honestly they make it pretty hard to hack with as a device. It could succeed as a device if they let devs easily create cool applications with it.
- mattnewton 10y agoThis is the comparison to the titan X / mapD article I was looking for. Still looks like the gpu is very competitive. Sort of meta, but Mark's job seems awesome. Gets all these toys and writes about configuring them. (The actual configuring is probably a pain but still)
- anonu 10y agoI've been using kdb+/q for a long time (7 years now) - and can attest to its speed. Objects are placed in memory with the intention that you will run vectorized operations over them. So both the memory-model and the language are designed to work together. Lots of people complain about the conciseness of the language and that it is "write-once" code. I tend to disagree. While it might take a while to understand code you didn't write (or even code you wrote a while ago), focusing on writing in q rather than the terser k can improve readability tremendously. My only wish is that someone would write a free/open-source 64-bit interpreter for q - with similar performance and speed to the closed version. Kona (for k) gets close https://github.com/kevinlawler/kona https://github.com/kevinlawler/kona
- gricardo99 10y agoSomewhat related is Kerf, also by Kevin Lawler. http://www.kerfsoftware.com/ http://www.kerfsoftware.com/ It seems more like a kdb+ competitor than open-source alternative, and isn't using q.
- srpeck 10y agoSome other k-inspired languages to have a look at: - https://github.com/johnearnest/ok https://github.com/johnearnest/ok - https://github.com/zholos/kuc https://github.com/zholos/kuc - https://github.com/ngn/k https://github.com/ngn/k - http://t3x.org/klong/ http://t3x.org/klong/ - https://github.com/tlack/xxl https://github.com/tlack/xxl
- vegabook 10y agoAll lots of fun, but kdb has an eye watering cost of 200k dollars per year per server. Here's hoping some combo of Apache Arrow (also cache aware, much more language stack flexibilty), Aerospike (lua built in), Impala, and others, can finally take on this overpriced product, which has had a lack of serious competitors for 20 years, owing to its (price inelastic) finance client base.
- atemerev 10y agoThe binary file of the database (all-inclusive and statically linked) is around 300 kilobytes (!), which makes is probably the most expensive (non-custom) software per kilobyte.
- geocar 10y agokdb is not statically linked. edit: $ file q/l64/q q/l64/q: ELF 64-bit LSB executable, x86-64, version 1 (SYSV), dynamically linked (uses shared libs), for GNU/Linux 2.6.18,
- kpierre 10y ago> Apache Arrow (also cache aware, much more cross platform) kdb+ is available for raspberry pi, is that cross platform enough? https://kx.com/2016/06/08/kx-releases-raspberry-pi-build-with-libraries/ https://kx.com/2016/06/08/kx-releases-raspberry-pi-build-wit...
- vegabook 10y ago32 bit. Please be serious. You know full well that for all non-toy work kdb is exhorbitantly expensive.
- bladecatcher 10y agoAre you sure about that? I think it's significantly lower than 200k
- gravypod 10y agoAt what point will CPUs out parallelize GPUs and will we be able to move vidoe rendering back onto the CPU? I see that as being something I'd very much like.
- thedarkknight0 10y agoWow, they are some incredibly impressive numbers. Great write up. HT to Mark
- mrcactu5 10y agofor reference the city is New York City taxi rides from 2009-2015 and there's about 500GB of metadata http://tech.marksblogg.com/billion-nyc-taxi-rides-redshift.html http://tech.marksblogg.com/billion-nyc-taxi-rides-redshift.h...
- dunkelheit 10y agoPretty wide array of technologies covered in these benchmarks. I wonder how ClickHouse will fare, should be very competitive.
- stuntprogrammer 10y agoIf the combination of such languages, high-performance hardware, and large scale compute problems is interesting.. the startup I work for in Mountain View is hiring...
- hpcjoe 10y ago:D Hope things are well by you!
- stuntprogrammer 10y agoAnd you too sir!
- pvitz 10y agoHas somebody here experience with Jd and could comment on the status or the performance? Thanks!
- eggy 10y agoThe interpreted Jd is fast, but you need the compiled, commercial license Jd, or you used to, for the speed test. I love J compared with K, but that is because I found it first, and the differences between J and K are minimal, but a different enough to keep me using J.
- scottlocklin 10y agoI think jd is mostly plain old interpreted J. There are a few shared libraries to add functionality to core J, but it's mostly just J. You can probably get a non-commercial license if you ask nicely.
- sndean 10y ago> You can probably get a non-commercial license if you ask nicely. This was my experience. I got a nice email from Eric Iverson in response along with the activation key.
- mtanski 10y agoI've build columnar OLAP databases and database engines in C++ for work. Now I'm doing it in my free time. Based on my experience the Phi and it's architecture is very exciting for OLAP databases workloads. Reasons: - Even in a OLAP database you end up with quite a few places that have very branchy code. Research on GPU friendly algorithms on things like (complex) JOINS and GROUP BY is pretty new. Additionally complex queries will functions and operations that you might not have a good GPU implementation for (like regex matching) - Compression. You can use input data that compressed in anyway that there is a x86_64 library for. So you can now use LZ4, ZHUFF, GZIP, XZ. You can have 70+ independent threads decompressing input data (it's OLAP so it's pre-partitioned anyways). (Technically branching, again) - Indexing techniques that cannot efficient implemented on the GPU can be used again. (Again branching) - If you handle your own processing scheduling well, you will end up with near optimal IO / memory pattern (make sure to schedule the work on the core with local memory) and you not bound PCIe speed of the GPU. With enough PCIe lanes and lots of SSD drives you process as near memory speeds (esp. when we'll have Xpoint memory) So the bottom line is if can intelligently farm out work in correct size chunks (it's OLAP so it's prob partitioned anyways) the the Phi is fantastic processor. I'm primarily talking about the bootable package with Omni-Path interconnect (for multiple).
- andrewstuart2 10y agoI get the distinct feeling this is not the usual price for a Xeon Phi. Still, might keep an eye to see if it comes back into stock. https://www.walmart.com/ip/INTEL-SERVER-CPU-SC7120P-XEON-PHI-COPROCESSOR-7120P-1-2G/147170233 https://www.walmart.com/ip/INTEL-SERVER-CPU-SC7120P-XEON-PHI...
- tmostak 10y agoIt looks like year is extracted from pickup_datetime at ETL, so hence its not a fair comparison against the other databases that do this at runtime in Q3 and Q4. In something like MapD Q3 would be nearly as fast as Q1 (~20ms) without the extract function, which involves relatively complicated math.
- 1024core 10y ago% cat startmaster.q k).Q.p:{$[~#.Q.D;.Q.p2[x;`:.]':y;(,/(,/.Q.p2[x]'/':)':(#.z.pd;0N)#.Q.P[i](;)'y)@<,/ Looks like line noise... :D