3 ms·
I definitely miss having the Morgan Stanley site license for K at my disposal. I can't quite justify the money to get a license, just for me, so I've been slowl
by oddthink 13y ago
I definitely miss having the Morgan Stanley site license for K at my disposal. I can't quite justify the money to get a license, just for me, so I've been slowly assembling a poor-man's substitute built on numpy, pandas, and hdf5.
K itself isn't as consistent as I'd like (odd corners like the auto-casting arrays, especially dict arrays to tables, and skipping NaNs in sums), but it's a great place to start.
J really needs a table or dictionary data type, or at least the tools to make one yourself.
- stiff 13y agoMind to share some of the background of yours? How did you get to know it, what were you doing with it, how did you learn it, what did you like about it?
- oddthink 13y agoSure. At MS, I used it while working on mortgage prepayment and default models. Before we got the K license, I had been using SAS and R. There, I had been dumping aggregated (or sometimes not) CSV files from the database (Sybase) and running my model fits based on those. I had a bit of exposure to A+ then, trying to debug bits of the interest rate model subsystem. I hated it. We were mostly using A+ as the wire protocol at that point, and I had endless problems getting the APL fonts set up, and so on. Non-ASCII was a huge black mark. Once we got the license and expanded the team working on the models, we decided to switch over to using kdb+ for the basic data store. That switch was like going from night to day. I could operate over the entire dataset, rather than just on a subset. I could query things on the fly with real aggregations. Things that were a huge pain (and slow) using SQL were suddenly fast and easy, like "calculate the average default rate grouped by FICO in 25-point buckets, for loans issued after 2005" or "backfill missing data from the first non-missing observation." The first example can be done in SQL, but it takes a whole lot more typing and is (IMHO) much more error-prone. The last, I'm sure you can do it, but it's a nuisance. I first got comfortable with kdb+ as a database query language. I read "q for mortals" by Jeff Borror, and we actually had Jeff around for a while for questions. After that, it we started actually using q to estimate the models. It was great at the data aggregation and was a pretty natural fit. It didn't have anything built-in, but so we wrote a few functions for multinomial logistic regression, and that was good enough. That part didn't have a huge advantage over using something like R. However, the interaction with the large data store was far superior, better enough that it was worth it to write our own routines to do the estimation. We even used q to run the models in production. That was more of a reach, but it worked pretty well. So, I liked kdb+'s ability to handle large data sets in a speedy way. It makes a fantastic query language. I liked q's conciseness. It doesn't matter much for actual in-production code, but if you're doing research, every character counts. It seems shallow, but I think it's a real effect. Writing a half-line of q to do a paragraph of SQL is a huge win. I liked q's array operations, but by then I'd already been using Numeric and numpy and such for a decade, so those weren't really new. It was nice to be able to write them easily, as opposed to something like "np.concatenate((foo,bar))", but it wasn't really novel. Basically, it was kdb+ that sold me on q.
- stiff 13y agoThanks, that's very interesting. Q for mortals is online and seems to be a good read, if anyone's interested: http://code.kx.com/wiki/JB:QforMortals2/contents http://code.kx.com/wiki/JB:QforMortals2/contents
- fhars 13y agoIf you can live with K3, you might want to take a look at kona: https://github.com/kevinlawler/kona/ https://github.com/kevinlawler/kona/