4 ms·
Actually, it was a combination of three things: 1. OtterTune Start-up (https://ottertune.com https://ottertune.com) 2. Biological Daughter (https://twitter.co
by apavlo 3y ago
Actually, it was a combination of three things:
1. OtterTune Start-up (https://ottertune.com https://ottertune.com)
2. Biological Daughter (https://twitter.com/andy_pavlo/status/1187841279260004355 https://twitter.com/andy_pavlo/status/1187841279260004355)
3. Pandemic
When the pandemic first started, I had a bunch of CMU students reach out to me saying that their summer internships were rescinded and that they were looking for a project to work on so that they wouldn't have a gap in their CV. I ended up taking on any student that could program C++ even if they hadn't taken my DB class before. It as my way of trying to help. But our research group grew to about 35 people. That was not sustainable and the code quality suffered greatly.
We ended killing the project and now all our self-driving work is done in the context of Postgres (https://db.cs.cmu.edu/papers/2023/p27-lim.pdf https://db.cs.cmu.edu/papers/2023/p27-lim.pdf).
I also now realize that building the DBMS engine first then building the query optimizer second is the wrong order. Our future project is going to start with the optimizer first.
- stakhanov 3y agoAwesome, thanks for the quick response.
- pcthrowaway 3y agoOh hey, please fix the link to macrobase; it redirects to pornhub
- brazzledazzle 3y agoOh wow
- erichocean 3y ago> Our future project is going to start with the optimizer first. What's your opinion of recent attempts like LingoDB, that move the query optimizer into a traditional compiler stack, in this case, MLIR?
- apavlo 3y agoLingoDB is an interesting system. Jana has done great work with it. I like projects that take unorthodox approaches to old problems. The problem with (most) query optimizers is that they take a one shot approach at optimization. I think an optimizer should be built from the groundup to support adaptive query optimization. Something similar to Berkeley's Eddies project from 20 years ago.
- zinclozenge 3y agoDo you know if there is anybody taking this approach? Alternatively, what would you consider to be the current state of the art when it comes to query optimizers?
- erichocean 3y agoOne way to do that might be to merge the query optimizer and executor together and then execute them simultaneously in a self-adjusting code framework.[0] The relevant optimizer variables would then be set initially (using whatever mechanism already exists e.g. heuristics, stats, etc.) and then while the query executes, you continually update those optimizer variables. The self-adjusting code property would cause the query to self-adjust as it ran, while still producing the same end result. I'm sure there are details I'm missing here, but I do believe the general approach could do implemented in LingoDB (or similar) as a compiler transformation, so the actually cost-to-develop this approach would remain tractable. [0] I suspect you'd need to model the whole thing as a streaming network so that you can update the network parts as you go, effectively re-wiring the streams while not invalidating earlier results. So SAC+logic to map from one stream architecture to another. JITs that support de-optimization have to do something similar (with a lot of careful upfront design), so that's at least plausible.
- zinclozenge 3y agoThere's also mutable that compiles to WASM and lets it get JITed by v8 https://github.com/mutable-org/mutable https://github.com/mutable-org/mutable.
- 3y ago
- cmrdporcupine 3y agoI'm curious, wondering if you could explicate why you feel starting from the query optimization end is key? I have my (amateur) guesses, but would love to hear your expert opinion.
- zX41ZdbW 3y agoI'm doing a similar thing - invite every student who is interested, without interviewing or skill tests: https://github.com/ClickHouse/ClickHouse/issues/42194 https://github.com/ClickHouse/ClickHouse/issues/42194 It works if you target for ~10% outcome if you have a good CI system with a decent test coverage and a ton of fuzzing.
- zX41ZdbW 3y agoMoreover, making as many as possible people to learn database engineering and production C++ experience - is one of my goals with ClickHouse.
- lifepillar 3y agoI see that MVCC is still your preferred way of doing CC, and what academic research is mostly focused. I am wondering whether that’s an advantage for in-memory databases specifically. I was once discussing MVCC vs 2PL with an experienced Sybase and SQL Server guy, and he claimed that, when transactions are implemented properly and the database is well-designed (no surrogate keys, in particular), 2PL leads to better performance and no deadlocks, while “readers do not block writers” leads to lots of aborted transactions in a heavy OLTP workload. I verified that (I should still have the code around): lots of conflicts in PostgreSQL vs smooth concurrent execution with no retries in Sybase and SQL Server. I have since heard similar opinions from other SQL Server practitioners: they disable MVCC and rely only on good ol’ 2PL.
- apavlo 3y agoSee our 2014 paper on evaluating CC protocols on in-memory system with high contention / core counts: https://www.vldb.org/pvldb/vol8/p209-yu.pdf https://www.vldb.org/pvldb/vol8/p209-yu.pdf All the protocols regress to the same. This evaluation was only with stored procedures though. It would be worth doing a similar investigation with conversational DB protocols (e.g., JDBC, ODBC).
- sundar28 3y agoThanks for the reminder.. this was a great paper, and I'm curious - why would conversational protocols indicate any difference? Also - curious if any such tests have used libraries such as seastar?
- iamcreasy 3y agoCan you kindly share any reading list or anything similar for someone who wants to get into database research at doctoral level? I am a recent MS(Statistics + CS) graduate working as a Data Engineering, and I am looking for material to learn about the current research landscape and getting ready for grad school application.