6 ms·
> KQuery does not yet implement the join operator. Whilst I applaud this book writing initiative, completing it could easily become a lifetime's work! It will
by refset 3y ago
> KQuery does not yet implement the join operator.
Whilst I applaud this book writing initiative, completing it could easily become a lifetime's work! It will be a fascinating journey to follow along with in any case.
Apache Arrow could (and hopefully will) really shake up the database industry in the years ahead. Whatever eventually supplants Postgres is quite likely going to be based on Arrow - polyglot zero-copy vector processing is the future.
Aside: for anyone looking for a more theoretical overview of databases and query languages, this ~free Foundations of Databases book still holds up well http://webdam.inria.fr/Alice/ http://webdam.inria.fr/Alice/
- c0balt 3y agoIdk, but I'd rather see a new iteration of PostgreSQL, similiar to Hydra[0], with a new engine become the upstream instead of a whole new database. There's a lot of experience about db operation and how to approach MVCC encoded in PostgreSQL that shouldn't be underestimated. [0]: https://github.com/hydradatabase/hydra https://github.com/hydradatabase/hydra
- refset 3y agoAnother Postgres-based project in this vein that makes use of Apache Arrow: https://heterodb.github.io/pg-strom/ https://heterodb.github.io/pg-strom/ > PG-Strom is an extension module of PostgreSQL designed for version 11 or later. By utilization of GPU (Graphic Processor Unit) device which has thousands cores per chip, it enables to accelerate SQL workloads for data analytics or batch processing to big data set. > PG-Strom has two storage options. The first one is the heap storage system of PostgreSQL. It is not always optimal for aggregation / analysis workloads because of its row data format, on the other hands, it has an advantage to run aggregation workloads without data transfer from the transactional database. The other one is Apache Arrow files, that have structured columnar format. Even though it is not suitable for update per row basis, it enables to import large amount of data efficiently, and efficiently search / aggregate the data through foreign data wrapper (FDW).
- cmrdporcupine 3y agoThe MVCC implementation inside Postgres is probably not the greatest, in terms of overall performance. This is a good read https://db.cs.cmu.edu/papers/2017/p781-wu.pdf https://db.cs.cmu.edu/papers/2017/p781-wu.pdf "Postgres ... configurations lead to the worst performance, and the major reason is that the use of append-only storage with O2N ordering severely restricts the scalability of the system.... "
- __all__ 3y ago> Whatever eventually supplants Postgres is quite likely going to be based on Arrow - polyglot zero-copy vector processing is the future. Can you elaborate this? I understand it's a very opinionated statement but still I don't see how "polyglot" and "vector processing" could be considered the future of OLTP and general purpose DBMS.
- refset 3y agoPolyglot means not having to fight with marshaling overheads when integrating bespoke compute functions into SQL, or when producing input to / consuming the output from queries. This could radically change the way in which non-expert people construct complex queries and efficiently push more logic into the database layer, and open the door to bypassing SQL as the main interface to the DBMS altogether. Vector processing means improved mechanical sympathy. Even for OLTP the row-at-a-time execution model of Postgres is leaving a decent chunk of performance on the table because it doesn't align with how CPU & memory architectures have evolved.
- __all__ 3y agoThanks! Honestly, I can't envision a near future where SQL is not the main interface. Happy to see the future proving me wrong here though! Despite I can buy the arguments about how having a better data structure to communicate between processes (in the same server) could help, it's a bit difficult to wrap my mind around how Arrow will help in distributed systems (compared to any other performant data structure). Do you have any resources to understand the value proposal in that area? Same for vector processing, would be great to read a bit more about some optimizations that would help improving Postgres leaving out pure analytical use cases.
- refset 3y ago> it's a bit difficult to wrap my mind around how Arrow will help in distributed systems Comparing with the role of Protobuf is perhaps easiest, there's a good FAQ entry [0] which concludes: "Arrow and Protobuf complement each other well. For example, Arrow Flight uses gRPC and Protobuf to serialize its commands, while data is serialized using the binary Arrow IPC protocol". This will be increasingly significant due to the hardware trends in network & memory (and ultimately storage too) compared with CPUs. I posted about that in a comment a few days ago [1], but it's worth sharing again: > here’s a chart comparing the throughputs of typical memory, I/O and networking technologies used in servers in 2020 against those technologies in 2023 > Everything got faster, but the relative ratios also completely flipped > memory located remotely across a network link can now be accessed with no penalty in throughput The graphs demonstrate it very clearly: https://blog.enfabrica.net/the-next-step-in-high-performance-distributed-computing-systems-4f98f13064ac https://blog.enfabrica.net/the-next-step-in-high-performance... > would be great to read a bit more about some optimizations that would help improving Postgres leaving out pure analytical use cases Unfortunately I don't have a good reference on that to hand but I'll take a look around and reply again soon. [0] https://arrow.apache.org/faq/#how-does-arrow-relate-to-protobuf https://arrow.apache.org/faq/#how-does-arrow-relate-to-proto... [1] https://news.ycombinator.com/item?id=37365816 https://news.ycombinator.com/item?id=37365816 [2] https://www.singlestore.com/comparisons/postgresql/ https://www.singlestore.com/comparisons/postgresql/
- hardware2win 3y ago>completing it could easily become a lifetime's work! How so?
- refset 3y agoI meant that mostly in jest, but in reality the pace of both database research and commercial development is happening faster than ever, so an explanation of "How Query Engines Work" is a moving target. Join algorithms and join planning are especially hot topics with a lot of excitement based on advances in machine learning.
- cmrdporcupine 3y agoIsn't Arrow biased towards analytics workloads? Like sibling commenter brings up, I'm not sure what that brings to the table for OLTP. High performance / ergonomic OLTP in general seems to be somewhat neglected in the last decade or so of DB innovation, because in part I think the needs of the industry have been around mass processing of click/log/event-stream data -- firehose of privacy violation :-). So "insert quickly then let me analyze later" is becoming a highly polished stone, but I think databases still present a very rough demeanour around transactional workloads and the "hip tools" out there reflect that. I'm personally all for supplanting SQL but only if what replaces it has a sound relational basis. As for the Alice book, it's great, but it's quite focused on Datalog. Which I think is awesome and an area of my own interest but not in the mainstream of what people think about in terms of "databases" though I wish it was. Most practitioners in software development unfortunately think of a database as "persistence" and not "data management"... And so the heaps of abuse brought on through ORMs etc.
- refset 3y ago> Isn't Arrow biased towards analytics workloads? Like sibling commenter brings up, I'm not sure what that brings to the table for OLTP. That's right, I was thinking more about the network effects, not performance - see my other reply: https://news.ycombinator.com/item?id=37432563 https://news.ycombinator.com/item?id=37432563 > So "insert quickly then let me analyze later" is becoming a highly polished stone, but I think databases still present a very rough demeanour around transactional workloads and the "hip tools" out there reflect that. > I'm personally all for supplanting SQL but only if what replaces it has a sound relational basis. 100% agreed. I think the status quo UX is far more rough than the performance. Unfortunately it seems performance is what drives most investment these days unless you are a producer of hip tools. > As for the Alice book, it's great, but it's quite focused on Datalog True, although I posted more for the tour of Relational Algebra fundamentals. I guess the SQL section is probably the most out of date part of the book :)
- cmrdporcupine 3y ago"True, although I posted more for the tour of Relational Algebra fundamentals. I guess the SQL section is probably the most out of date part of the book :)" I think a better, more accessible, book for the relational foundations is maybe Chris Date's "Database in Depth: Relational Theory for Practitioners". Though it goes out of its way to avoid dirtying itself with SQL and uses some of his "Tutorial D" concepts instead, it gives a good "practical" look at the relational model that I think "practitioners" would find understandable, though it's not without its eccentricities. "> I'm personally all for supplanting SQL but only if what replaces it has a sound relational basis. 100% agreed. I think the status quo UX is far more rough than the performance. Unfortunately it seems performance is what drives most investment these days unless you are a producer of hip tools." I applied for and took a job at RelationalAI last fall because I loved what they were doing with building out a system that was really strong on the relational fundamentals. They have a kind of Datalog-ish language (but better/more practical imho) that is quite nice. The talk that Martin Bravenboer did for the CMU lectures really sold me on it. It was what I was looking for for years, as purely relational alternative to SQL (without.. going off the rails into wonky badly thought through NoSQL network-hierarchical-graph land.) Couldn't make the job work for personal reasons, but was very promising at first. Unfortunately their direction seems to have gone elsewhere, basically becoming a graph analytics plugin inside someone else's SQL DB: https://relational.ai/blog/pr-snowflake-summit https://relational.ai/blog/pr-snowflake-summit -- I understand why the path, but it's less exciting. Anyways it's interesting to see how people talk about PostgreSQL and SQL generally on this forum. Like, as if we can't do better. But if someone tries to do better they go off making something without understanding relational foundations and you end up with something unsound like Redis or MongoDB.
- tomnipotent 3y ago> is quite likely going to be based on Arrow I don't see the connection. Apache Arrow isn't going to make a b-tree or LSM faster or more efficient. It's not going to make a point look-up query faster against columnar storage, or a range scan faster against row-based storage. It doesn't make distributed quorum faster, or otherwise impact consistency and fault tolerance. Removing or reducing SerDe overhead is great, and for analytical workloads where SerDe can be 30-50% of clock time then something like Apache Arrow is pure magic. For the remaining 9X% of use cases it's not adding any more value then you'd see from protobuf.
- refset 3y agoI agree Arrow by itself doesn't address any novel/fundamental OLTP challenges, but I'm not arguing that the thing which eventually supplants Postgres will succeed because of best-in-class OLTP performance - anyone needing that today is not choosing Postgres anyway (same as ever). The real proposition is having a modern, general purpose workhorse underpinned by an ecosystem with strong network effects and polyglot APIs. Assuming analytical systems continue to gravitate towards Arrow I believe the OLTP world will be dragged along also.