3 ms·
That's all great, but sadly impractical. I looked at one of the first statements: > GenDB is an LLM-powered agentic system that decomposes the complex end-to-en
by vladich 7mo ago
That's all great, but sadly impractical.
I looked at one of the first statements:
> GenDB is an LLM-powered agentic system that decomposes the
complex end-to-end query processing and optimization task into
a sequence of smaller and well-defined steps, where each step is
handled by a dedicated LLM agent.
And knowing typical LLM latency, it's outside of the realm of OLTP and probably even OLAP. You can't wait tens of seconds to minutes until LLM generates you some optimal code that you then compile and execute.
- menaerus 7mo agoNo, that's not how I believe they intended it to work. They generate the workload-specific engine up-front and not when the query arrives.
- vladich 7mo agoThen why they write the opposite?
- menaerus 7mo agoIf you look into the results, you will see that they are able to execute 5x TPC-H queries in ~200ms (total). The dataset is not large it is rather small (10GB) but nonetheless, you wouldn't be able to run 5 queries in such a small amount of time if you had to analyze the workload, generate the code, build indices, start the agents/engine and retrieve the results. I didn't read the whole paper but this is why I think your understanding is wrong.
- vladich 7mo agoIf they count only query execution time, not everything else, it would make sense though. It also could be practical, if your system runs just a few predefined and very optimized queries.
- menaerus 7mo agoTo my understanding this is akin to what profile-guided optimization (PGO) in C or C++ does.
- vladich 7mo agoConsidering it's just s single Phd student who does this work, I don't believe such a task can be realistically accomplished, even as a PoC / research.
- menaerus 7mo agoWhy not? Even without LLMs it is technically feasible to build custom database engine that performs much better than general database kernels. And we see this happening all the time, with timeseries, BLOBs, documents, OLTP, OLAP, logging etc. The catch is obviously that the development is way too expensive and that it takes a lot of technical capability which isn't really all that common. The novelty which this paper presents is that these two barriers might have come to an end - we can use LLMs and agents to build custom database engines for ourselves™ and our™ specific workloads, very quickly and for a tiny fraction of development price.