4 ms·
- Blurring (safely!) the line between database and the app using it: transparently switch between bringing data to compute, or compute to data. - Comprehensive
by cafxx 5y ago
- Blurring (safely!) the line between database and the app using it: transparently switch between bringing data to compute, or compute to data.
- Comprehensive auto tuning: automatic index creation, automatic schema tuning, dynamically switching between column/row-oriented, etc. User specifies SLOs, database does the rest.
- Deeply related to the previous two points: perfect horizontal scalability
- Configurable per-query ACID properties (e.g. delayed indexing)
- All of the above while maintaining complex SQL features like arbitrary joins, large transactions, triggers, etc.
Sure, some of these are in some form in some existing databases. But none offer all of them.
- vaughan 5y ago> Blurring (safely!) the line between database and the app using it What do you mean by this? > Comprehensive auto tuning...automatic schema tuning You should be able to maintain a logical database schema, and then flip a toggle for things you want to have denormalized (and eventually have the db just do it automatically). Or maybe even take a blob of unstructured data and automatically normalize it.
- cafxx 5y ago> > Blurring (safely!) the line between database and the app using it > What do you mean by this? Consider this app pseudocode: resultset = db.query("SELECT ... WHERE ...") for row in resultset if some_complex_condition(row) do_something_with(row) If some_complex_condition filters out a lot of rows then all the data movement has been wasted: it would be much more effective (if possible) if some_complex_condition was pushed down to the database. But if some_complex_condition is compute intensive then potentially pushing it down to the db may actually slow things down, as the db becomes the bottleneck... so the optimal solution, if it exists, 1) is likely to change over time because of changes in the workload and 2) is unlikely to be determinable before runtime. At the same time, consider the case of the table being selected being almost read-only. In this case, and if it fits, it may make sense to move the data directly to the app, keep it in sync when it changes, and have the query processing (the `SELECT ... WHERE ...`) happen directly in the app. This hints at what I meant by blurring the line: dynamically shifting where computation happens, and where the data lives, depending on the available resources, the workload, and the data. This is partially doable today e.g. with Hazelcast, in that it allows app instances to be part of the "database" (or, as they call it, the "in-memory grid").
- vaughan 5y agoIt is interesting that we use declarative SQL to query from the DB, but then when we get it into the app, we are manually doing further filter, mapping and joining of data. And then denormalizing it into a cache in a separate data store. Losing some benefits of the declarative nature of our queries. And this also applies to the frontend as well as the backend.