4 ms·
I can't make sense of lots of this. - Without self-replicating grey goo, infinite scalability is surely more a property of some kind of networked computer rent
by samhw 4y ago
I can't make sense of lots of this.
- Without self-replicating grey goo, infinite scalability is surely more a property of some kind of networked computer rental business (like AWS) rather than a database.
- What does 'serverless' mean exactly? My understanding is that it denotes a stateless application which is executed to serve a request but doesn't run continually as a daemon. Essentially the aforementioned computer rental business provides the event loop and the program provides the event handler. I fail to see how this is compatible with a database, which is definitionally very stateful. (And that encompasses much, much more than just the data.)
- 'Intelligence': Databases already are intelligent and self-optimising, and have been at least since MySQL/Postgres: https://dev.mysql.com/doc/refman/8.0/en/cost-model.html#cost-model-operation https://dev.mysql.com/doc/refman/8.0/en/cost-model.html#cost...
- 'Fundamentally reliable': This idea could reasonably be described as, uh, 'not novel'.
- 'Distributed globally, locally available': As far as I can tell, this collection of words is entirely devoid of any meaning. It sounds like it came out of a random passphrase generator.
- 'Scale should not come at the cost of performance'. While technically semantically meaningful, this is not novel or interesting, and I'm pretty sure this has been a pleasant daydream for database designers since databases were stored in punchcards. As far as I can see, this is comparable to saying 'houses should not come at the cost of money'.
This feels more like a laundry list of daydreams rather than a meaningful narrowing-down of how future databases will be architected. "It should be infinitely scalable, usable by a toaster, and it shouldn't need a computer to live on. It should be serverless and stateless but also self-optimising and with connection pooling. It should be consistent, available, and, uh, partitions, it should be cool with those too. It should be usable by anyone and perfectly tailored to their needs as well as to the opposite needs. Also..."
- imachine1980_ 4y agoWhat does 'serverless' mean exactly? 1) i think server-less, means ops-less, no manager of server even at scale,maybe you need optimization query but not to deploy databases, migration and that. 3)Scale should not come at the cost of performance is easier said than done, is the same idea automatic cache an thinks like this 2) i think the same i will still use monolith like PostgreSQL use a doubt have more intensive workflows than medium.
- samhw 4y ago> maybe you need optimization query but not to deploy databases, migration and that Well, how do you indicate to it that you want to create a database? Does the database read your mind too? How do you indicate that you want to provision some more because there's a big event coming up? > Scale should not come at the cost of performance is easier said than done That's essentially my exact point, yeah. It's not a novel aspiration, and they don't contribute any kind of solution. It's like saying "computers should cost less and also be gooder". > i think the same i will still use monolith like PostgreSQL Yeah, I think that's probably sensible for many people, though I'd also welcome more innovation. Kleppmann's idea of 'unbundling the database' – i.e. the modern database becoming fragmented into several distinct components – is I think very promising and very probable. (Of course, your business and your production environment may not be somewhere you wish to be a hotbed of experimentation.)
- tshaddox 4y ago> Well, how do you indicate to it that you want to create a database? Does the database read your mind too? I'm not sure what you mean. You still instantiate resources with serverless products. With an AWS Lambda function, you go to the Lambda web console, click "Create a function," and type in the code for that function (of course this can also done via the AWS API). There's no mind-reading going on. For AWS Aurora, you still go to the RDS web console, click "Create a database," choose Aurora as the database engine, etc.
- samhw 4y agoI was joking, with that particular sentence. My point was that there is someone managing the server irrespective, and that there's not a particularly clear metaphysical distinction between 'create my database in this way' and 'instruct someone else to create my database in this way'. Not in the computing world, where everything's already under 17 layers of abstraction.
- voz_ 4y ago> How do you indicate that you want to provision some more because there's a big event coming up? You literally do not. As load starts to increase, you scale up automatically. The word elastic has been used to represent this pattern in the last few generations of cloud and/or computing infra. What's the alternative? Manually ssh into some box and crank up mysql instance by hand? like its 1996?
- infogulch 4y agoMaybe the the article would be better titled "Aspirations for a Future Database".
- samhw 4y agoYeah, I'd agree with that. But even then, it would be more valuable if it committed to at least some meaningful, non-platitudinous positions. Something with which at least one person on the planet might disagree. Not "it would be nice if databases were fast, scalable, reliable, personally tailored to everyone on the planet, ...". For example, my predictions for future databases would be: - They will take over more functions of the average backend codebase: instead of database users, they will have a concept of application users, along with their privileges, and simple CRUD logic will be executed by the database. - Horizontal scaling will be less important than we currently think. Consensus will be handled at a lower level, by networked filesystems or storage engines. (Zookeeper, in the Java world, is a proto-example of what I mean.) This is one instance of the trend that... - Databases will be 'unbundled' (Kleppmann's term). Many of the dull uniform bits will be shared rather than reimplemented. This will happen either through libraries or - more likely and preferably - through separate pieces of software, implementing an interface, which the user will compose. (Rebundling will occur for users who just want a click-and-tick experience.) - Self-optimising will - I agree with the article here - widen in scope. Users won't have to perform housekeeping tasks like creating indices on commonly-queried fields. - Databases won't target a filesystem but block storage. This will accompany a convergence of disk (NVMe) and RAM (NVRAM) towards persistent random-access storage of state. Databases, along with applications, won't think in terms of a rigid distinction between "what's in my process's memory" and "what do I have to expressly commit to the disk with a syscall". Tuple spaces are a precursor.
- infogulch 4y agoAgreed. And I like your list a lot better, thanks!
- ertian 4y agoI suspect the intention is something more like IPFS (https://ipfs.io https://ipfs.io), built on some distributed data structure like a DHT. With that in mind: - "Infinitely scalable" in the sense of the internet, I suppose? If we had a way of paying for storage and indexing in a decentralized way, then you could switch from one 'provider' to another, or host your own data. Data would not be siloed in the way it is by AWS. - I assume "serverless" means you don't have some specific upstream server you must hit with requests. You could submit queries to a local job, which could pass them to any node in the network. - "Fundamentally reliable" because data is replicated, and jobs can be executed by any member of the network. - "Distributed globally" because you could store data from any internet-connected device, "available locally" because, again, you don't have some single-point-of-failure server you need to connect to, or VPN you need to join, or whatever: your computer would be a node on the network, as capable of accessing data and running jobs as any other node. And "intelligence" and "no performance cost" are just aspirations.
- throwusawayus 4y ago> I suspect the intention is something more like IPFS (https://ipfs.io https://ipfs.io), built on some distributed data structure like a DHT. the company's products are all built on top of sharded mysql .. unless they plan to throw away everything they have done so far, i do not think these assumptions are correct!
- ertian 4y agoAh, fair enough. I didn't look into the company, I've just come across suggestions along this line.