8 ms·
Database internals are becoming less important than developer experience
- orlovs 5y agoCall me old, title theme for me deeply resonates with foundations chapters of loosing important knowledge.
- waynesonfire 5y agonothing to do with age. this is published by "PlanetScale is a MySQL compatible, serverless database platform powered by Vitess." -- they have vested interest in promoting the complexities of DBs. In my view the author has absolutely zero basis to make such a claim.
- cl0ckt0wer 5y agoI would think that having bad internals would create a bad dev experience.
- wmf 5y agoUnfortunately MongoDB proved the opposite. You don't notice the data loss and schema problems until later.
- cl0ckt0wer 5y agoThat is a pretty bad experience in the long run.
- DemocracyFTW 5y ago> You don't notice the data loss Wait, is that a feature or a bug?
- Chiron1991 5y agoFriendly reminder that this post is published on the PlanetScale blog, a company that sells a database SaaS. Beware of the bias. I personally would argue with every single point this article makes, except scalability.
- pm90 5y agoAgreed. The recent splurge in VC money for dev tools startups is going to lead to a lot more articles like this one. I hope developers read it with an eye towards that bias.
- hodgesrm 5y ago> I personally would argue with every single point this article makes, except scalability. Maybe you could be more explicit about what you don't like about their ideas? I personally do like a lot of their ideas, such as the following: > In the future, I’d expect to see a tighter coupling between the frameworks we’re using for reactive frontends – React, Vue, etc. – and the database, via hooks or otherwise. This builds on the behavior that made MongoDB so phenomenally popular, as the article points out. Data management is pervasive in modern applications and anything that makes it easier for devs to implement is goodness.
- eatonphil 5y ago> MySQL, MongoDB, Firebase, Spanner; there has literally never been a better time to be a database user at any level of complexity or scale. But there’s still one common thread (ha!) among them – the focus is on infrastructure, not developer experience. It was my impression that everyone picked (and still picks) MySQL, MongoDB, and Firebase _because_ they were the easiest to use. It seemed like developer experience was by far the most important thing to them (compared to sane behavior initially in the case of Mongo and MySQL, some of which has since evolved).
- IncRnd 5y ago> It was my impression that everyone picked (and still picks) MySQL, MongoDB, and Firebase _because_ they were the easiest to use. I've found that to be the case, except for enterprise development, which has different concerns than how quickly code gets written to use a database.
- eatonphil 5y agoAre you saying that enterprise developers pick MySQL, MongoDB, and Firebase for other reasons?
- 5y ago
- seibelj 5y ago> NoSQL databases are maturing, for sure – we’re starting to see support for transactions (on some timeframe of consistency) and generally more stability. After years of working with “flexible” databases though, it has become clearer that rigidity up front (defining a schema) can end up meaning flexibility later on. So funny to me that NoSQL boosters have only recently understood that designing sane schemas and knowing what order your data is inserted is important for data integrity. It's like an entire generation of highly paid software devs never learned fundamental computer science principles.
- blacktriangle 5y agoThat's exactly what it is. "Self-taught coder" really isn't the right word for what many are, as it implies some form of intentional individual study. More like "self learned to duck tape shit together thanks to Stack Overflow" but we don't have a catchy term for that.
- dragontamer 5y agoTo be fair, relational algebra is hard. That being said: going back to 1970 to read the original "A Relational Model of Data for Large Shared Data Banks" by Codd (the paper which started the relational-database + normalization model) is incredibly useful. But yeah, all of this debate about "how data should be modeled" was the same in 1970 as it is today. ----- SQL doesn't quite fit 100% into the relational model, but its certainly inspired by Codd's relational model and designed to work with the principles from that paper. And strangely enough, legions of authors and teachers and courses do a worse job at explaining relational databases than Codd's original 11 page paper.
- munk-a 5y agoI had a specific class on relational algebra in uni and it is up there with algorithm design and analysis in the realm of classes that actually provided me the most long term value. Relational algebra is a lot easier once you start viewing it as relational algebra - a declarative expression of intent that can be manipulated and re-expressed similar to other purely mathematical statements. Then, when performance tuning becomes the watchword, you take that flexible expression and slice and dice it according to how the DBMS you're working with requires to align it with performance. You always want to think of your queries as complex summoning spells that draw in different necessary resources in some particular patterns and then impose an expression form on that blob of data - then you'll skate through all things SQL.
- fabian2k 5y agoUnderstanding a limited amount of database internals has been very useful to me. There is one aspect of using databases that you simply cannot abstract away and that is performance. If you ask your database a question in a way it is not suited to perform or that isn't supported by indexes performance is not going to be good. And these performance differences are not small once your database has a decent size. And if you tables are really large it's not a question of fast or slow but fast enough or so slow it's indistinguishable from the database being down. Of course to some extent you can simply throw hardware or money at the problem. This certainly works for smaller inefficiencies, but sometimes knowing the database will give you orders of magnitude better performance. Hardware and money also don't scale indefinitely.
- abanayev 5y agoAt some point, DBaaS systems should be able to understand and make inferences about your use cases, to the point where indexes and other performance optimizations are automated whenever you register a new query or something. This would be the new era of database systems, and as the article points out is increasingly true about all “infrastructure” concerns.
- pgwhalen 5y agoOne of the problems with this is that there will always be trade offs. It’s hard to imagine a database understanding the appropriate compromise between read and write speeds for your specific application, for example.
- deleted 5y ago[deleted]
- TekMol 5y agoAs a developer, I have to say that sqlite gives me the best experience. Everything else pales in comparison. Create a database? sqlite3 mydata.db Where is the database? In the current directory How is it structured on disk? It's a single file How do I backup the DB? cp mydata.db /my/backups/mydata.db Do I have to install a server? No Do I have to configure anything? No During setup and deployment I usually I dabble a while with the whole GRANT *.* ON localhost IDENTIFIED BY PASSWORD or something. How do I do that with sqlite? It just works Do I have to close / protect any specific ports? No, it's just a file Which field types should I use for ... ? None. It just works.
- waynesonfire 5y agocalibre uses sqlite and it doesn't support hosting the db on a network drive. https://manual.calibre-ebook.com/faq.html#i-am-getting-errors-with-my-calibre-library-on-a-networked-drive-nas https://manual.calibre-ebook.com/faq.html#i-am-getting-error... i guess that's a bad -dev- user experience?
- TekMol 5y agoYou can certainly host it on a network drive if the network filesystem has the right features and behaviour. The same goes for a local filesystem. Sqlite has certain features it requires the filesystem to have. That is independent of how that filesystem stores the data physically.
- fmakunbound 5y ago> if the network filesystem has the right features and behaviour Which network filesystems are those?
- dogma1138 5y agoSQLite actually works just fine on Windows files shares (and yes with multiple clients since Windows file shares do support file locks) but I wouldn’t recommend it as a remote DB/multi client solution.
- deleted 5y ago[deleted]
- tyingq 5y agoThat does make some sense, especially for databases that have been around a while. Lots of internals were written with hard drives in mind rather than SSDs, much lower amounts of memory, and so on. On the other hand, it's nice when your database works well in even a very limited or old environment.
- klysm 5y agoGiven that’s pretty much the only way you can differentiate in this market, it makes sense for planetscale to believe that.
- deleted 5y ago[deleted]
- jdblair 5y agoThe author skips the first decade of database systems in the 1960s. The oldest databases were not relational! They were hierarchical or navigational. Hierarchical databases were much like a filesystem, but for records instead of files. Navigational databases allowed data to be linked in a network. Look up CODASYL for detail. The relational database design was first proposed in the 1970s.
- gagejustins 5y ago~hello everyone, author here~ I know posts with ThOuGhT LeaDeRshIp titles like this are usually annoying, but I thought it would be interesting to write down some of the lessons I've been gathering as I've spent more time covering and using specific databases. My background is in data science / analytics with a couple of years of more traditional full stack here and there. Broadly we've seen this pattern with infrastructure in general – it's a lot easier to set up a server than it used to be, all things considered. Now obviously if you're a tiny startup, you're more comfortable outsourcing everything to Heroku, and if you're a hyperscale enterprise, you probably want more control on exactly what your database is doing. The thesis here is that on the tail end (hyper scale), things are getting more under control and predictable, and developers there want the same "nice things" you get with platforms like Heroku. Elsewhere in the ecosystem, more and more parts of the stack are getting turned into "simple APIs" (Stripe for payments, Twilio for comms, etc.). And perhaps most interestingly, as serverless for compute seems kind of stuck (maybe?), it may be the case that serverless for databases – whatever that ends up meaning – is actually an easier paradigm for application developers to work with.
- jasonwatkinspdx 5y agoI think the article's thesis is a false and misleading dichotomy. It's absolutely true that a low friction developer experience is necessary for a database product to be successful. But this in no way implies that database internals are being commoditized or relegated to minor importance. Snowflake is a particularly bad example as taking a clean sheet and novel approach to internals is the very fulcrum that creates the easy developer experience. Admittedly it's been a while since I looked at vitess, but my recollection is that it's cross shard functionality is so limited as to make claiming internals no longer matter a bit dubious. The reason there's only a handful of spanner style systems is exactly because the internals both matter and are quite daunting to get right.
- felixhuttmann 5y agoI agree. It is also amazing how different the database systems are that are competing against each other today: Partitioning: 1) DynamoDb: Partitioning is explicit and one of the most important parts of schema design 2) Spanner, Cockroach: Database automatically partitions the key ranges. 3) Postgres: You will probably never reach the scale where you need to partition your dataset! Transactions: 1) Spanner, firestore - no stored procedures, client-side transactions are important 2) Dynamodb: No stored procedures, no client-side transactions, only transactions where all items involved are known by primary key in advance. 3) Fauna, Supabase: Stored procedures are the way to go! You do not need application code, access your database from the client! 4) Postgres: We have everything, use what fits your particular use-case! If database internals did not matter, why are they all doing something different and are sometimes quite opinionated about it?
- fmakunbound 5y ago> Database internals will eventually just not matter Of course you need to know the internals of your database. If you've ever come across a project where the team treated a key/value, or document database as a relational one (probably because the query syntax looks similar), then you will know just how important database internals are.
- withinboredom 5y ago> databases will win based on superior developer experience. I guess rethinkdb really was ahead of it's time.
- askdba 5y agoI work at PlanetScale and this post aligns with my thoughts on https://askdbablog.wordpress.com/2018/04/09/learning-the-fundamentals/ https://askdbablog.wordpress.com/2018/04/09/learning-the-fun... whether you agree or disagree on some parts.
- _jezell_ 5y agoLooks like they let the marketing folks write tech articles again.