7 ms·
> It turns out my day to day work doesn’t require deep knowledge about database internals and I can mostly treat them as a black box with an API 2. Because of
by ivanhoe 5y ago
> It turns out my day to day work doesn’t require deep knowledge about database internals and I can mostly treat them as a black box with an API 2.
Because of such attitude of the previous dev team my client ended up with DB integrity ruined. Previous guys somehow didn't know they should use transactions when updating/deleting stuff in the DB, because hey, it's just API you call, who cares of mambo-jumbo happening behind the curtains, right?
I myself had a similar fuckup with MongoDB a decade ago, because I used it "because it's fast", but failed to read the small letters, and wasn't aware that it achieves that speed by buffering the data for async save, and so has no guarantees it will actually be saved at all. So I lost a lot of data when traffic suddenly jumped, and since those were affiliate clicks lots of people got very pissed about not getting paid, and it almost ruined both my client's business as well as my own, as it was our core client.
Lack of understanding of (or even caring to know about) DB internals is also why like 90% of projects I've seen has wrong indices, or often no even a single index set at all. Because, hey, it's someone else's job to think about it, but in reality budgets are limited and team don't have a dedicated DBA or even a capable devops person to notice it, and it just ends up in the production like that.
So, in summary: no you don't need to know every in-and-out of every technology out there, no one can do that, but there's a valid requirement that one has to be familiar to those that they actually use - at least those that can seriously bite one's ass and cost a lot of money. And all DB related stuff is very much like that, that is probably the most critical part of the whole system, so no, it's definitely not a black box.
- kgeist 5y ago>Because of such attitude of the previous dev team my client ended up with DB integrity ruined Same thing here, our legacy codebase avoided using transactions, and to this day, after migrating to a new codebase, we still get reports from clients here and there about errors which stem from inconsistent data
- jhgb 5y agoDamn, and I've always felt like I must have some kind of OCD whenever I'm getting a refresher on MVCC/MGA isolation levels and their guarantees and the associated auxiliary locking (say, select for update with lock in Firebird, for example) just to be sure I'm not writing something stupid. Simply not using transactions at all is something I can't even fathom. That's basically no better than blindly mutating shared in-memory data structures in a multi-threaded program. Who would write something like this?
- Aeolun 5y ago> Who would write something like this? Many, many people.
- joshuanapoli 5y agoExactly. Maybe not everyone needs to be an expert in databases and distributed systems, but if you’re developing web technologies, then someone needs to be the expert. They will need to have the vocabulary to communicate with the rest of the team.
- hannofcart 5y agoI think it is possible that an engineer knows about the need to use transactions while doing certain kinds of DB updates while not knowing the exact definition of the ACID acronym. I suspect the author of this article would have no objections to probing if a candidate understands transactions and indices.
- EdwardDiego 5y agoExactly, I've spent years optimizing RDBMS and columnar databases and I couldn't at this moment fully explain the definition of ACID to you without having a wee google refresher. I mean, I learned ACID's definition, nodded because it made sense, realised that MySQL's default (at the time) MyISAM storage engine wasn't ACID, it flat out failed at C, so encouraged my LAMP stack using friends to switch to InnoDB, then switched to Postgres myself. Likewise, I use a lot of distributed systems, but I don't think in terms of the CAP theorem on the regular. Rather I just think about all the fun ways each distributed system can die in, and while ZK is CP (although I'm dubious about the P), and Kafka is CA, the failure mode is roughly the same - you lose enough cluster members and shit breaks. I know that my k-safety (k being how many nodes I can lose and still have a system) for ZK is N - (N +1)/2 (assuming truncating int division), and as for Kafka, well... it really depends on the failure. But the usual conservative design for a Kafka cluster is n * 3 replicas on a 3 AZ stretch cluster with min.insync.replicas set to num.replicas - 1,leading to a k safety of 1. I'm a fan of the 2.5 stretch cluster, if I have to do a stretch cluster, I generally prefer separate clusters replicating. So for a 2.5, you'd go 2 AZs with M brokers and N ZK nodes, where M % 2 == 0 && N % 2 == 0, with one ZK in the third AZ as the tiebreaker. You can now lose an entire AZ, and still have a quorum, while paying 33% less inter-AZ traffic from your Kafka clients. Basically, theory underpins all my work, but the theory and the reality are quite different things.
- edw519 5y agoThank you, ivanhoe! When I read that sentence, I gasped and thought the same thing. I've spent a career writing thousands of lines of application code that would have never been needed if the database had been designed properly. But it wasn't because the deployer didn't think they needed "deep knowledge" that day. We stand on the shoulders of giants who figured out the fundamentals just for us. The worse the violation (different data types in the same column, ugh!), the more application code (10X) will be needed. The more code written, the more more maintenance (another 10X) needed. And none of it ever gets replaced, just migrated to something else years later (if we're still in business). ASIDE: As for the interview question, I would never ask what "ACID" meant. I'm more interested in understanding than recall. I would just give the candidate a poorly designed database and ask what they would do to fix it.
- emilsedgh 5y agoAnd that's why you have senior engineers and junior ones. If you are trying to hire, expecting every junior dev to know all this stuff means you'll skip a lot of candidates who can contribute a lot to your business. They shouldn't be trusted with writing things from scratch. But if your codebase is mature enough, getting db connections, making endpoints atomic and stuff like that already should have a convention so your junior engineer wouldn't have to do it on their own.
- TheRealDunkirk 5y ago> I've spent a career writing thousands of lines of application code that would have never been needed if the database had been designed properly. And then there's the time that I rewrote a small database application, but couldn't really reduce the byzantine schema, as that's just how our business works. However, I did reduce the main query from 727 lines to 93, and made it run twice as fast. And then, of course, the project was canceled because of politics, and the group still struggles with their terrible application, 3 years later. I should really write a book about terrible management in Fortune 250's.
- deleted 5y ago[deleted]
- ZephyrBlu 5y agoThe author says, "I wouldn’t be able to explain from the top of my head what the ACID term means". I think there is a big difference between not knowing something off the top of your head and being completely unaware of it. If you have awareness of something, you can make decisions based on that awareness without necessarily having a deep understanding. When you have no awareness is when the big issues arise. For example, with your MongoDB fuckup if you had even a vague idea about the underlying behaviour you might have been able to consider the failure scenario ahead of time and prevent it from occurring. A deep understanding of how MongoDB works probably wasn't required to prevent this incident. Same for indexes. If you're aware of them, you can trigger the thought process of "this query is pretty slow, maybe I should look into adding an index because I know that speeds things up". You don't necessarily have to understand how a B-tree works for an index to be effective.
- dspillett 5y ago> The author says, "I wouldn’t be able to explain from the top of my head what the ACID term means". > I think there is a big difference between not knowing something off the top of your head and being completely unaware of it. The next level is being aware that the “I” doesn't quite apply in most databases with their default transaction settings, true isolation/serialisability being sacrificed for performance, though the distinction rarely matters is can be significant. > You don't necessarily have to understand how a B-tree works for an index to be effective. How a b-tree operates can be safely kept as a black-box thing, the important level of detail that many miss in my experience is how compound indexes can (and more importantly, can't) be used. Others include that in most databases defining a foreign key relationship does not create a supporting index by default, and that careless use of functions in joining & filtering clauses often blocks index use. If you have a DB specialist on the team there is probably no need for other devs to be intimately aware of even these details, but they should at least be aware enough of their existence than they know to ask the local expert (or look things up elsewhere) to check their working when the issues might affect what they are working on. As a bit of a local DB specialist who has met some bad ones, I should point out that it is important that we, and other local “exports”, be approachable so the other devs feel they can run things by us for verification/improvement.
- 5y ago
- p2t2p 5y agoYou're right but you or more precisely that interviewer is also wrong. Understanding the way your storage works is crucial, yes. But you ensure people understand it not by asking "Explain ACID to me" but by asking to design a certain system and observing them figuring out requirements and asking questions. You then ask follow up question about consistency and reliability, you probably even can ask about benefits of using transactions but I wouldn't because there's plenty of databases that can be used reliably that don't have transactions. Life is more complex than ACID, that's it. Investigate if person understands life, not that they were able to memorise bunch of magic acronyms. For the life of me I don't remember what SOLID is and I totally fail to parse any written explanation of liskov substitution principle but whenever I see code that violates those I immediately go "please don't do this, that'll cause problems in the future".
- MauranKilom 5y ago> I totally fail to parse any written explanation of liskov substitution principle Off the top of my head, "all child classes should be usable interchangeably when dealing with an interface/base class". I guess that's too imprecise to qualify as "written explanation"?
- Koshkin 5y agoThis doesn’t sound correct to me. The original statement is, Let ϕ(x) be a property provable about objects x of type T. Then ϕ(y) should be true for objects y of type S where S is a subtype of T. Here, objects of a base class are replaced with objects of derived classes.
- rfrey 5y agoI interviewed a guy once who used rho instead of phi when I asked this question. What a bozo.
- marcosdumay 5y ago> This doesn’t sound correct to me. Yet the GP says exactly the same thing as you said. He just ignored that the property belongs to the base class.
- raxxorrax 5y agoPutting althe modification resulting from an API call within a transaction doesn't do anything bY itself.
- whack 5y ago> Because of such attitude of the previous dev team my client ended up with DB integrity ruined. Previous guys somehow didn't know they should use transactions when updating/deleting stuff in the DB, because hey, it's just API you call, who cares of mambo-jumbo happening behind the curtains, right? There's a difference between knowing the internal implementation, and the pros/cons of different usage patterns. What you're talking about is having mastery over an interface. This is distinct from knowing the internal implementation. As someone who used to work as a hardware engineer, I can assure you that there's a tremendous amount of hardware internal-implementation-details that you're ignorant of. But that doesn't get in the way of you being you a great developer, as long as you spend some time understanding the interface and the recommended ways to use the hardware black-box. The same applies to what the OP is describing.
- dan-robertson 5y agoI think your comment is basically just reacting to the title or a snippet of the article and not really giving the OP due respect. For example transactions are clearly a part of the api for a SQL database so not having deep knowledge and not using transactions are not equivalent things. I think the article also touches on the fact that when you look up the definition of ACID you get a rough description of the acronym and then a description of how it is mostly a marketing term and all the ways real systems can deviate from it. So really the deep knowledge is not so much what ACID means as what it doesn’t. The article is basically complaining that knowing what ACID stands for is not really a fundamental and that knowing how to use a database is. So I don’t think you really disagree with the OP much either.
- joshspankit 5y agoCame here to address his part of the article. My statement is that it was a terrible example, but the core principle stands. Yes, we should all definitely understand how the storage layer works (and anything else that could corrupt data) but no one needs to have the laundry list of acronyms memorized.
- addicted 5y agoThe author’s point still stands. If you’re implementing interactions with a fairly new database (let’s face it. 10 years ago MongoDB was basically experimental) then it is indeed incumbent on you to understand in greater details what you’re doing. But if you’re building a CRUD web app, that’s maybe only used internally by your company and backed by SQL Server you probably have to never give it a thought. Either way, lack of knowledge of a particular acronym, one which is described as basically meaningless in texts that describe it, is not a deal breaker of any sort. And it doesn’t apply just to ACID but most acronyms in tech today, such as REST/SOLID, etc like the author points out.
- kragen 5y ago> But if you’re building a CRUD web app, that’s maybe only used internally by your company and backed by SQL Server you probably have to never give it a thought. Until you get complaints that the hour-long reconciliation transaction you run in the morning is locking out all the users, or there's an API a customer hits that's erroring out 1% of the time because Sybcoughuh, Wonderful Microsoft SQL Server has detected a deadlock, and this is causing us to lose customer orders, or... No, you really need to know what ACID means, and what tradeoffs are made to balance ACID with other considerations in your chosen database implementation. Lumping ACID in with vague heuristics like SOLID on the basis that they're both acronyms is just plain ignorance.
- throw1234651234 5y agoJust to play devil's advocates, our juniors LOVE putting everything in a transaction and locking up the client's legacy database.
- ratww 5y ago> Previous guys somehow didn't know they should use transactions when updating/deleting stuff in the DB, because hey, it's just API you call, who cares of mambo-jumbo happening behind the curtains, right? Sure, you're right, but I'd argue that transactions are is hardly "database internals". Transactions are day 2 of "using a database 101". You don't have to know what ACID means or how transactions are implemented to use them correctly. I'd rather hire a candidate that knows how to use transactions but has no idea what ACID is than one who knows what ACID means and how it's implemented but doesn't know when or why they should use transactions in real life. (Of course the second case is rare, but this is just an example).
- chousuke 5y agoI get the feeling that many people do not even know what problems databases solve. It seems to be common to treat them as an extension to the application's memory that somehow isn't lost between restarts, and the properties, operational caveats and other capabilities of the database are not even considered.
- ratww 5y agoSo true. I see a lot backend candidates these days who learned only how to use the ORM of their framework of choice, plus migrations, and that's it.
- siva7 5y agoI don't think that transactions count as deep database knowledge. Those are the absolute fundamentals that most developers need on a daily basis and not some dusty computer science stuff. If you have never heard of them it is fair to say that you should take a step back and learn the basics of your profession before being allowed to touch a production database. But the thing is of course - you don't know what you don't know.
- gpderetta 5y agoGreat point: > you don't know what you don't know aka knowledge is what you have after you forgot everything you learned.
- notTheAuth 5y agoThe author covers your perspective: > But, that’s only because you are a noob and you don’t deal with the things that I have deal with. > Someone on the internet In conclusion: not all software problems require the same implementation. I don’t run a DB for some deterministic applications I have that can regenerate the same outputs such I don’t need to store it all. I could, in theory, regenerate an entire data set with a seed. Maybe app developers are ignorant and lazy for not designing systems that do that! What does your one story have to do with “software engineering” as a field? Sounds like people are conflating problem solving and personal career demands.
- cmiles74 5y agoWhen I read it, my first thought was that "an API" in this case was SQL. Understanding why we should use database transactions is not the same, to my mind, as understanding database internals. When the author said they can treat the database as a "black box", my thinking was the they didn't necessarily need knowledge of the database's implementation. People can understand how to effectively work with a database in the general sense (i.e. use transactions) without a deep knowledge of every implementation. When we start talking about MongoDB, I think that's really interesting because it is so different from our traditional, SQL-first databases like PostgreSQL, Oracle, etc. Our understanding of SQL-first databases isn't all that useful when we try to apply it to MongoDB. At that point I think we are in agreement: it's unreasonable to think we can treat a MongoDB database as a black box if our base knowledge is only SQL and SQL-first databases. We probably do need someone who has already dug in and dealt with the implementation, how MongoDB really gets things done.
- lmilcin 5y agoOne of projects I was called to help some time ago set up a service with over a thousand servers. They reached out to help improve the performance because in their opinion the network and network filesystem were too slow. They invested incredible amount of time in learning technology and scaling their application but forgot about need to learn fundamentals -- data structures and efficiency. Fast forward 1,5 years, the application as I left it ran on a single server using about 10% of its capacity. A second server is just a hot standby backup. The service went literally from being able to process tens of transactions to a hundred thousand transactions per second on a single node. More than that, we threw away most of the exotic technology that was used there -- greatly improving team productivity. The implementation is simpler than ever with layers upon layers of microservices replaced with regular method calls and a lot of infrastructure basically removed without a need to replace it with anything. People who do not learn from history (fundamentals) are doomed to repeat the same mistakes.
- benlivengood 5y agoSomething people seem to forget when designing systems based on what they've learned is that modern individual machines have approximately the power of the top supercomputer ~20 years ago, and a few racks of modern machines can in many ways match the top supercomputer of ~10 years ago. Approaches to solve large problems are continually changing and it's almost always worth a big-picture design session when looking at a familiar-seeming problem. By the time students graduate from a 4-year program computing is about 10X better than when they started. I still remember some interviews within the last decade where single-core machines with a few GB of RAM were a default assumption for whiteboard designs, or that spinning disks were the default.
- lmilcin 5y agoWell, I catch myself that I am not "updating" the state of my understanding of hardware. Just recently I caught myself putting a lot of effort into solving a problem that only exists if persistent storage is too slow to be used for calculations. Then I facepalmed myself hard when I realized that I can just move that entire 500GB data structure to an NVMe and treat it almost as if it was in memory. But in general I think that the problem isn't that people are not "updating" their understanding. Even 10 years ago it wasn't a huge problem getting hundreds of thousands of transactions per second on a single modest machine. The problem rather is people relying on more and more layers of abstractions for vary small gains. Example: I get that Python is a nice language (for somebody that does not know Lisp). But is it worth it to choose Python for a little bit improvement in productivity for a problem that requires a lot of throughput, to then suffer performance issues, to then spend many times more effort on trying to improve performance? I don't think so. Or more in my space: is Spring Data (Java) worth the very incremental productivity improvements if it completely destroys your application performance? The application I described in my parent post used Spring Data MongoDB which kinda means it was fetching entities one by one which is extremely costly. By replacing it with bare MongoDB reactive driver and ensuring data is being streamed in large batches (why read one user data if you can read 10k at a time) and getting rid of costly aggregations in favour of application side processing we have improved throughput by many orders of magnitude WHILE reducing load on MongoDB. Granted, there is a little bit of additional complexity (on the order of 10% more of application code) but just the performance improvements mean that the team can breathe and focus on other problems like modelling the domain correctly.
- cultofmetatron 5y ago> Previous guys somehow didn't know they should use transactions when updating/deleting stuff in the DB, because hey, it's just API you call, who cares of mambo-jumbo happening behind the curtains, right? I would consider transactions part of the api. I use them all the time in postgres and I'd hardly call myself an expert in the low level details
- tootie 5y agoI don't think use of transactions or any other query or modeling hygiene are "internals". You can just as easily consider SQL to be an API (or really a DSL) for the black box of DB internals. In most cases, you're better off not looking inside the box because DB vendors only guarantee behavior at the API level. You can always identify and improve bottlenecks if/when they actually impact application function. I can write a load of crummy queries, expose them in an API and then just apply some clever caching and the DB will just never be a bottleneck no matter how little I understand it.