4 ms·
I think there are two different kinds of consistency, and it's important to not conflate them. There's consistency that's internal to a system. Do all of the f
by ddulaney 4y ago
I think there are two different kinds of consistency, and it's important to not conflate them.
There's consistency that's internal to a system. Do all of the foreign keys line up correctly? Have I lost any data that was provided to me? Here, we can aspire to be 100% correct. I don't think the examples in this article conflict with that.
Then there's consistency that's external to a system. This can be between this system and other systems, or between this system and reality. Did the operator enter the correct data for this item? Did the external system change and fail to tell me? Here there is no way to be 100% correct using only the tools that are inside the system. You need external audits, reality checks, periodic reconciliation.
Critically, all of the author's examples are about external consistency (accounting matching reality; inter-system communication), but their conclusion seems to be that because external consistency can't be fully achieved, we should be OK to abandon internal consistency within single systems. I think that's too strong a conclusion.
- taeric 4y agoI think I agree with you. That said, internal and external are a touch inadequate. Specifically, for a large enough system, internal consistency will look more like external from a smaller system's perspective. To that end, it is all about costs. If the cost of keeping consistent is not above the budget, do so.
- gfody 4y agoincluding unknowable cost/opportunity risk
- GauntletWizard 4y agoThere's a ton of places where foreign keys are used wrong - deletion is acceptable, and good error handling for that case is the right thing to build anyway. That's one of the key arguments that nonconsistency advocates are arguing for. If you build a music playlist, the behavior of the program if the mp3 files referenced within should be to skip that track, not to crash. I've heard too many people arguing that they should be allowed to emit nasal demons if there's a data error, and they're wrong even in many well structured and tightly integrated datasets, but especially wrong in "web-scale" datasets
- bcrosby95 4y agoDeletion is acceptable, and if you have everything in a consistent system they will both exist or not. "Nonconsistency advocacy" doesn't make a lot of sense to me here. Are you advocating not relying on consistency in systems that guarantee consistency? That is a waste of time. Are you advocating eschewing consistent systems? Well, then you have more work, so only if I need to. And yes, you should handle data errors here because inconsistency is consistent with the system you've chosen.
- GauntletWizard 4y agoMy point is: Removing a music track from the list of playable tracks is not a good reason to remove it from playlists. Having dangling pointers there, and handling them appropriately is the right thing to do. This is true in a ton of cases where neither soft-delete or cascading delete make sense. Rarely, if ever, have I actually encountered the latter.
- ddulaney 4y agoHmm... That's an interesting example. It definitely makes sense to enable features like that. However, I'm not sure it's a place where you want to abandon foreign-key consistency. To display the song on the playlist (greyed-out, Spotify-style) you need its name and info, and you probably want to keep artist and album links valid as well. It sounds like this is actually a perfect use-case for a soft-delete: keep the song metadata in the songs table and mark it as nonplayable. This keeps the artist page, album page, and any links to the song itself valid, just greyed out when you get there.
- stevesimmons 4y agoThere's a great book "Data and Reality" that discusses these subtle but very crucial differences. Discussed here in HN a year ago: https://news.ycombinator.com/item?id=30251747 https://news.ycombinator.com/item?id=30251747
- ddulaney 4y agoThank you so much for reminding me! I remember Hillel Wayne’s post about it from a while ago, and I had only gotten through the first couple of chapters before stuff got in the way. It might be a good long weekend read. Great suggestion!
- josephg 4y agoHm, I'd cut the cake a little differently. I think there's two kinds of consistency we can strive for: - Strict consistency. Every pointer in my b-tree must point to a b-tree node, and not random data which could cause the program to crash. In a financial world, a bank should never print money. - Fuzzy "good enough" consistency. In the examples in the article, all of the financial transactions should end up close enough to being reconciled. There's value in both kinds of consistency. When talking about data entered by humans into a database, there's always going to be a bit of slop involved. Someone mistyped a digit. A few records weren't entered at all. But when building software systems, designing with strict consistency guarantees is such an unbelievably massive win. I can't overstate how important it is. The entire ladder of abstraction in modern computers from transistors all the way up to this website is only possible because each layer "underneath" the layer we're standing on is solid and deterministic. If CPUs made even 1 error in every billion operations, our computers wouldn't boot at all. Tony Hoare talks about his invention of null pointers as his "billion dollar mistake". Null pointers take something that should be strictly consistent (references) and make it fuzzy. Rust is exciting lots of people in the systems programming space because it takes things that are fuzzy in C (aliasing, memory management, thread-safe variables, etc) and makes them strictly consistent. All the value in unit testing comes from how they make our systems more strictly correct. Every time I've relaxed consistency guarantees internally in systems I've worked on (or heard coworkers doing the same) we've come to regret it. At a startup several years ago, we needed to build an external search index for a database. The database updated live - and we had a change feed that updated the browsers live as records changed. The engineer in charge did a "good enough" job - he wrote a scrappy script that was only mostly correct. But it sometimes left the index inconsistent with the data. We got constant reports from our users about items not showing up in the search results. He would dutifully go back and fiddle with things to try and fix the problem. Eventually one of our senior engineers went in and rewrote the whole indexing script to be strictly correct. We never heard a peep about it after that - it just worked, every time. Even putting aside the frustration of our users, writing it correctly was a big win for us in terms of maintenance. Once it was correct, we didn't need to keep pulling engineering time away to fix problems. Maybe data consistency is overrated in databases. But I think if anything, consistency internally in computing systems is underrated. We take for granted how well computers work. But our capacity to make computers do anything depends entirely on those consistency guarantees. It seems ridiculous to disregard its importance.
- sebk 4y agoProf. Daniel Abadi talks about the two meanings of consistency in this blogpost: https://dbmsmusings.blogspot.com/2019/07/overview-of-consistency-levels-in.html https://dbmsmusings.blogspot.com/2019/07/overview-of-consist... I'd argue that in your second case, there's no consistency "external" to the system; you've now defined a larger, distributed system that includes your external parties, and now you're in consistency-as-in-CAP territory, rather than consistency-as-in-ACID.
- klabb3 4y agoYeah I tend to agree. Another way of looking at it is that consistency helps curb complexity, and complexity is our biggest nemesis. If breaking consistency is necessary, so be it. Sometimes it’s the only way forward, especially so for horizontal scalability. But at the same time, if we don’t have consistency within small units, we get circumstantial complexity for no good benefit. As usual, balancing trade offs is better than following dogma, and within our current technological paradigm I don’t see any alternative but to minimize inconsistency for large distributed systems. Perhaps in the future we find something that eliminates a bunch of these problems altogether, but that’s not today.
- m463 4y agoSort of reminds me of memory consistency and gpus. In the beginning, there was less need for it - suuuper fast memory, but you might have a bit error on a pixel glitch for one frame and ... no big deal. Then GPUs started to be used for compute and... well, we might not only need good memory, we might need ECC...