4 ms·
The C in ACID is consistent, not serializable (i.e. it's not ASID). Consistency means that declared consistency constraints are enforce -- no dirty writes, no
by jimstarkey 15y ago
The C in ACID is consistent, not serializable (i.e. it's not ASID). Consistency means that declared consistency constraints are enforce -- no dirty writes, no violations of unique indexes, referential integrity, and anything else that the bells and whistles supports.
Let me give a simpler example: Database with one table of one field, a number. One transaction: count the number of records and store that number.
A serializable system will force one zero, one one, one two, etc. But a consistent system can have two zeros and no ones. Why? Because that's what each concurrent transaction saw? Nothing wrong with that. But if application semantics dictate that each value must be distinct, then put a unique index on the number and the system will enforce uniqueness. Automatically enforcing "auto-magic" constraints that nobody cares about is why serializability destroys scalability of distributed system.
In NuoDB all messaging is asynchronous and batched, making it very fast and efficient.
For more than you want to know, see http://www.gbcacm.org/sites/www.gbcacm.org/files/slides/SpecialRelativity[1]_0.pdf http://www.gbcacm.org/sites/www.gbcacm.org/files/slides/Spec...
Someplace there's even an audio recording which I recommend if you're into self-abuse.
- jhugg 15y agoSo running with your example, assume I have a transaction that adds 5 to a column value and then reads the value back. If I start with a value of 0, then run my transaction twice, can both transactions return 5? Or is it guaranteed that one will return 10?
- deleted 15y ago[deleted]
- andrewcooke 15y ago[sorry. replied earlier incorrectly]. one of the two transactions would abort in this case. snapshot isolation needs to check that any data mutated in the transaction were not also mutated externally. in the example you are replying to, there is no mutation, and so no problem, but in your example there is. see http://en.wikipedia.org/wiki/Snapshot_isolation http://en.wikipedia.org/wiki/Snapshot_isolation note that it's only mutated values that are checked for conflicts, and only against other mutations (this is why it is efficient - the number of checks required is small). so you can get weird behaviour when multiple values are read while different transactions change each - there's a good example in the link above. this is called "write skew". the whole approach is, in a sense, exploiting poor phrasing of the ansi sql-92 standard, which doesn't actually require serialisation even though that is the most natural way to interpret it (as far as i understand things). so you can think of MVCC as "exploiting a loophole" that leads to a more efficient system, but one that is less intuitive. on the other hand, this is not new - it's already the standard behaviour for postgres, oracle, sql server, etc.
- julochrobak 15y agoThanks for the clarification. However, this does not sound like something I could use in practice. If I have two transactions and one of them is aborted because of the changes done by the other transaction what am I as a developer supposed to do? Retry? I hope not, because a retry is in other words serializing the execution. One after another. So, if I have a system with a lot of concurrency (i.e. bank accounts and transfers) I'd better not use NuoDB because I'd get a lot of transfers aborted, not good. If I have a system with a very little write conflicts I'd go for serializability because that gives 100% consistency and will be fast anyway due to very little conflicts. From the wikipedia link you posted, it is fairly clear that the Snapshot isolation is good when you don't need consistency. For that you'd have to either abort every time there is a conflict or introduce write-write conflict (ie. serializing).
- andrewcooke 15y ago(1) serializability will not be "fast anyway". it will be slow. this is the problem with, for example, mongodb's global write lock. (2) you would not get "lots" of aborts with bank accounts because each transaction is, typically, to a different account. you're only going to get a problem when two processes try to change the same person's account at the same time. (3) what this provides is standard-compliant ACID sql, the same as postgres, oracle, sql server, etc. if you use any of those and don't have retry in your code for when transactions fail then you're already in a mess. i am not associated with this project, but what you're saying doesn't really make sense. as far as i can see, you're criticising it for being the same as everyone else in the standards compliant, sql world.
- julochrobak 15y agoI'm not criticising it for being the same as everyone else. I'm just not sure how the whole thing works. I don't see enough information provided on the site, yet a lot of statements about providing full ACID and scalability. re 1) and 2). I think they are related. The serialization can be done on account level. There is no need to have a single global write lock. This also implies that if the bank account transfers are typicaly to different accounts the serialization will be as often as the aborts. Hence very rare and the system will perform equally good in both cases. However, the serialization does not require further retries and doesn't force the application programmers to workaround the problems.
- julochrobak 15y agoI'm lost here as well. Let's imagine I have COUNTER table of one numeric field VAL. All I do in transactions is: update COUNTER set VAL = VAL + 1 Let's assume the initial state is one entry with VAL = 0. If I run the update statement in two concurrent transactions I'd expect the result to be 2 when both transactions complete. However, I don't understand how that can be achieved without serializing those transactions.
- jimstarkey 15y agoThe system recognizes that two concurrent transactions are trying to update the same record. If the first commits, the second will return an update conflict. The application can then restart the transaction. These are the same MVCC semantics I used in Rdb/ELN (1984), Interbase (1986), Firebird (1999), and MySQL Falcon (2006). The implementation, however, is wildly different.
- julochrobak 15y agothank you for the explanation. I'm really sorry but I don't agree that this is the way going forward. You claim that serializability is not needed to achieve consistency. I agree up to the point that it's only possible if the conflicts are recognized and transactions aborted. What is the benefit here? As an application developer I need to restart the transaction and hope it's not going to fail again - hence I'm serializing the execution on my own... On top of all this, when it comes to constraints in relational model they can be complex and the probability of failing transactions is just going to grow. I can see this working only if I start relaxing on my constraints and redirecting transactions in such a way that conflicting transactions are coming to the database already serialized. What about restarting transactions within NuoDB transparently as soon as an update conflict has occurred?
- bsmorris 15y agoWe'd be happy to address these issues if you are interested. It's not clear where we're misunderstanding each other, but for clarity NuoDB manages such things as atomic operations and update conflicts without reference to the application. It's an ACID database. Give us a call or attend a webinar (eg tomorrow) if you want to take a deeper look at it. Barry Morris, NuoDB
- deleted 15y ago[deleted]
- deleted 15y ago[deleted]