4 ms·
What do you find so rediculously good about this? I find it complicated. If I understood it properly, it's just making inserts asynchrounous (removing durabilit
by julochrobak 15y ago
What do you find so rediculously good about this? I find it complicated. If I understood it properly, it's just making inserts asynchrounous (removing durability) and batches data into single statement (removing atomicity). Or did I miss something?
- willvarfar 15y agoIts still ACID. Its the same callee code as you use in, for example, twisted already to have DB async. It just gets way better performance because the DB speed is measured in requests/sec and not how big those requests are. So merging many requests into a single one where possible means you get more throughput.
- marshray 15y agoBut it's not ACID. For example, if the server or process running his database worker queue were to crash, records would be lost. Of course, it's quite likely an acceptable trade-off for for these records. Relational DBs sold everyone on ACID in the 90's, but a lot of data doesn't really deserve its cost. He hasn't found a way around Brewer's Conjecture (and I wish I had a nickel for every time it's rediscovered :-).
- willvarfar 15y agoBut then the callback to the request would timeout and the callee code waiting on the web request would be in no misunderstanding that the thing was committed or not. Things in the queue get acknowledged - and sometimes resultsets or IDs returned, in the case of SELECT and INSERTs that return auto keys.
- marshray 15y agoOK, I see. So your batching related inserts to cut down on the per-transaction overhead, but not actually returning anything to the client until the commit is confirmed. I think this is a good method, as long as you're careful about these things. (I assume you've taken care of them, just talking about this method in general). 1) If there were multiple transactions in the web request, previously the second ones likely wouldn't have been run if the first were to fail. This would likely change that and it could even necessitate having rollback for the second transaction. 2) Now there is a potential "thundering herd" problem. When the transaction batch commit notification is received, all 1000 web requests made over that 100 ms period become ready to execute and start competing for network/disk/database resources simultaneously. A disk that was just sitting idle now has 1000 different things to do at once. If everybody hits the network at once it can cause temporary packet loss. When packets are lost, connections can be dropped or delayed by retransmits. Sometimes client retries and retransmits can happen synchronously. In the worst case, such a system can end up oscillating wildly between excessive load and excessive idle.
- willvarfar 15y agoYes this is what twisted's "interactions" are and how I talk about throwing rollbacking-exceptions and such in the article. Yes, you do have to be careful about saturating links generally. The thundering herd thing is a rich man's problem :)
- julochrobak 15y agoI'm sure the statements to the database comply to ACID properties. Nevertheless, from the client perspective the batching simply means that request from client one is batched together with request of client two - i.e. the atomicity is broken. If the operation fails due to problems with data from client one it impacts client two. I guess my point is, why use a database which comply to all ACID properties if you don't need them.
- willvarfar 15y agoYou missed something; the only part of ACID its relaxing is isolation: http://en.wikipedia.org/wiki/ACID#Isolation http://en.wikipedia.org/wiki/ACID#Isolation "This property of ACID is often partly relaxed due to the huge speed decrease this type of concurrency management entails." Indeed. If you have the viewpoint that DB validation is for catching programming errors, not invalid data (i.e. you validate everything before it hits DB anyway), this relaxation can be reasoned about client-side anyway.