5 ms·
I've seen a lot of "Cloud" database services lately, and while I love the general idea, from what I've seen in real life, for any serious db use, the impact of
by modoc 15y ago
I've seen a lot of "Cloud" database services lately, and while I love the general idea, from what I've seen in real life, for any serious db use, the impact of network latency is a HUGE problem. So unless this service happens to be residing in the same data center as your application it doesn't seem like a viable option.
The difference between a local DB with <2 ms latency to a remote DB with ~50 ms latency is huge. I've seen application start times go from 1-2 minutes to 20-30 minutes by pointing at a remote DB.
So if your app is even somewhat DB intensive, you really need <2 ms latency to your DB, and if it's not DB intensive, you probably don't need all the scaling and other infrastructure strengths these services bring to the table.
- michaelbuckbee 15y agoI think you are correct. I'd guess that the major target of this would be applications already running on AWS (which would put them mostly in the same datacenter). Obviously, a bunch of caveats about availability zones, etc.
- joshu 15y agoI think this is enormously exacerbated by roundtripping. Even if it's just 2ms away, and you go back and forth a few dozen times, you are doing a ton of waiting. This is one of the problem with joinless NoSQL databases, as well, since you HAVE to roundtrip to join.
- randomdata 15y agoYou don't have to. For instance, in CouchDB, you can use view collation to retrieve related documents in a single request.
- joshu 15y agoThat sounds like a join by another name :)
- randomdata 15y agoIf you use the term loosely, I suppose so. :) But you're not actually joining any data. You return one document, and then exploit the properties of sorting to return the related document(s) next in the collection. In SQL-speak, each relation comes as a distinct row and are only connected by the order in which the rows are returned. There is a writeup on the specifics here, if you are curious: http://wiki.apache.org/couchdb/View_collation http://wiki.apache.org/couchdb/View_collation It is a pretty clever trick that can save you a round trip to the DB, but doesn't have nearly the power that a real join does.
- einhverfr 15y agoyou can do the same in PostgreSQL with joins. For example: select i.* , array_agg(ROW(l.*)::text) from invoice i join invoice_line l ON i.id = l.invoice_id where i.id = 1023; This returns a single row for invoice, along with an array of tuples representation of all invoice lines. You could also use casts and stored procedure languages to convert to xml or whatever if you want.
- misterbwong 15y agoI tend to agree with this sentiment, though for some read heavy applications you can mitigate this problem via a distributed cache layer in your web tier.
- bgentry 15y agoGiven that these databases run in AWS, there's no reason you can't run your application in the same datacenter.
- milkshakes 15y agowell, except that AWS doesn't let you coordinate AZ's between customers [1] [1] http://aws.amazon.com/ec2/faqs/#How_can_I_make_sure_that_I_am_in_the_same_Availability_Zone_as_another_developer http://aws.amazon.com/ec2/faqs/#How_can_I_make_sure_that_I_a...
- ceejayoz 15y agoI get 1 ms or lower latency between AZs.
- stock_toaster 15y agoI get about 1.5ms to 1.6ms ping latency between AZs and about .3ms to .4ms between nodes in the same AZ (west-1).
- cperciva 15y agoMaybe not officially, but I believe "is in the same AZ" is an equivalence relation, so once you've figured out the appropriate permutation function...
- aurynn 15y agoWell, Postgres can shine here if you can do a fair amount of processing inside of a stored procedure. The sproc takes care of lifting, the client can be (fairly) limited in the transformations it applies, and the network latency of repeatedly going to the DB is lessened. This is, naturally, very application-dependent.
- einhverfr 15y agoIndeed, you can do a fair amount of processing without even hitting stored procs in Pgsql. Plain SQL gives you tremendous power in Pg. I find stored procedures add useful semantic sugar around these, but especially with newer versions things like common table expressions, arrays of tuples, and the like, gives you tremendous power.
- eftpotrm 15y agoNot just Postgres; any database server in which you can do either stored procedures or submit arbitrary batches of SQL statements rather than just individual queries will let you do this. I routinely do in MS SQL Server, for example, and have worked on multiple projects where the ability to do this was utterly critical to overall performance. To be perfectly honest, I'd consider any database server where I couldn't do this to be something of a toy because of the restrictions it places on overall app performance. It might be OK in SQLite or Access but a real database? Sorry, no, come back when you've finished the thing please.
- eftpotrm 15y agoPardon me, but if an app's operations time will go from 1-2 minutes to 20-30 by changing from a local to remote DB then you've either got an utterly abysmal network or enormously too many round trips to the database and the app could benefit hugely from being rewritten to move critical functions to running in SQL directly on the database server. Now, I'd still rather have a local database but if you do that properly then almost all pages in your typical webapp shouldn't require more than a single round trip. IMHO.
- modoc 15y agoFWIW: Remote DB was ~50ms away. Large scale eCommerce applications tend to be DB heavy-ish. Have a complex catalog structure with a few million SKUs, add in a few tens of millions each of users, orders, coupons, and push thousands of orders an hour, and you have a LOT of DB traffic. Add in things like real time inventory checks on product pages, dynamic shipping option/cost calculations based on the user's address (if logged in) or IP address (if not logged in), determining applicable coupons, cross-sell, up-sell, CTAs, dynamically based on the user's history and profile data, etc... You make a lot of db calls. Even caching stuff on the app tier you have to load that stuff in from the db at least once, etc... Again this is all dependent on your application and needs, but saying you should rewrite it to move more logic into the DB isn't really useful in many scenarios. If you need things like dynamic cloud scaling, read only slave replication, etc... chances are you're doing a lot of DB transactions as well, and that latency can and will kill you.
- einhverfr 15y agoStill, you are talking about round-trip overhead. As a counter-example, LedgerSMB is very databse-intensive but we do what we can to make sure round-trips are minimized. This means making sure that everything that needs to be queried together is queried together, in the same query. I suppose if you are doing a lot with ORMs and the like though that may not be an option.
- eftpotrm 15y agoNo reason you can't do an inventory check at the same time as returning the basic catalogue details, it won't be more accurate in a separate call 50ms later. If you're generating the page for a logged in user then you know who that logged in user is and what they're asking for, so you can calculate coupons and the like in the same call too without needing to resend the same parameters to the database to look up, and if not feed in a different parameter for their IP address. Now, I'm sure there's all sorts of other things in apps beyond my experience where you're better off doing multiple calls to the database, but certainly nothing you've outlined in those two couldn't be handled by a decent database infrastructure in a single call returning (potentially) multiple results sets. One thing I learnt many years ago in data intensive apps. Communications latency between your app and your database will kill performance if you let it. Shipping data back and forth repeatedly will hang, draw and quarter it. Absolutely, aggressively, pare the number of external calls you have to make to the bone and you'll see a very significant performance boost.