6 ms·
> Most database servers will be able to retrieve rows from well-indexed tables at far greater rates than low-performance application platforms' ORMs can transla
by nrser 13y ago
> Most database servers will be able to retrieve rows from well-indexed tables at far greater rates than low-performance application platforms' ORMs can translate those rows into usable objects. Modern database servers fetching rows from well-indexed tables can keep up with the query demands of the very highest-performance frameworks without saturating a database server's CPUs (with throughput measured in the tens to hundreds of thousands of queries per second per server).
could you be a little more specific? is this from real world experience or from some benchmarketing material somewhere?
if i recall from early Zynga (circa 2008), we would get MySQL (InnoDB) on decent commodity hardware (Dell 2U with 8 cores, 32G mem, 15K RPM disks in raid 10, cost about $7K USD) to around 6K QPS before it would eat shit.
that was with correct indexes and only consisting of simple selects and updates (no joins or other complex queries. we probably only had ~10M users in the database, but the `user_item` table at least was likely well over 100M rows).
we even used to have a saying: "first, check the database. if it isn't the database, it's probably the database." when the app crashed, it was almost, almost always the database.
this wasn't at the 'monster' scale that the games later grew to, just a handful of php scripts with a few million monthly actives, a scale that i think a lot of people hope to achieve with modern consumer apps.
i'm just saying that i've lived the daily "it's the database", and i agree that it's a major concern for people building on open-source stacks and commodity hardware for a wide audience. the honest truth is that everything else in the stack is really, really easy to scale in comparison.
- jacques_chester 13y agoThe obvious question is: did you try other RDBMSes?
- nrser 13y agono, that's sort of why i was asking for details. I'd be interested to know if postgres or some other open source solution could net us the 10-100x performance he's referencing. I've heard great things about postgres, but I've never used it under heavy load. I'm under the impression that some proprietary databases on expensive hardware can achieve those levels, and some friends that do oracle installations claim that the total cost-per-query is competitive with open source / commodity hardware setups, but I've never really considered using them at my companies, and it seems like a lot of start-ups are in the same boat.
- mattrobenolt 13y agoFWIW, PostgreSQL is our primary data store. It works well. ;)
- nrser 13y agolike i said, i've heard good things. what kind of QPS have you guys seen a box go up to? anywhere near the 'tens to hundreds of thousands' that [bhauer](https://news.ycombinator.com/user?id=bhauer https://news.ycombinator.com/user?id=bhauer) threw out there? 45K req/sec is big, btw. props, yo.
- jacques_chester 13y agoYeah, the proprietary databases can be very expensive. I was told at an IBM internship day that DB2 was doing billions of transactions / day for UPS in the mid 1990s ... on clustered[1] mainframes. Not something most startups can easily get their hands on. Basically the technological bifurcation seems to exist near the limits of a high-value corporate credit card. Which naturally changes over time. The biggest shift in data economics in the last 5 years has been the emergence of SSDs. For startups the 2nd biggest shift has been the surged in performance of PostgreSQL; they've made huge strides in scaling it up over the last few years. [1] http://en.wikipedia.org/wiki/IBM_Parallel_Sysplex http://en.wikipedia.org/wiki/IBM_Parallel_Sysplex
- asdasf 13y ago>if i recall from early Zynga (circa 2008), we would get MySQL (InnoDB) on decent commodity hardware (Dell 2U with 8 cores, 32G mem, 15K RPM disks in raid 10, cost about $7K USD) to around 6K QPS before it would eat shit. Mysql is slow for concurrent access, but not that slow. I don't have mysql installed, but postgresql running on openbsd (a famously poorly performing OS) in a virtualbox virtual machine (hooray overhead) only being given one core and running off of a slow laptop hardrive is beating that for me.
- nrser 13y agohow big is your data set? i remember MySQL degrading when the tables got too big, i think around 100M rows or so. the data set probably didn't fit in memory either. i don't know how much of a difference it makes, but we were also moving all the data across the wire (vs. what i'm guessing is local access on your machine), and probably about 50% writes. i might be off on the 6K number... this was all some years ago, so it's a little fuzzy... but i don't think it was 'tens or hundreds of thousands' by any means.
- bhauer 13y agoAdmittedly, a queries per second measurement comes with a bunch of variables. Reads? Writes? Fetching by primary key, an indexed field, an unindexed field? Joining? Does the data set being accessed fit in the memory of the database server? In our benchmarks [1], the hot data set does fit in memory, rows are small (by design, to avoid saturating the Gigabit connection) and they are fetched by primary key. The results top out at ~145,000 queries per second (cpoll-cppsp shows 7,252 requests per second and this test exercises 20 queries per request). The database server's resources are not fully utilized by this test. That is, in this test, the web framework--even this blisteringly fast C++ framework with its optimized MySQL driver--is the bottleneck. In the updates test [2], the same framework scores 1,963 requests per second and this test exercises 20 reads and 20 writes per request for a total of 39,260 select queries and 39,260 update queries per second. Outside of the database server using a Samsung 840 Pro SSD, our benchmarks run on early-2011 vintage i7-2600K Sandy Bridge workstations. This benchmark is not a database benchmark, however. We capture numbers with MySQL, Postgres, and Mongo but it is not our intent to measure the performance of databases. Another benchmark is Databench [3], which is measuring the impact of ORMs on bank-transfer-like transactions. Again, not a benchmark focused specifically on measuring the database itself, but interesting nevertheless. On an hi1.4xlarge Amazon instance using the Prevayler library, that hardware achieves 84,860 transactions per second. Databench uses Postgres. [1] http://www.techempower.com/benchmarks/#section=data-r6&hw=i7&test=query http://www.techempower.com/benchmarks/#section=data-r6&hw=i7... [2] http://www.techempower.com/benchmarks/#section=data-r6&hw=i7&test=update http://www.techempower.com/benchmarks/#section=data-r6&hw=i7... [3] http://databen.ch/ http://databen.ch/
- bhauer 13y agoCorrection: Now that I'm at home, I took a closer look at Prevayler since I wasn't familiar. It appears to be a persistence platform in its own right rather than an interface to the database. So disregard. I'm not precisely sure how to read the DataBench results, but presumably the rows that cite Postgres in the name are more relevant.