17 ms·
AWS Outperforms GCP in the 2018 Cloud Report
- verdverm 8y agoIs there anything open source so I can reproduce these results?
- awoods187 8y agoAll of the benchmarks we used to test are open source. TPC-C https://www.cockroachlabs.com/docs/stable/performance-benchmarking-with-tpc-c.html https://www.cockroachlabs.com/docs/stable/performance-benchm... Sysbench https://github.com/akopytov/sysbench https://github.com/akopytov/sysbench Stress-ng https://kernel.ubuntu.com/~cking/stress-ng/ https://kernel.ubuntu.com/~cking/stress-ng/ iPerf https://github.com/esnet/iperf https://github.com/esnet/iperf PING https://linux.die.net/man/8/ping https://linux.die.net/man/8/ping
- verdverm 8y agoLooking for the experimental settings as it would be difficult to reproduce without detail. Could you post the scripts to GitHub?
- planckscnst 8y agoThe report states they are using iperf, ping, stress-ng, and sysbench; these are all open-source. There might not be enough details to reproduce exactly, but I think there is enough there that you should be able to produce similar results.
- planckscnst 8y ago> At first glance it appears that GCP has a tighter latency spread (when compared to the network throughput) centered on 0.2 ms Comparing a distribution of two completely different metrics is not particularly meaningful. You can change the histogram buckets to a different size and it will look just as spread out. When I first read it, I thought it was saying GCP has a tighter spread than AWS. This was confusing, especially since the chart immediately above that seems to have AWS and GCP numbers flipped.
- gamegoblin 8y agoI can't find the the text you quoted in the article. Did they edit it out?
- planckscnst 8y agoYes, it looks like the text was edited, but the table is still incorrect.
- orangechairs 8y ago(editor here) Table was corrected last night, and the text was edited per the comment above and a few keen eyes internally who also caught the confusing sentence.
- planckscnst 8y agoI wonder if they will consider expanding the cloud report to things that are not necessarily relevant to Cockroach labs use-case. In my work, latency to the block store (especially outliers) is very relevant. It would be great to see latency distributions of various workloads (sequential/random and read/write/mixed) to the patform's distributed block store. edit: I missed that 95th percentile latency for read and write was included. That is helpful; I also typically look at p99.99 and p100.
- awoods187 8y agoWe can consider expanding this in our testing next year! Glad you found it helpful
- londons_explore 8y agoYou can tradeoff talk latency vs cost yourself easily. For example, you could have two copies of your data in the block store, and issue reads to both simultaneously, and use whichever returns first. Suddenly, your 99% latency becomes your 99.99% latency... You can do the same for writes (albeit a bit more complex).
- planckscnst 8y agoYes, the cost/latency trade-off is exactly why this can be interesting. Significantly fewer or less-impactful outliers can save a lot of money (or allow a better SLA with the same cost).
- kyrra 8y agoIs this table labeled wrong? https://d33wubrfki0l68.cloudfront.net/f61cd6683f5c13f8d2b506a8cc45c8e9e8e70430/2d211/uploads/2018/12/aws_v_gcp_network-latency-table.png https://d33wubrfki0l68.cloudfront.net/f61cd6683f5c13f8d2b506... All the text around it says GCP is worse, but the table shows GCP is better.
- elmo1788 8y agoVery promising results!Look forward to seeing what’s in store for 2019.
- elmo1788 8y agoVery promising results. Look forward to seeing what’s in store for 2019!
- brian_cunnie 8y agoI've also benchmarked GCP vs. AWS [0], and, for the tests that I ran, found that GCP outperformed AWS by a factor of 3:1. Specifically, a GCP instance n1-highcpu-8 with a 256GB pd-ssd disk, clocked in at 11,728 IOPS vs an AWS c4.xlarge with a 256GB gp2 disk, clocking in at 3,634 IOPS. To put that in context of the blog post, it means your setup can drastically affect your results. Using local NVMe disk, for example, yields excellent results at the expense of increased risk. Also, AWS's io1 disk is very expensive—after my first io1 bill from AWS, I never used that disk type again. [0] http://engineering.pivotal.io/post/gobonniego_results/ http://engineering.pivotal.io/post/gobonniego_results/
- deleted 8y ago[deleted]
- virtuallynathan 8y agoI would compare against the C5 instance type; it uses the newer "Nitro" hypervisor.
- peferron 8y agoJust looking at the report's second graph, c5d.4xlarge has twice the throughput of i3.4xlarge, which is incredible. I wish AWS would release a new generation of "i" instances, with the same large amount of local storage as i3, but built on the same platform (including Nitro and processor choices) as c5/m5.
- malisper 8y agoTo be fair, an i3.4xl has 10x the disk and 4x the memory as a c5d.4xl at 1.5x the cost. If you are optimizing for GB/cost, i3s are still a reasonable choice. There's also the i3.metal which has no hypervisor. I'm curious how the performance of one of those compares to a traditional i3.
- derefr 8y ago> I wish AWS would release a new generation of "i" instances, with the same large amount of local storage as i3 I figure they can't. [Economically, I mean.] I always figured the "local per-instance storage" on most instance-types was actually not literally local, but rather a set of disks allocated from a per-rack iSCSI disk server. (The host "wastage" if this wasn't true would be quite large.) The "i" instance-type hosts—especially the ones for the large instance-types—were likely just making a claim for the entire disk-server of their rack (leaving the rest of the rack to be schedulable only by instances that need no instance disks.) Since iSCSI and "bare metal" don't go together, on Nitro, you'd actually need to build instances with real physical reserved disk pools that the hardware can see. That may even mean putting the disks inside the computer (shock horror!) and thus needing complex 2Us that need to be "recycled" (= having the still-good stuff fished out of them when they die) and may need to be opened up by ops folks for more than one reason, rather than just having "throwaway" compute blades + equally "throwaway" hotswap disk pools. In such a 1990s-reminiscent setup, i3.metal seems like a sensible upper bound for how big such an instance could get, economically. Maybe you don't need the 2Us; maybe you can build the disk pool as something like a disk server (i.e. a dedicated PCIe backplane leading to RAID-controller daughterboards? something something Thunderbolt?) That'd lower the TCO of these host machines a bit, but it'd still be questionable how much usage such instances would get—and when they're unscheduled, despite the disks being a separate physical box, those disks would still be unavailable for any other instance to use, even ones scheduled to machines in the same rack. (Or maybe they could get really fancy, and have a disk-server rack that can present itself as a RAID controller over PCIe, such that, most of the time, it can just be a disk server, but when it gets reserved by a Nitro-i-instance, it can switch off its Infiniband cards and switch on its PCIe client interface cards, and moonlight as a RAID array. If AWS does get bigger i-type instances, I'd wager that this is what they would have built to achieve that. That or custom RAID controller cards that present Infiniband-rDMA targets as if they were local NVMe devices, and don't allow host configuration, only BMC configuration.)
- WestCoastJustin 8y agoThere might be a networking cap & disk I/O issues with the instances you picked on GCP vs AWS. The GCP instance has 8 Gbps vs 10 Gbps for AWS. I don't really know without seeing the graphs from the instances, if you hit a cap, but this could make a difference in both transfer speeds and latency #'s for GCP. Also, for your local disk test, on GCP, disk size makes a difference to get the best performance. The larger the disk, the better the performance. PD disk read/write performance also comes out of the available network bandwidth! So, the instance you picked on GCP was at a disadvantage right from the start [3]. This likely explains the I/O Experiment graph and the "67x difference in throughput" as you're likely hitting caps, both in terms of network bandwidth, and disk performance compared to AWS. Seeing anything where it is x67 difference is a pretty big red flag that something strange is going on and needs further investigation. GCP's n1-standard-16 = 8 Gbps max [1] AWS's c5d.4xlarge = 10 Gbps max [2] I guess the problem with comparing clouds, it is never apples vs apples, and I don't fault you for picking what do you (as it is not obvious). GCP typically gives you (core count / 2) = # Gbps network bandwidth. A good followup to your comparison might be to investigate why they #'s are different. Does adding more cpus, memory, network bandwidth increase performance? [1] https://cloud.google.com/blog/products/gcp/5-steps-to-better-gcp-network-performance https://cloud.google.com/blog/products/gcp/5-steps-to-better... (see section #3). [2] https://aws.amazon.com/blogs/aws/ec2-instance-update-c5-instances-with-local-nvme-storage-c5d/ https://aws.amazon.com/blogs/aws/ec2-instance-update-c5-inst... [3] https://cloud.google.com/compute/docs/disks/performance#size_price_performance https://cloud.google.com/compute/docs/disks/performance#size... (see the table re: disk size to bandwidth)
- halbritt 8y agoI don't understand the throughput numbers given. 5.6GB/s for GCP and 9.6GB/s for AWS would be 44gbps and 76gbps respectively. I don't don't know of any instances offering that kind of throughput. I've personally validated GCP's statement that they offer 2gbps/core up to 16gbps. I can get 16gbps consistently between any two n1-standard-8 using iperf. This generally makes network IO in GCP much cheaper.
- jsnell 8y ago
- kharms 8y agoFor someone who has worked with both: which offers a better developer onboarding experience, for small web/data apps? Ease of learning and use wise.
- KirinDave 8y agoThere are many dimensions in which working with GCP is a lot better for small and medium apps, if only the proximity to Firebase. These results are, to my mind, quite questionable. They certainly don't line up with my personal experience and the measurements I've done in the last year. AWS is much harder to use correctly right now, to me.
- social_quotient 8y agoI prefer aws to google and msft for account mgmt purposes. google and msft are so dead set on binding your logins to your global accounts which might be used for other things. I know someone might argue that aws and amazon.com are sharing the same account, I guess they "can" but its easy to make a new aws account with just an email - I think we manage about 12-15 different aws accounts. We tend to isolate major clients or projects by making entirely new accounts. I don't feel like I'm working uphill to keep multiple aws accounts from merging or somehow binding to my personal shopping account for amazon.com. With msft and google it feels like I am always almost "tricked" in to binding multiple accounts and login states together. I would argue that once you have an account up and running most of these guys are similar with differences. The ui on aws isn't glamorous but I find its utilitarian simplicity pretty easy to deal with and get most things done.
- manigandham 8y agoThat seems to be dependent on whether you're working on your own personal/company accounts or on behalf of other clients. For owned accounts, I prefer GCP and Azure because the logins are seamlessly integrated into GSuite and Office 365 so we can manage IAM on an individual basis in one place.
- snuxoll 8y agoOn the note of account management, you cannot be a Google Cloud Partner without GSuite or Google Cloud Identity accounts, period, end of story. This whole setup is just ludicrous, I have to pay Google money just to have a partner login? Even Microsoft offers a "good enough" tier of Azure AD that can be used to sign up for their partner portal, and AWS just uses normal Amazon accounts.
- nodesocket 8y agoDeveloper experience on GCP is vastly superior to AWS. - Pricing on GCP is much easier, no need to purchase reserved instances, figure out all the details and buried AWS billing rules. Run your GCP instances and automatically get discounts. AWS reserved instances requires knowing your instance types, knowing that you can purchase the smallest type of an instance class and combine, knowing that you can only purchase 20 reserved instances per zone/per region in an account. So many gotchas. - GCP projects by default span all regions. It is much easier if you run multiple regions, all services can communicate with all regions. Multi-region in AWS is sort of a nightmare, setting up VPC peering, can't reference security groups across regions, etc.. - Custom machine types. With GCE, you simply select the number of cores you need and memory. No trying to decipher the crazy amount of AWS instance types T2, T3, M5, M5a, R5, R5a, C5, C5n, I3... - Instance attached block storage is easier to grok and in my experience is much faster than EBS. The bigger the disk on GCE, the more IOPS. No provisioned IOPS madness.
- briffle 8y agoAnother great point to GCP: -- You can setup an 'organization' that will hold multiple projects, in a hierarchy of folders. You can set permissions/roles on a folder, and have them propagate down to all projects underneath.
- _wmd 8y ago> no need to purchase reserved instances GCE offered committed use discounts for quite some time (note: completely different from sustained use discount that is automatic), by ignoring this discount tier, the results from this post look significantly worse The idea that "developer experience" is paramount is the entire reason why there is an entire sub-industry of vendors dedicated to cost optimization, following in the wake of choices made with completely the wrong business priorities in mind
- yovagoyu 8y agoAWS has cheaper discounted instances though too.
- halbritt 8y ago
- sethvargo 8y agoHey all - Seth from Google here. Thank you to the authors who worked on this report. These types of reports help us better understand the ways in which our customers and partners utilize our platform. Our team is reviewing the report and will provide a response as we conduct our own benchmarks. Varying factors impact these types of benchmark analyses, many of which are difficult to isolate and control. As an example, I refer to some of the benchmarks others have posted in this very thread. As a technical practitioner, I'm positive there are areas in which cloud A outperforms cloud B and vice versa, giving users choice and flexibility. As an employee of Google, I can assure you that we are committed to providing best in class performance and availability on our platform. Thank you for your patience as we review these findings and craft our responses.
- peferron 8y agoWhere do you intend to publish your response? GCP blog [1]? I'd like to make sure I won't miss it, and this HN thread may be buried by the time you come up with a response. [1] https://cloud.google.com/blog/ https://cloud.google.com/blog/
- cybernoodles 8y agobump
- ahmedalsudani 8y agoBumping does not work on HN. If you want to give something visibility, upvote it.
- sethvargo 8y agoHi there - sorry for the delayed response. I'm still working on getting an answer here, but I didn't want to give the impression that I was ignoring the question. My suspicion is that we'll either work with the original authors or publish on the GCP blog, but I can't confirm any of those options at this time.
- socceroos 8y ago
- ta_271828 8y agoThis is interesting given that I heard on the grapevine that some major cloud players are actually using AWS on the back-end even though they are advertising say as.. "GCP".. wonder if anyone can confirm or deny...
- WestCoastJustin 8y agoBy nature large companies have massive global teams and there is no single provider for anything. Team A could using AWS, while team B cloud be using GCP, and team C is using Azure. Just because team B says they are using GCP doesn't mean the others are lying. Or, that there is anything weird going on.
- ta_271828 8y agoActually I meant to say that the news on the grapevine is that the big cloud players (GCP etc.) are potentially outsourcing demand for cloud services in excess of their capacity to AWS...
- deleted 8y ago[deleted]
- WestCoastJustin 8y agoTo be blunt. No. This is totally absurd and would be trivial to detect if true. On the legal side, you'd be breaking all sorts of ToS, Security, and Privacy agreements (you'd have to disclose this via a data processor clause). On the technical side, latency would also be so obvious to detect this too. They are totally different hardware/software platforms and would have different characteristics (as proven by this thread). On the business side, no AWS/GCP/Azure CEO will ever do this, staff would totally be aware too. This is 100% bogus. I actually did not even want to reply and this make zero sense but felt invested. Whoever told you this doesn't know what they are talking about.
- JaRail 8y agoThat said, hosting providers who historically ran services on only their own hardware could certainly be load-balancing to cloud hosting providers. This is certainly not what the comment you responded to implies. However, it is something I'd expect people to get confused about..
- derefr 8y ago> What about network throughput variance? On AWS, the variance is only 0.006 GB/sec. This means that the GCP network throughput is 81x more variable when compared to AWS. What would cause this particular effect? It's very interesting. Is it, perhaps, that with GCP you're hitting the capacity of the network, while with AWS you're being artificially capped at that speed on a network that could theoretically go faster? Or maybe it's just different strategies for bandwidth-limiting instances employed by AWS's SDN layer vs. GCP's? Probabilistic packet-drop (to force TCP window scaling) vs. artificially-induced nanosecond-scale egress latencies?
- wmf 8y agoGCP Andromeda is software-based network virtualization which tends to have lower performance and higher performance variability. https://www.usenix.org/node/211244 https://www.usenix.org/node/211244 AWS Nitro/ENA is hardware network virtualization which is faster and more consistent.
- Rafuino 8y agoHmm it'd take forever to dig into the respective documentation at AWS and GCP, but from a quick look the CPU frequency alone is quite different (3.0 GHz for AWS's Xeon Scalable and 2.0 GHz for GCP's Xeon Scalable), and we don't know anything about CPU cache sizes, etc. etc. That's problematic to start. Then we have very little info to go by on the underlying storage performance. More problematic, I don't know how large the working data set is for TPM-C (i.e. how many warehouses are being simulated?), so I can't tell how much of the storage is being used. I assume it's larger than the 60GB of DRAM offered on the GCP instance (thus spilling into the storage), but with the CPU differences and unknown storage performance, I don't know what to make of this report.
- raboukhalil 8y agoChoosing a cloud provider is about more than just performance. For me, I lean towards GCP because of the combination of awesome UX, custom VM configurations (which also means GPUs attached to a custom # CPUs), no-bidding spot instances, and the multi-regional cloud storage offering that replicates data across many regions for much cheaper than AWS. I wrote about this last year if you're curious: https://medium.com/@robaboukhalil/a-tale-of-two-clouds-amazon-vs-google-4f2520516a38 https://medium.com/@robaboukhalil/a-tale-of-two-clouds-amazo...
- yovagoyu 8y agoYeah, because who cares about performance when you have a pretty UX?
- hueving 8y agoYou're confusing UX with UI. You can have an excellent UX that's still just a CLI.
- yovagoyu 8y agoGoogle uses "gsutil" and "gcloud" whereas AWS just uses "aws". Both are clear but because Google arbitrarily has 2 tools I constantly have to remember which one had which feature. It's not a big thing but little stuff like makes it more cumbersome to use.
- ernsheong 8y agoThere's also bq for BigQuery stuff. If they did merge... it would be "gcloud bq..." or "gcloud storage...", my fingers protest :)
- withhighprod 8y agoAWS sales and support team is way ahead GCP
- ravedave5 8y ago
- nogbit 8y agoLaunch EKS on AWS, wait 20-30min. Launch GKE on GCP, wait 5min.
- withhighprod 8y agoAm I the only one thinking the huge sales team (yes the solution architect org) is the key for this game?
- kerng 8y agoHow does this compare to the number 2 in the cloud, Microsoft Azure?
- iamgopal 8y agoSo, in conclusion, When Using cockroachdb, with using 32gb ram instead of 60gb ram, and with different throughtput setting, you may consider AWS as it provides slight better cost because of using lower resources. And since there are no other big player in the market, we will test only two of them, and our only choice will be AWS.
- null000 8y agoHonestly I would have expected performance on a cloud provider to be measured in x per $. You're pretty much renting everything, the hardware etc is really difficult to compare, and you can usually throw more machines at the problem anyway, so measuring a single machine vs a single machine doesn't make much sense if the two might cost vastly different amounts.
- riking 8y agoDid anyone consider comparing the cost of this vs. Cloud Spanner? Because CockroachDB has the explicit inspiration of being Google's Spanner without the special hardware, so... why not just use Spanner with the special hardware instead?
- ernsheong 8y agoMost small companies don't reserve for 3 years. The on-demand monthly cost is 388 (GCP) vs 562 (AWS) per instance for the instances in the report (omitting SSD costs).
- ernsheong 8y agoFurthermore, GCP encrypts its disks by default. So I'd expect a slight degradation of performance.