3 ms·
I thought this would be interesting to the audience here. Uber is well known for its scale in the industry. Here are the latest numbers I compiled from a plet
by enether 2y ago
I thought this would be interesting to the audience here.
Uber is well known for its scale in the industry.
Here are the latest numbers I compiled from a plethora of official sources:
- Apache Kafka:
- 138 million messages a second
- 89GB/s (7.7 Petabytes a day)
- 38 clusters
- Apache Pinot:
- 170k+ peak queries per second
- 1m+ events a second
- 800+ nodes
- Apache Flink:
- 4000 jobs
- processing 75 GB/s
- Presto:
- 500k+ queries a day
- reading 90PB a day
- 12k nodes over 20 clusters
- Apache Spark:
- 400k+ apps ran every day
- 10k+ nodes that use >95% of analytics’ compute resources in Uber
- processing hundreds of petabytes a day
- HDFS:
- Exabytes of data
- 150k peak requests per second
- tens of clusters, 11k+ nodes
- Apache Hive:
- 2 million queries a day
- 500k+ tables
A lot of thought and investment (Uber's annual R&D budget is around $2B) has been put behind this data infrastructure, particularly driven by their complex requirements which grow in opposite directions:
1. Scaling Data - total incoming data volume is growing at an exponential rate
1. Replication factor & several geo regions copy data.
2. Can’t afford to regress on data freshness, e2e latency & availability while growing.
2. Scaling Use Cases - new use cases arise from various verticals & groups, each with competing requirements.
3. Scaling Users - the diverse users fall on a big spectrum of technical skills. (some none, some a lot)
Uber leverages these open source technologies largely in a Lambda Architecture, with Presto used to bridge the gap between both, allowing users to use SQL to query and join across all data stores.