6 ms·
I wonder how much their setup costs. Naively, if one were to simply feed 100 PB into Google BigQuery without any further engineering efforts, it would cost abou
by randomtoast 2y ago
I wonder how much their setup costs. Naively, if one were to simply feed 100 PB into Google BigQuery without any further engineering efforts, it would cost about 3 million USD per month.
- AJSDfljff 2y agoGood question. I thought it would be a no brainer to put it on s3 or similiar but thats already way to expensive at 2m/month without api requests. Backplace storage pods are an initial investment of 5 Million, thats probably the best bet you could do and on that savings level, having 1-3 good people dedicated to this is probably still cheaper. But you could / should start talking to the big cloud providers to see if they are flexible enough going lower on the price. I have seen enough companies, including big ones, being absolut shitty in optimizing these types of things. At this level of data, i would optimize everyting including encoding, date format etc. But i said it in my other comment: the interesting questions are not answered :D
- orf 2y agoThe compressed size is 20pb, so it’s about 500k per month in S3 fees
- francoismassot 2y agoIndeed. They benefit from a discount, but we don't know the discount figure. To further reduce the storage costs, you can use S3 Storage Classes or cheaper object storage like Alibaba for longer retention. Quickwit does not handle that, so you need to handle this yourself, though.
- AJSDfljff 2y agoI would probably build my own storage pods, keep a day or a week on cloud and move everything over every night.
- jcgrillo 2y agoLogs should compress better than that, though, right? 5:1 compression is only about half as good as you'd expect even naive gzipped json to achieve, and even that is an order of magnitude worse than the state of the art for logs[1]. What's the story there? [1] https://news.ycombinator.com/item?id=40938112 https://news.ycombinator.com/item?id=40938112
- Aurornis 2y agoThey provide some big hints about the number of vCPUs and the size of the compressed data set on S3: > Size on S3 (compressed): 20 PB There are also charts about vCPUs and RAM for the indexing and searching clusters.
- gaogao 2y agoYeah, doing some preferred cloud Data Warehouse with an indexing layer seems fine for this sort of thing. That has an advantage over something specialized like this of still being able to easily do stream processing / Spark / etc, plus probably saves some money. Maybe Quickwit is that indexing layer in this case? I haven't dug too much into the general state of cloud dw indexing.
- fulmicoton 2y agoQuickwit is designed to do full-text search efficiently with an index stored on an object storage. There are no equivalent technology, apart maybe: - Chaossearch but it is hard to tell because they are not opensource and do not share their internals. (if someone from chaossearch wants to comment?) - Elasticsearch makes it possible to search into an index archived on S3. This is still a super useful feature as a way to search punctually into your archived data, but it would be too slow and too expensive (it generates a lot of GET requests) to use as your everyday "main" log search index.
- BiteCode_dev 2y agoClick house does have it, but it's experimental.
- Daviey 2y ago"Object storage as the primary storage: All indexed data remains on object storage, removing the need for provisioning and managing storage on the cluster side." So the underlying storage is still Object storage, so base that around your calculations depending if you are using S3, GCP Object Storage, self hosted Ceph, MinIO, Garage or SeaweedFS.
- onlyrealcuzzo 2y agoA lot. 1PB with triple redundancy costs around ~$20k just in hard drive costs per year. That's ~$2.5M per year just in disks. I'd be impressed if they're doing this for less than $1.5M per month (including SWE costs). Obviously, if they can, saving $1.5M a month vs BigQuery seems like maybe a decent reason to DIY.
- BiteCode_dev 2y agoWhy per year? If they buy their own server, they keep the disk several years. The money motivation to self host on bare metal at this scale is huge.
- onlyrealcuzzo 2y ago> Why per year? If they buy their own server, they keep the disk several years. The cost per year is much higher - that's using a 5-year amortization.
- BiteCode_dev 2y agoSeems high. You can get a spinning disk of 18TB (not need for SSD if you can parallel write) for 224€. Let's round that to $300 for easy calculations. To store 100 petabytes of data by purchasing disks yourself, you would need approximately 5556 18TB hard drives totaling $1,666,800. Of course, you'll pay more than the disks. Let's add the cost of 93 enclosures at $3,000 each ($279,000), and accounting for controllers, network equipment ($100,000), and power and cooling infrastructure ($50,000, although it's probably already cool where they will host the thing), that would be a about $2.1 M. That's total, and that's for the uncompressed data. You would need 3 times that for redundancy, but it would still be 40% cheaper over 5 years, not to mention I used retail price. With their purchasing power they can get a big discount. Now, you do have the cost of having a team to maintain the whole thing but they likely have their own data center anyway if they go that route.
- JackSlateur 2y ago
- francoismassot 2y agoGood question. Let's estimate the costs of compute. For indexing, they need 2800 vCPUs[1], and they are using c6g instances; on-demand hourly price is $0.034/h per vCPU. So indexing will cost them around $70k/month. For search, they need 1200 vCPUs, it will cost them around $30k/month. For storage, it will cost them $23/TB * 20000 = $460k/month. Storage costs are an issue. Of course, they pay less than $23/TB but it's still expensive. They are optimizing this either by using different storage classes or by moving data to cheaper cloud providers for long term storage (less requests mean you need less performant storage and usually you can get a very good price on those object storages). On quickwit side, we will also improve the compression ratio to reduce the storage footprint. [1]: I fixed the num vCPUs number of indexing, it was written 4000 when I published the post, but it corresponded to the total number of vCPUs for search and indexing.
- rcaught 2y agoSavings plans, spot, EDP discounts. Some of these have to be applied, right?
- Onavo 2y agoAt this level they can just go bare metal or colo. Use Hetzner's pricing as reference. Logs don't need the same level of durability as user data, some level of failure is perfectly fine. I would estimate 100k per month or less, maximum 200K.