4 ms·
This is a complex misunderstanding... First, we are getting better throughput from S3 than I we were using a SATA SSD. (and slower than a NVMe SSD). This is a
by fulmicoton 5y ago
This is a complex misunderstanding...
First, we are getting better throughput from S3 than I we were using a SATA SSD. (and slower than a NVMe SSD).
This is a bit of a secret.
Of course, single sequential throughput on S3 sucks. At the end of the day the data is stored on spining disk
and we cannot do anything against the law physics.
... but we can concurrently read many disks using s3. Network is our only bottleneck.
The theoretical upper bound on our instances is 2GB/s.
On throughput intensive 1s query, we observe an average of 1GB/s.
Also you are not accounting for replication. S3 costs include battle tested, multi-DC replication.
Last but not least, S3 trivially decouples compute and storage.
It means that we can host 100 different indices on S3, and use the same pool of search server to deal with the
CPU-bound stuff.
This last bit is really what drives the price an extra 5x down for many use case.
- 2Gkashmiri 5y agoHow about setting up minio on these hertzner setups? You get benefit of s3 on cheap hardware without aws costs
- fulmicoton 5y agoAbsolutely! I want to try that.. We are especially interested in testing the latency minio could offer.
- plater 5y ago"S3 costs include battle tested, multi-DC replication." Sometimes we pay a bit too much for this multi-replication, battle tested stuff. It's not like the probability of loosing data is THAT huge. For the 4x extra cost you could easily take a backup every 24h. "It means that we can host 100 different indices on S3, and use the same pool of search server to deal with the CPU-bound stuff" You can do that with NFS. It's amazing how much we are willing to pay for a bunch of computers in the cloud. Leasing a new car costs around $350/month. You could have three new cars at your disposal for the same price as this search implementation.
- curryst 5y ago> For the 4x extra cost you could easily take a backup every 24h. It's also worth considering the cost to simply regenerate the data for something like this that isn't the source of truth. You'll lose any content that you indexed that has disappeared from the web, but that seems like a feature more than a bug. > You can do that with NFS. You're going to be bound by your NIC speed. You can bond them together, but the upper bounds on NFS performance are going to be significantly lower than on S3. Whether that's going to be an issue for them or not, I don't know, but a big part of the reason for separating compute and storage is so that one of them can scale massively without the other.
- layla5alive 5y ago100Gbps NICs are cheap, relative to the price of the cloud...
- deleted 5y ago[deleted]