3 ms·
How is it that the S3 API is remotely fast enough to make this work? As search engine that operates at any kind of scale needs to skip through very large files
by ccleve 4y ago
How is it that the S3 API is remotely fast enough to make this work?
As search engine that operates at any kind of scale needs to skip through very large files to evaluate a query. You need very low-latency, high-bandwidth access to disk. A search engine instance that accesses files on a local SSD is an order of magnitude faster than one that puts files on EBS.
They make some mention of local caching, but the devil is in the details here. Does all data get copied to local cache? What is the performance here?
- ianbutler 4y agoProbably somewhat similar to how Trino/Presto/Bigtable/Spanner works but targeted at search, decompose the query into a set of highly parallelizable steps and execute them simultaneously over the set of data using some type of specialized storage format for rapidly indexing into the file, some really nice heuristics, and then drop all the ones without a potential for a hit, aggregate the rest and then do maybe a more classical search over the vastly reduced set of potential files in memory. I know Presto isn't focused on search, but Athena (AWS branded Presto) can do some really fast queries over S3, the issue is coldstart time on the compute, for a similar solution focused on search maybe you keep the compute always warm and work from there.
- remram 4y agoIf you're doing a single round-trip it's really not bad. You don't get that big an impact compared to the round-trip to the user. If you are doing multiple dependent loads, e.g. loading an index that tells you which part of the data to load which tells you which other related table to look into (e.g. a complex join)... that would be bad.