4 ms·
I just did some tests and the perf is actually worse than that — but it’s actually still perfectly fine for our use cases. - 15m row db will return a search fo
by bomewish 2y ago
I just did some tests and the perf is actually worse than that — but it’s actually still perfectly fine for our use cases.
- 15m row db will return a search for an extremely common term in ~3s (9m rows have that term);
- When adding terms, so that only ~200k rows have the combination, response is ~2s
- That 15m db was in 16 shards. (I’m actually not sure why it should be so slow; doesn’t actually make sense given that a single 1m row db is not that slow, might look into it but not a huge priority)
- For a ~200k row db, search is ~150ms for three terms that appear in ~1.5k entries
- Obviously the network io is the slower piece anyway
This perf is fine for our use case — a research system for internal users. Certainly worth the trade-off of not having to deal with a more complex database like elasticsearch, or even PG with the tantivy extension (which we might switch to some day).
For sharding — we only shard the huge one. To make sure bm25 rankings are still _roughly_ the same as they would be if we did not shard, we just randomised the rows assigned to the num_core shards. It’s a 16 core machine so 16 shards.
We know/ensure each db file gets its own core because we use aioprocessing in python. It handles both multicore and async. Running a query with htop shows all the cores light up (unlike with unsharded).
It’s all about trade-offs in the end. We cache the first search then pre cache the next page (and so on) so the user only has to wait 2-3s when searching the big db for that initial search and after that it’s pretty snappy. Most searches are on much smaller dbs (thousands to hundreds of thousands of rows) and results there are often 10ms or something. No user complaints so far.