4 ms·
Bioinformatician here. I appreciate that you need to make money somehow. I do find this a little hard to believe, though: > full-text search of 20MM PDFs at ou
by xaa 9y ago
Bioinformatician here. I appreciate that you need to make money somehow. I do find this a little hard to believe, though:
> full-text search of 20MM PDFs at our traffic levels is expensive to operate
Assuming you convert them to text once, index them, and put them in a standard FTS engine, I'd guess it is on the order of 100GB-1T of text (max), plus some more for the index (basing these estimates with my experience text mining PubMed Central and MEDLINE). So it can all fit on a pretty standard server. Maybe at 100 req/s it would take a few. Yes, you'd want replication.
The number of servers required to get good latency FTS is the part of this that I'm least familiar with. Anyone have a ballpark, given these estimates, on what kind of hardware would be required? (I could easily be wrong, and indeed this is very expensive. If so, I'd be curious about ballpark numbers)
- minxomat 9y agoI maintain a P2P on-premise FTS search. Though this one indexes many types of text (plain, HTML, PDF, DOCs). One 8c server (running about 10 workers) can handle 8 to 10qps, depending on the depth required. This is on an index of 20 million documents. If the number of workers is constant, doubling the index will have the qps. 2 million docs take about 50GB of disk space (20 million = 500GB, 1TB with redundancy). It's better to go with SSD arrays here, since random IOPs are much higher than for other workloads. This can skyrocket cost. So for this (our) system, it could be as cheap as $1k for the hardware, e.g. using the Foxconn Purus cloud server: http://www.bargainhardware.co.uk/cheap-e5-2600-lga2011-sixteen-core-cloud-foxconn-server-configure-to-order/ http://www.bargainhardware.co.uk/cheap-e5-2600-lga2011-sixte...
- xaa 9y agoThanks so much. Fascinating. I've made many of these types of app but never to "web scale". Bioinformatics apps are a bit niche, and we take the view that our non-paying academic "customers" can wait however long it takes to finish the query, in the unlikely event that there is high load. Just to make sure I understand correctly, that's about $1-2K of one time cost per 10qps (w/o SSD, and not counting power and maintenance, etc)? When I first saw "cloud server", I thought that was a per-month rental cost, but the link is for actual in-house hardware. If this is even close to correct, my suspicions seem confirmed. Except for one thing. I have no idea how many qps a site like Academia would have. 100qps was completely out of my ass, but it seemed hard to imagine it being any more than 1-2 orders of magnitude higher, at most. Any guess on that?
- minxomat 9y agoYeah, in my example, the $1k is for one server (cluster node). These servers (Purus, Quanta etc.) are commonly used to rapidly build enterprise clouds (you usually buy them by the rack). It's the closest thing you get to plugging a network cable into a bunch of Xeons. The cost of one system breaking is negligible. This is not counting the colo costs. You can also do this with virtual public clouds (AWS, linode, GCP et al), but you'll of course pay a premium for the infrastructure. This might be worth it though, because you can now scale within seconds to handle qps bursts. Usually, latency can be lowered by going baremetal (see e.g. Algolia). Academia should be able to handle more qps than our system, because the queries are really trivial in comparison. With decent caching, an 8c should be able to do 50 to 80qps. That's what I get from a few experiments when I switch my test cluster into restricted mode (basically just substring search). Of course I can only speak from my experience, not how this can be applied to Academia's existing infrastructure. Testing large search engine deployments can be really, really frustrating.
- eecc 9y agoHi, what do you mean with P2P FTS, can you please elaborate? Solr, Elastic, some other custom thing you wrote? Language? I'm curious...