4 ms·
What's been a shame is that there is still no open source Search Engine despite this being a "solved" problem. Like not even something like there is a docker im
by guix992 8y ago
What's been a shame is that there is still no open source Search Engine despite this being a "solved" problem. Like not even something like there is a docker image that you throw at your cluster that gets you faster and faster search results. That's the real shame.
We should've commidified the core of search engine by now with programmatic and API access as commonplace and yet here we are where search engine software is still dominated by proprietary services.
- Gaelan 8y agoI’m guessing that it’s impractical to self-host something big enough to be useful.
- guix992 8y agoI don't believe that is a good enough limitation in this day and age of docker and kubernetes. You take the search engine docker image and keep scaling the number of nodes that are running it in the cluster till your search result time is fast enough for you.
- cabaalis 8y agoI am sure this already exists even as I type it, but this sounds like a great job for a distributed system running a federated search protocol.
- guix992 8y agoThe hard part is already solved, you don't even have to crawl the web to build the index. There is already a periodically refreshed index of the web that you can download: commoncrawl.org Now someone just needs to configure, Apache Lucene as a proper docker image that can consume this index.