4 ms·
Offline search is the obvious alternative. Every cell in your body has its own data and search engine. It doesn't need a Google, Microsoft, Yandex, and Baidu. H
by hux_ 9y ago
Offline search is the obvious alternative. Every cell in your body has its own data and search engine. It doesn't need a Google, Microsoft, Yandex, and Baidu. How come?
Downloading a dump of Wikipedia or Stackoverflow and indexing it however you like on mid range hardware is trivial today. Look at what Kiwix or Zeal(offline doc browser) do today. It's just a matter of time before people start packaging up data and indexes customized to your info needs downloadable like a song from iTunes.
- ChuckMcM 9y agoHaving built a search engine I can tell you that your minimum index size is on the order of 5 billion documents. You can back fill long tail searches with one of the big players. For minimum hits on the backfill 10 billion documents are better. To get 10 billion high quality documents you will want a “crawl frontier” of probably 500 billion URIs, which, given the power log rule of data you will need to actually crawl about 20% of those every 7 days. Of course you can hold much fewer pages in your index if you are not particularly broad in your searches. The gotcha there is that when you need that thing you don’t normally need, you either have to go online or go without. Not impossible of course, just that the scale may be larger than you expect.
- hux_ 9y agoThe value of the long tail has been oversold I feel. My estimate based on my own usage for the past year and a half or so, of curating my own local indexes is about 30-40 million docs. This is the equivalent of having your own personal Library of Alexandria (text and images no video). I don't have numbers but I work offline a lot and my guess is 70-80% of my queries are probably getting satisfied by my local dumps.
- ChuckMcM 9y agoNot disagreeing with you, I was just trying to connect that the index size is a function of the breadth of information. For example, I have digitized roughly 300 volumes (books) which collectively represent about 100,000 pages of information. That collection which represents a big chunk of my originally print reference library creates a relatively small n-gram index of about 22 GB. But it is a small sample of generally available reference material and doesn't include the 22 years of digitized Scientific American articles (much harder to parse out for indexing when starting with the PDF form). But it still answers a lot of reference queries quickly and accurately for things I am interested in. For things that I become interested in and have yet to have started curating a set of references for, its worthless. As a result my experience is that the closer I get to my long term interests the more likely I am to find something in my library to answer the question, things that are more temporal (news, new research) are not there at all generally, and things that are only now of interest are similarly not represented. The thing that Search engines do so remarkably well is that they cover a very wide swath of interesting material preemptively. To host that locally would be a more significant effort for me.
- hux_ 9y agoAh full text search...sorry my bad. I have been referring to much simpler indexes. All mostly sub 1GB that fit in memory. For Stackoverflow for example title, URL and tag indexes. Wikipedia - title, URL, categories, geotags. This has been working out somehow for my general use. It has the feeling of working inside a library with card indexes. Lot of decent work can still get done. I agree with your points on why the bigbois are valuable and relevant when it comes to temporal and new constantly changing info. But what I am finding is (probably unique to my usecases) is I have enough info on disk to keep me occupied and productive for long periods totally offline.