5 ms·
Exploring a ‘Deep Web’ That Google Can't Grasp
- rozim 18y agoGreg Linden's (MSFT) comments on a recent Google paper on this: http://glinden.blogspot.com/2009/01/how-google-crawls-deep-web.html http://glinden.blogspot.com/2009/01/how-google-crawls-deep-w...
- jaspertheghost 18y agoThere's many startups attempting to do this including pipl.com, http://cazoodle.com/ http://cazoodle.com/ among others. Here's some research about it: http://www-sal.cs.uiuc.edu/~kcchang/ http://www-sal.cs.uiuc.edu/~kcchang/
- sam_in_nyc 18y agoI believe I've seen this type of crawling in action in request logs. For example, Yahoo might try to request "news.ycombinator.com/user?id=britney_spears", even though it's not linked to from anywhere.
- smanek 18y agoThat haystack is infinitely large. Gah. I really hate it when people misuse "infinite".
- kristiandupont 18y agoWhile it is not really infinite, I think that in this situation, the term is justified as it is semantically very close to it.
- smanek 18y agoIt really isn't ... the semantics of infinity are completely different than the semantics of 'really, really, big.' An infinite data store would, by definition, contain my DNA, correct (and incorrect) predictions about the universe (down to the molecule) for all time, the true value of Pi, and every piece of knowledge that has ever or will ever exist. An actual infinite is just a ludicrous concept.