4 ms·
> SearchHut indexes from a curated set of domains. The quality of results is higher as a result, but the index covers a small subset of the web. [citation need
by birken 4y ago
> SearchHut indexes from a curated set of domains. The quality of results is higher as a result, but the index covers a small subset of the web.
[citation needed]
The quality of the results right now are not very high, and in theory I don't understand why one would believe a search engine with a hand picked set of domains would be expected to outcompete a search engine that can crawl the entire web and determines reputation by itself. This also ignores the fact that a lot of domains have a mix of high quality content and low quality content, for example twitter or medium. If you are going to rely on domain-level reputation then your search engine is going to be way behind the search engines that can judge content more specifically, which is all of the other search engines.
If you were to tell me curated domains is just a bootstrapping method and as the search engine evolves it will change, fine, but right now the search engine is so simplistic that the theory of how it might be good is really the only point. And if that underlying theory is dubious, and the infrastructure is simplistic and obviously won't scale, then I don't know what is interesting or novel about this right now. Doesn't seem worthy of reaching the top of HN.
- p-e-w 4y ago> If you are going to rely on domain-level reputation then your search engine is going to be way behind the search engines that can judge content more specifically, which is all of the other search engines. Then why do Google and DuckDuckGo return 90% garbage for most queries? "All of the other search engines" have completely failed to keep pages from the results that are not only low-quality, but outright spam.
- mda 4y agoThey definitely do not return "90% garbage for most queries". This 8s an unsubstantiated claim I see often i HN and honestly not backed by any real data. e.g. You can check your search history and see it yourself.
- p-e-w 4y agoI just tried searching for "python str" on Google. I expected the top result to be a link to the official Python docs for the `str` type, then ideally some relevant StackOverflow questions highlighting common Python issues with strings, bytes, Unicode etc. Instead, the top result was W3Schools. Then came the Python docs, then 5 pages somewhere between blogspam and poor-quality tutorials. Then a ReadTheDocs page dating to 2015. And that was it. No more official Python resources, no StackOverflow. In the middle of the results some worthless "Google Q&A" dropdowns that lead to more garbage quality content. So for this query, using my definition of "garbage", the "garbage percentage" is somewhere between 80% and 90+%, depending on how many Q&A dropdowns you waste your time opening.
- jwilk 4y agoFor me, https://docs.python.org/3/library/stdtypes.html https://docs.python.org/3/library/stdtypes.html is the top result.
- p-e-w 4y agoThe fact that the ranking of results for queries that have nothing to do with location-based services depends on where you are located (and, possibly, on whether or not you are logged in) is one of the worst things about Google. And the fact that you can't seem to disable that behavior is even worse.
- goldsteinq 4y agoI just tried searching for “python str” on searchhut and the top result is Postgres docs, then Wikipedia article for empty strings and then Drew’s blog. Official Python docs isn’t in the index at all.
- jwilk 4y agoFor me the second hit is: https://docs.python.org/3/howto/clinic.html https://docs.python.org/3/howto/clinic.html So at least some official Python docs are indexed.
- 4y ago
- birken 4y ago> Then why do Google and DuckDuckGo return 90% garbage for most queries? If you can give me a list of 10 normal-ish queries where 9 out of the first 10 results on Google or DDG are "garbage", then I'll concede your point. I think you are creating an impossible standard for search engines, then using it to deem the current ones as failures. While at the same time ignoring that this new search engine is, as present, unusable with no realistic argument for why it might eventually be better.
- p-e-w 4y agoSee my reply on the sibling comment for an illustrative example.
- slimsag 4y ago> Notice! This product is experimental and incomplete. User beware! Seems like your expectations are misplaced. Being at top of HN is not an indicator of quality, just interest.
- deleted 4y ago[deleted]
- yjftsjthsd-h 4y ago> why one would believe a search engine with a hand picked set of domains would be expected to outcompete a search engine that can crawl the entire web and determines reputation by itself. Because SEO manipulation is a well developed field, ensuring that the search engines trying to determine reputation automatically will (and does) end up with bad results.
- p-e-w 4y agoIndeed. Whatever "smart" algorithm you use to rank results, you can be certain that half the web will turn into adversarial examples once your engine becomes popular enough.
- petercooper 4y agoIf you were to tell me curated domains is just a bootstrapping method and as the search engine evolves it will change, fine This makes me think of a possible approach. Curate a giant set of domains that almost exclusively host high quality content. Crawl said domains. Use all of the crawled data as a training set to create a model with which to ascertain the quality of random Web pages from other domains. Then spider everything and run it against the model.