7 ms·
Even with a pool of proxies, I would expect an instance of this "metasearch engine" to quickly get banned by the other search engines. The same IP running thous
by halflings 9y ago
Even with a pool of proxies, I would expect an instance of this "metasearch engine" to quickly get banned by the other search engines. The same IP running thousands of queries and scraping its content (which is against their ToS) should be easily detectable.
- VoidWhisperer 9y agoThis is self-hosted, so I'm assuming it's running under the assumption that each person hosts their own instance and uses that instance. The number of queries coming from the instance in that case wouldn't look too out of the ordinary.
- lawl 9y agoThen that defeats the purpose of trying to be privacy focused if your search queries aren't mixed with other people's queries.
- y4mi 9y agoI can't grasp how you got that idea. Do you not know what self-hosted means? the engine craws the web and saves its data locally. this locally saved data can be queried/searched. So yes, in your search engine, there will only be your own searches. But these searches are only visible to your own servers/services.
- jbg_ 9y agoThis is not how searx works.
- halflings 9y ago> I can't grasp how you got that idea. By reading the link? > the engine craws the web and saves its data locally ...Saves an index of the whole web, locally? This is not what SearX does. It queries other search engines.
- y4mi 9y agoyes, i mixed up the engines and commented without verifying which it was. i'm sorry for that. i just thought it wasn't necessary to edit as somebody sufficiently pointed out how mistaken i was 9 days ago.
- jbg_ 9y agoI also run searx self-hosted, configured to proxy all its queries through Tor. Occasionally one of the engines doesn't return results (probably due to blocking), which is barely noticeable since several others still work, but normally all the engines including Google return results. Since searx doesn't store cookies returned by the search engines, and I'm using it through Tor, I think this is a significant improvement over sending all my search queries to Google directly from my laptop.
- snowpanda 9y agoBuilding on that issue, I'd like to add that it would be nice to have a feature that alerts a user that certain a search engine is denying requests. It's visible in the logs or settings somewhere, but usually I find myself wondering for a while why my search queries aren't accurate before heading off to figure out why. Still a great project though, I use it every day.
- jbg_ 9y agoAt least for me, next to each result is a list of the engines that returned that result. I run searx through Tor, so I occasionally find that Google stops returning results for a few minutes. It doesn't happen often, but it's easy to tell when it does because none of the first page results have "google" next to them, while of course normally most of them would.
- iuguy 9y agoI've been running multiple SearX instances for goodness knows how long and this has never happened to me. I'm not aware of this happening either. SearX uses multiple sources for queries, so you'd have to be banned by quite a few search engines to stop it being useful. Also relevant is filtron[1], an application firewall built-in to SearX that rate limits searches. [1] - https://asciimoo.github.io/searx/admin/filtron.html https://asciimoo.github.io/searx/admin/filtron.html
- saas_co_de 9y agogoogle will give you a captcha every once in a while but they never actually stop you from using their service.
- userbinator 9y agoIt will also sometimes ban you completely (not even the CAPTCHA works, solving it just gets you another) for ~2h. I've triggered it manually, usually when trying very specific queries and multiple variations in quick succession and also going through to the "end" of the result pages.
- dajohnson89 9y agogetting banned in that case must've been extremely aggravating.
- jazoom 9y agoI'm curious. How does DuckDuckGo do this?
- amelius 9y agoI wouldn't worry about it. If they get banned, they will probably apply some ML technique to circumvent any CAPTCHAs to get access again. Also, this can run from the user's computer so it would actually be quite hard to detect that the results are being aggregated, and stripped from ads.