11 ms·
Show HN: No Trash Search
- etchalon 5y agoThis is the approach I imagine Apple would take if they were to ever launch a Search Engine. A large corpus of handpicked sites.
- deleted 5y ago[deleted]
- rickdeveloper 5y agoI built this website a couple of months ago because I was annoyed by how hard it was to find useful things on Google. As "Google no longer producing high quality search results in significant categories" [0] is currently #1 on the front page I figured I'd share this project again. I hope it's useful to some people. 'No Trash Search' is very focussed on STEM and not "for daily use". It's surprisingly good when you're looking for certain kinds of information. Under the hood it's little more than a programmable search engine [1] with a whitelist of ~120 sites. [0] https://news.ycombinator.com/item?id=29772136 https://news.ycombinator.com/item?id=29772136 [1] http://programmablesearchengine.google.com http://programmablesearchengine.google.com
- BlueTemplar 5y agoWhile I can understand the appeal, restricting your search engine to only ~120 websites out of hundreds of millions (?) is basically giving up on the Web. (BTW, any good search engines these days that aren't indirectly using Google or Bing ?)
- narrator 5y agoThere's http://yandex.com http://yandex.com . It's great if you want to search controversial subject matter and controversial results that Google wouldn't give you. The reverse image search is also amazing.
- quocanh 5y agoWhich results are different than Google's?
- jhugo 5y agoMost. Yandex is great, especially for programming searches. It generally ranks GitHub, Stack Overflow and other content-heavy sites highly. Google has been taken over by weird clones of GitHub and SO lately, Yandex has no such trash. It completely boggles my mind that the useless GitHub and SO clones rank first page on Google. Do engineers at Google not use their own product?
- skinkestek 5y agoRegarding stackoverflow there is a fair chance they can congratulate themselves: If I am right they played stupid games and won stupid prizes. More specifically they have allowed rampant deletionism for years so while I am fairly certain the questions and answers originated on Stack Overflow it wouldn't surprise me if a good number of of those aren't visible on Stack Overflow anymore which would explain why they rank higher in Google. Done right this would actually be a service. Sadly some of them seems to mix together various questions and answers in the same page to generate text matches for unusual queries.
- ramphastidae 5y agoEngineers at Google build what the ads and sales teams tell them to.
- imglorp 5y ago> Google has been taken over by weird clones of GitHub and SO lately Do you have an example search leading to a GitHub clone?
- vgalin 5y agoFrench is my mother tongue, but I've quickly learned during my studies that using English keywords in my STEM-related searches would simply lead me to better (and more abundant) results. A few weeks/months ago however, while I was trying to solve an issue whith a colleague who would search using french keywords, I noticed that some websites featured on the first page of the Google results were off. In short, they were machine-translated versions of Stack Overflow threads. And they would appear in most of the searches using french keywords. Those websites also appeared rarely in my searches while I was using English keywords, but most of the time I never bothered opening them. But now I notice them every time. Some examples: When searching for "wget set http proxy" on Google, the fourth result leads me to qastack.fr, and the ninth to it-swarm-fr.com, both are websites featuring scrapped and machine-translated threads from Stack Overflow. When searching deliberately in french for "Eclipse CDT stdout ne s'affiche pas" ("Eclipse CDT stdout not displayed [in console]"), the first result leads me to askcodez.com and the fourth one to qastack.fr (askodez is the same as the other two). I have never stumbled upon Github clones, yet, however.
- version_five 5y ago> restricting your search engine to only ~120 websites out of hundreds of millions (?) is basically giving up on the Web. Sure - the web is now a cesspool optimized for advertising and attention. The traditional search engines made a lot more sense at the dawn of the internet when it was more about discovery. Now, for the most part, it's closer to an information retrieval tool, where a finite list of established sites have the bulk of what one is looking for. It only makes sense to have a tool that lets one navigate the established, legit internet, and not have to deal with all the crap. That doesn't mean there is no use case for google as it is, but some more focused competition is a no brainer.
- deleted 5y ago[deleted]
- 1vuio0pswjnm7 5y ago"(BTW, any good search engines these days that aren't indirectly using Google or Bing ?)" The code for Gigablast is open-source, including the crawler. I could be wrong but I do not think search.marginalia.eu nor wiby.me use Google or Bing. The comment about "hundreds of millions" is interesting. Assume hypothetically a search engline claimed to be searching millions of sites for a given query but in truth it was actually only searching 120 sites that it had determined answered this query (i.e., was the most popular answer source) for the majority of users. How would a user verify the search engine's claim about searching millions of sites was true. What if the search engine only allowed the user to retrieve a maxmimum of about 230 results, not matter how many sites it claimed to search.
- imachine1980_ 5y agoGigablast resource tend to be full of trash in my short experience whit it
- 1vuio0pswjnm7 5y agoAll the search engines have trash. I retrieve results from a variety of search engines and mix them into a simplified SERP with zero cruft that can be read very quickly. Some call searching multiple search engines "meta-search". The main differences with mine is 1. it is all done client side (there is no remote "meta-search" engine) and 2. searches can be "continued" where they left off at any time. This allows one to avoid rate limits. There are always trash results, every search engine has them in their SERPs, but I find that the more results and the more varied the results the better the chance of finding useful, non-trash ones. Gigablast allows returning at least 100 results at a time. Few search engines allow 100 results at a time that anymore. Google still allows it but will not allow a user to retrieve more than 200-something results total.
- jerf 5y ago"How would a user verify the search engine's claim about searching millions of sites was true." Search for things specifically on those pages, by very specific phrases and such. Of course you have to find them yourself first for that verification. I can say having set up some very teeny tiny websites here and there that the googlebot is hooked up to a lot of stuff. I'm not even sure how it found a couple of them as quickly as it did. Things like "if someone adds an RSS feed to Feed.ly" seem to do the trick. None of them were sites trying to "hide" or anything and I expected them to be found eventually, but they got found much faster than I expected. Or maybe they just scan new domain registrations, though it seemed to me it wasn't that that triggered it.
- fuckcensorship 5y agoCheck out marginalia[1], made by another user on HN. [1]: https://search.marginalia.nu/ https://search.marginalia.nu/
- marginalia_nu 5y agoYeah I do my own crawling, and offer results from around 200k sites (although it's indexed 700k domains, most of which are crap).
- ColinHayhurst 5y agoTry Mojeek https://blog.mojeek.com/2021/03/to-track-or-not-to-track.html https://blog.mojeek.com/2021/03/to-track-or-not-to-track.htm... Disclosure: team member. Feedback good or bad appreciated
- karlzt 5y agoFeedback: https://news.ycombinator.com/item?id=29786112 https://news.ycombinator.com/item?id=29786112
- blaerk 5y agoI think https://www.qwant.com/ https://www.qwant.com/ use their own, just started using it so I can't really say much about it other than it seems alright compared to ddg and google(?)
- BlueTemplar 5y agoLast time I checked, it just used an old index from Bing ?
- DantesKite 5y agoHey I was looking for something like this. Thanks.
- throwawayboise 5y ago> Under the hood it's little more than a programmable search engine [1] with a whitelist of ~120 sites So back to what web search was in the 1990s, roughly: an index from a curated selection of sites.
- rdiddly 5y ago120 sites is pretty hilarious and sad. "Here you go, the worthwhile part of the internet!"
- dataflow 5y agoYou might want to add cppreference.com to your list of programming sites.
- lionkor 5y agoSeems to be in there now :)
- SilasX 5y agoFYI, I think this is just the case where you should prefix the submission title with “Show HN:”. Can mods update it so it shows with the others? @dang? https://news.ycombinator.com/show https://news.ycombinator.com/show https://news.ycombinator.com/showhn.html https://news.ycombinator.com/showhn.html
- bruhhh 5y agothe github repo is just few html pages i dont see anh source code.. i wont trust this search engine...
- vizzah 5y agoNothing should be in the repo. There is no source. Check author's link [1].
- rickdeveloper 5y agoGlad you're interested in the source code! As explained in my other comment, this website is a wrapper around google programmable search. The actual searching happens on Google servers, and I can see why people have problems with that. The code you see on the website is the same as the repo, though. It's actually hosted by GitHub! You can verify this by opening the web inspector in any browser or looking at the `.github.io` portion of the URL. You can learn more about Programmable Search here: https://developers.google.com/custom-search/docs/overview https://developers.google.com/custom-search/docs/overview. NoTrashSearch uses the 'Programmable Search Element Control API', which is documented here: https://developers.google.com/custom-search/docs/element https://developers.google.com/custom-search/docs/element and can be used with very little code!
- version_five 5y agoI think your site is great, I've thought about something like this before but didn't realize how simply it could be implemented. Stupid question though: where is the list of whitelisted sites? Is that something you set up separately with google? I scanned though the code and expected to find a list somewhere, but obviously you do it in a different way
- rickdeveloper 5y agoThanks! Yes, it's configurable in a private dashboard. I created a pastebin with all sites at the time of writing this comment if that's helpful: https://pastebin.com/qLC0wQ0t https://pastebin.com/qLC0wQ0t. If you're looking to create your own search engine, go here: https://programmablesearchengine.google.com/cse/all https://programmablesearchengine.google.com/cse/all.
- razemio 5y agoGoogled "Best smartphone 2021" which resulted in crappy result. Maybe I am missing what significant categories actually means?
- rickdeveloper 5y ago(obviously: this is subjective, so what's significant to me may not be to you.) Honestly I just created this search engine for myself to find things more easily while programming or studying (I study biology, cs and ai; and philosophy in my free time, so expect results the best results for queries related to those subjects). I think those subjects also appeal to the HN audience, that's why I shared it here. When I'm not doing those things, I just use Google or DDG because they have better results for day-to-day queries. That being said, I'm definitely interested in helping improve other people's search as well (reason I'm posting at all), so let me know if you have suggestions for sites to add!
- razemio 5y agoThanks for sharing. My initial comment sounded a bit harsh. I am sorry for that. I will look into fine tuning it for my needs. It is an interesting approach to a very annoying problem.
- rickdeveloper 5y agoNP :) It's exciting to have people using & commenting on something I build. > I will look into fine tuning it for my needs. It is an interesting approach to a very annoying problem. If a premium version were available with a customizable whitelist, would you pay for that? The API is around $5 / 1000 searches so it would cost about the same.
- razemio 5y agoYes, a product which cleans my search from most referral sites is something I would pay for.
- 5y ago
- baal80spam 5y agoFirst test search "python random" - returned just what I would expect instead of multitude of low-quality blogs like Google Search does. +1 from me!
- oburb 5y agotry "python.org random"
- thekyle 5y agoWhen I search for "python random" on Google, I get: https://docs.python.org/3/library/random.html https://docs.python.org/3/library/random.html as the first result, which is what I would expect.
- mda 5y agoGoogle easily finds correct results for "python random" no spam no bs.
- throw_me_up 5y agoGreat job! How are the ads implemented and do they cover the costs? I'm thinking of building a similar search engine for a completely different domain. I am a bit concerned about paying for it though.
- tyingq 5y agoIt appears to be google's custom search that you use to embed search on your own site. https://programmablesearchengine.google.com/about/ https://programmablesearchengine.google.com/about/
- rickdeveloper 5y agoThanks! Ads are added automatically by Google. The whole thing is little more than a wrapper around the 'Programmable Search Element Control API' which is an HTML element you can just insert into any site, like an iframe. Unfortunately this is the only way to make Programmable Search available at scale as the API is restricted to either 10 sites or 10K queries / day, even when paid! There is a paid version for the HTML plugin, but that would leak the API key and so it wouldn't work as a business. There is an option to get a share of the revenue generated by a search engine. Maybe it's time for me to figure out how that works. I was thinking of making a hosted, ad free, customizable version where people upload their own keys. Not sure if people would like that. As a side-note, it's super easy to remove ads with 1 line of CSS, but I wasn't sure how Google would feel about that so it's not in the online version. TamperMonkey is an extension that allows people to insert their own CSS on different websites. Hmm. You can view all offerings in the docs [0]. [0] https://developers.google.com/custom-search/docs/overview#summary_of_programmable_search_engine_offerings https://developers.google.com/custom-search/docs/overview#su...
- yanmaani 5y agoHow about using the Bing API? Isn't that more open? With caching, I think you might be able to reduce the load. Also, why is w3fools in the list? It's an awful site.
- 5y ago
- pacificmint 5y agoI need to buy a 3 hole punch, and when searching for reviews yesterday I had the same problem of lots of hits with affiliate links and low quality sites. I searched for “3 hole punch review” [1] here, and the results have zero relevancy. First one is a Chinese cell phone company, second a Wikipedia page for an episode of the office, third a thesaurus page with synonyms for ‘colorful’ and fourth a link to the Wikipedia page of Yellow Submarine. I can’t even imagine how you get there from “3 hole punch review” [1] https://notrashsearch.github.io/?q=3+hole+punch+review https://notrashsearch.github.io/?q=3+hole+punch+review
- babalulu 5y agoAnd as of this minute, your comment is the number one result for that search.
- rickdeveloper 5y ago"It's a feature not a bug" (TM) The site uses a whitelist of URLs to (attempt to) keep results relevant to science and programming. In the context in which I'm using this search engine, I have no interest in (reviews on) 3 hole punches. (That's not to say I never do, but in that case I'd use Google, Reddit, etc.) The fact that results don't show up here means that they also won't show up when I'm not looking for them, which is 100% of the time when I'm using this search engine. That's a plus for me personally. Best case would be to have relevant results in a single search engine, but that's not what I intended when building this site.
- cortesoft 5y agoI think your name might need a rework, then. The name of the site doesn’t match what it is doing.
- hallway_monitor 5y agoGuess we better rename eBay and Amazon while we're at it.
- jaytaylor 5y ago
- Supermancho 5y agoI did a search for "eternal crusade". The blob of the ads are still the top results. This is not the "no trash search" I'm looking for.
- UltimateFloofy 5y agoso you've used google programmable search to make googling better?
- rank0 5y agoThe first two results for “Rust awesome” were two ads which took up my entire screen on mobile. Both travel related. The GitHub repo was third and had to be scrolled to. Seems pretty trash to me.
- dzhiurgis 5y agoSo many listed their search engines that I feel at this point we need an aggregater
- GlitchMr 5y agoPretty nifty project. I'm curious what whitelist are you using, would it be possible to list allowed domains somewhere publicly. One suggestion that I have is to remove w3schools.com from the whitelist. MDN is a much better source for information about web development.
- rickdeveloper 5y agoHere's a pastebin of the list I made yesterday [0] in a format that allows uploading to your own instance of programmable search. I have updated a few things since (like removing w3schools :)). [0] https://pastebin.com/qLC0wQ0t https://pastebin.com/qLC0wQ0t
- schleck8 5y agoIf you are annoyed by Google, take a look at kagi and neeva, they are new takes on search engines - https://neeva.com https://neeva.com - https://kagi.com https://kagi.com
- twofornone 5y agoWhy does neeva require an email sign in?
- schleck8 5y agoThey have a waitlist, it's a beta
- skinkestek 5y agoHaven't tried Neeva but I can vouch for Kagi. Also search.marginalia.nu puts a smile on my face almost every time I use it :-) (I should try Neeva, I keep hearing good things about it.)
- llaolleh 5y agoI agree with this approach. You can't crawl everything from the getgo - focus on a very specific domain catered to a specific set of users.
- quickthrower2 5y agoThis is good, but also sad in a way. It means Jo Blogg's blog won't get discovered and they may have had some good information on the topic. One way to improve is a "bring your own list" feature, and the ability to include vetted lists. Maybe some kind of web of trust - if your friends have whitelisted a site, it is whitelisted for you too. If you find a problem with that site, you let your friend know to remove it. If they don't respond you can remove that friend from your trusted persopn list (maybe they got hacked?). Then maybe you can 'follow' a few lists of famous trusted people (e.g. paulg etc.) to build up a bigger slice of the internet you can search. A spammer will want to come in then and create something that white lists their spam sites, but they need to convince you to add their list! And when you see the spam you can just unfollow them. They can't succeed.