37 ms·
Search engines and SEO spam
- cblconfederate 5y agoGood. In fact, if we want people to visit websites other than google.com (and then read the answer in the snippet or the box in the sidebar) then it's good that google results are crap. Use google less.
- digitcatphd 5y agoMost of their complaints are not related to Google Ads, which means the poor results are not there because of profit motives. Moreover, they are more related to a specific type of search query, that likely a result of broad based ranking algorithms that loosely are the most efficient ranking system.
- new_here 5y ago> Maybe ultimately you open up spam fighting to your users. If you managed this well, you could harness a lot of energy. Doesn't Google already consider that if a user returns to the results page (or clicks a second link) then the first link visited was not satisfactory. Seems like a pretty elegant solution.
- tonyedgecombe 5y agoThat's to Google's benefit though, they get another chance to present some adverts to the user.
- james-redwood 5y agowww.neeva.com www.kagi.com Two privacy oriented search engines with results and features better than and surpassing Google (did I mention that they’re ad free?)
- mg 5y agoWhy not try writing a search engine specifically for some category dominated by SEO spam? I like to compare search engine results and wrote this tool to make it easy: https://www.gnod.com/search https://www.gnod.com/search There in fact are many vertical search engines. You can click on "more engines" to see the whole list.
- Debug_Overload 5y agoThe fact that you included Reddit, SO, Google Scholar etc is awesome (I thought it was only for main search engines). Thanks for sharing. Bookmarked.
- noduerme 5y agoOkay that's sort of what the !bang in DDG is for, and why it's a meta-engine. What's the blue sky ideal for a real, no-bullshit, everything search engine that doesn't fall prey to the constant flood of garbage? I have an exterminator who comes to my house every couple months, and sets up traps here, poison there. I don't have any rats in my house. I do see rats running across the yard sometimes. The exterminator explains it like this: Rat pressure. The rats overpopulate and there's "pressure" (like, uh, "memory pressure", which is also a fluid concept) so they try to get into your house more, through smaller holes, as a function of how much outside drama is going on, how many they are and how overpopulated, how scarce their food supply is, how cold it is outside, and whatever else drives rats into your house. (I love the dude who's my exterminator). Anyway, this is the same problem every search engine faces. The more surface area they expose, the more pressure they have building, the more ways people have to fake out their systems. We have to go back to the 1990s Yahoo! model. Curated content. A list of websites that are reputable. 1990s Yahoo is the future.
- loceng 5y agoCreating a "trustless" search crawler, where anybody can participate, and then applying an algorithm to determine trust or value feels like it'd be a never-ending arms race - that'd require AI and extensive/expensive resources that is likely better invested in developing real trust networks and curation; curators are corruptible and regulatory capture of policy is possible if the organization is infiltrated or poorly overseen. Carte blanche opening your system up to anyone to inject data seems like the wrong foot to start off on, whereas my curating a moderator, someone I personally know and feel good about, trust to whatever level, and hiring them - ideally making sure they're someone you respect and you're someone they respect, pay them well, and at scale will be able to pay for itself; this did just bring to mind however big pharma and pharmaceutical trials structure and how that system can be/is/has been captured - and so perhaps the pressures when dealing with multi-billion dollar market categories will always lead to shenanigans if ever trying to centralize too much, not allowing for de-risking and broader resource distribution via sales/profits to more parties than the "5-star" rated products. In a thread on HN, I think it was yesterday, a few people posted about review sites where some product reviews are free - but others you had to pay for. A system to facilitate such organizations could allow a highly competitive environment, where organizations develop/build a brand - build trust for their brand as being competent and thorough - and so then over the lifetime of a reader/customer, perhaps they'll spend $1,000 buying reviews (say 333 big purchases over 40 years that you're willing to pay $3 a hit for) to make sure they're ; mind you there will be organized that could be captured to say promote one conglomerate of products over another, perhaps even regionally, but I'm beginning to think it's a necessary layer to combat the shit show that is Amazon (et al) reviews. Ideally these systems and how the reviews present the information, and how thorough - the technical depth and breadth and testing done - will help educate those who dive into using this system, which will sharpen themselves while keeping reviewers on their toes and arguably strengthening their organizations and competency as well.
- SavantIdiot 5y ago> And boy would Google find it hard to follow you down that road. This is a good perspective. Where can Google not go? Places that don't lead to profit. They will try (cough Wave cough) but will give up.
- stanleydrew 5y agoBut isn't profit the point of any company? If you're going down a path that doesn't lead to profit, you'll fail whether Google follows you or not.
- SavantIdiot 5y agoNo. That's why Non-Profits, 501(c) corps in the US, exist. E.g., Linus wasn't looking for profit, and Linux ate the world.
- coding123 5y ago> Lots of people want to be amateur police. (pg) This is very true. How many times have I clicked on a site met with ads so bad that the browser slows down, and after 10 seconds the page gets covered up by more and more crap and then a paywall shows up sometimes too. Now here's the thing - a competitor to Google might detect you clicking back and then pop-up a special set of controls near the search result that lets you say: "too many ads" or "paywall". However, if such an engine were to start beating Google, I'm sure Google would implement it in their own way: automatically detect why you clicked back in such a short timespan.
- DenisM 5y agoSooner or later you’ll have to deal with ballot-stuffing - companies trying to bury their competitors by casting lots of negative votes. Perhaps ML will help in detecting such campaigns.
- cinntaile 5y ago> However, if such an engine were to start beating Google, I'm sure Google would implement it in their own way: automatically detect why you clicked back in such a short timespan. Do you seriously believe that Google doesn't use that as a datapoint already?
- numpad0 5y agoGoogle already detects immediate returns and knock off that link for you. What's problematic to me is I tend to reactively mash back and forward and link just goes from it.
- birken 5y agoThe funny thing is that if the people who worked on spam at Google were free to talk about it, I'm sure it would become evident that they know more about spam and anti-spam efforts than anybody else in existence. It's a ridiculously hard problem, especially when people are targeting you directly. But they aren't free to talk about it, because if they did it would just give more assistance to the spammers, and make the problem worse. I'm not saying that curated search results for particular verticals is a terrible idea (though I'm sure like anything the devil is in the details), but on the whole Google search is very, very good considering the constant assault they are under from spammers (which most other search engines are not, at least directly).
- stoicjumbotron 5y agoHighly OT, but if a technical person (not at a managerial level) involved in tackling spam at Google were to leave the team, are they allowed to work on the similar problem space at a different company?
- goatherders 5y agoThis x100000. There is no scenario - none - where thousands of engineers at Google working on search wake up in the morning and say "we sure have made it good enough wr2 SPAM. I think I'll have another Danish."
- compiler-guy 5y agoWhen the cafes were open, you can bet they said, "I'll have another Danish, and then get back to work on this problem that never seems to go away."
- techdragon 5y agoI sure wish I had problems that were totally unsolvable, they are so easy to measure progress on. /sarcasm I think it’s more likely that because they are just building hundreds of tiny tweak experiments and it’s someone else who desides what to build and if it even worked. Search quality is such a meta-problem that it goes beyond any real hope of simply working on it in anything beyond piecemeal trial and error fashion on their dataset.
- gomox 5y agoI believe that the only moat protecting $100B of AdWords revenue is the quality of the Google Search results. There is no meaningful switching cost to using a new search engine, and the spend inertia in ad spend is not very significant (i.e. any online marketing manager will happily spend 5% of their budget in a different search engine adwords-like program if they get better ROI, there is no incentive the be a "Google Ads only shop"). On the other hand, Google needs to maintain the ballistic trajectory of its revenue growth. So how can they fix search quality when they've minmax'd themselves into this situation in the first place? If they were to make the ads background yellow again, that would have negative short term effects that I doubt any career exec can stomach.
- dredmorbius 5y agoNo. The other moats are lock-in to advertising networks, website metrics, and effective control over Web standards through the Chrome browser. An alternative search platform might provide better search. It would be fighting Google on at least three other fronts. It might have some success, but it would be challenging. (As history largely demonstrates.) Even a rival tech monopolist, Microsoft, barely holds even with its own search offering (I use that indirectly via DDG), and scrapped its own web-browser development
- gomox 5y agoI mean, Google is big and it has that advantage like any big enterprise, but the search engine market is very permeable compared to what people traditionally refer to as a moat (say, trying to compete with YouTube as a video hosting platform or with Salesforce as a CRM). If you have a good search engine, people will flock to it, and search ads will be valuable. That's it. That's how Google became Google. The fact that Microsoft couldn't do it honestly doesn't mean much. Microsoft also couldn't do a phone OS, a portable music player, and many other things. They have a complex web of conflicting interests that $SEARCH_ENGINE_STARTUP does not.
- dredmorbius 5y agoJust for the record, I'm in at least mild agreement that search is looking increasingly vulnerable. Google are falling down here. It's just that "search" is really a web of interrelated services, capabilities, and revenue streams, and they tend to reinforce each other strongly. I'd like to see the monopoly disrupted.[1] But I don't think it's just a matter of "build a better mousetrap^Wsearch engine." Attack one corner, and Google will snipe at you from the others. And with the AdWords cash cow, they've got an immense revenue stream. ________________________________ Notes: 1. Well, mostly. Google's acquired so goddamned much personal data that the premise is frankly kind of terrifying as well --- a weakened Google with neither the revenues nor talent to defend that pile.... And I'd really like to see the toppling occur without simply raising a new monopoly in its place.
- fuckcensorship 5y agoGo read any default subreddit on Reddit to see what this idea would look like long-term, especially the "amateur police" part.
- baby 5y agoSearching code is also impossible on Google. If there’s a competing search engine for that I’ll use it at least for this use case.
- darinf 5y agoGive Neeva a try. We have improved ranking and some nice features around tech queries.
- baby 5y agohot take: why would I need to enter my email to do a search online? You already lost me :o
- aghilmort 5y ago* we're adding a code finder as a topic search engine on Breeze * publicwww.com is also a good source code search engine
- hammock 5y ago>This may not just be a problem with Google but possibly also the recipe for beating Google. A startup usually has to start with a niche market. Why not try writing a search engine specifically for some category dominated by SEO spam? >You might need to do a lot of manual spam fighting initially. That could be both the thing-that-doesn't-scale, and the thing that differentiates you by being alien to Google's DNA. (They must hate manual interventions; so inelegant). Is he describing...Yahoo circa 1994? A manually curated directory service.
- tonyedgecombe 5y agoI'm starting to think Yahoo circa 1994 might be better than Google today.
- behnamoh 5y agoI wouldn't just complain about Google. Google search results mostly reflect a deeper problem with the web today. I do miss the simplicity of the 2000s.
- edoceo 5y agodmoz! https://en.m.wikipedia.org/wiki/DMOZ https://en.m.wikipedia.org/wiki/DMOZ
- matt_heimer 5y agoI always used DMOZ more than Yahoo! Directory. It looks like [dmoz](https://en.wikipedia.org/wiki/DMOZ https://en.wikipedia.org/wiki/DMOZ) became https://curlie.org/ https://curlie.org/ which is still active.
- loceng 5y agoAnd makes me think that StumbleUpon had a similar curation ability, in that the value qualifier is how often [hopefully] real people interact with content - tracked by who's using SU and agreed to allow tracking; can't remember if sharing that was optional or not? The gamification of the system then would have to come through onboarding fake users, pretending/mimicking real user behaviour to send that signal into the system; not sure if SU ever ran into that problem or was actively paying attention to trying to identify and removing fake or suspicious signals from their output? I feel a much better system is easily within reach, it's simply getting the right structure to it, the right foundation, and then it will quickly take off due to the quality difference. I've already figured out a design pattern that Twitter and Facebook has indoctrinated us with, making us think it is normal - and keeping us blind to an actual normal way or organizing or communicating, but that isn't conducive to control or ad revenues - and so extending my future plans to include a better search-directory system would fit snugly into my efforts.
- curiousllama 5y agoGood idea. You could start with fitness. Lots of high-quality information out there that’s entirely, 100% inaccessible via google. Over COVID, I did the whole fitness thing from a few different angles (overhauled diet, trained for a marathon, now lifting weights a lot). I found I could only find good info by going directly to a trusted source - literally, typing http://www http://www. like I’m in the 90s or something. This is the exact issue a search engine should solve, but Google doesn’t.
- paulcole 5y agoIs your trusted source the same as my trusted source which is the same as my neighbor’s trusted source which is the same as my Australian cousin’s trusted source? If not, at least one of us is going to hate this new search engine.
- thewarrior 5y agoCould you share your trusted sources ? Human search engine :P
- ok123456 5y agohttps://www.boards2go.com/boards/board.cgi?user=tfannon https://www.boards2go.com/boards/board.cgi?user=tfannon
- curiousllama 5y agoJPG on Tik Tok, and Geoffrey verity schofield on instagram/quora/buy his book.
- nefitty 5y agoThis is a problem I'm working on. What sorts of unique things did you do that Google failed at? Maybe you read through discussion sites or got tips from books or something like that
- curiousllama 5y agoSocial media and trial and error. I knew a bit, and asked some friends for advice. I used that to find people who said stuff I thought made sense on Quora, Tik Tok, and Instagram (ie they agreed with what I knew to be true and false, so I could assume the other stuff they said was likely to be true as well). I tried what they said, found what worked, and went all in when I saw results. Importantly, this was bottom up: it was largely recommendation engines suggesting people I then filtered through for what I wanted (running, bodybuilding) vs didn’t (traditional weight loss). I couldn’t specify what I wanted, or it would be garbage SEO spam.
- Jenk 5y agoSingle-page thread: https://threadreaderapp.com/thread/1477760548787920901.html https://threadreaderapp.com/thread/1477760548787920901.html
- hooande 5y agoWhat are some search categories that are so dominated by spam that they are unusable? I'll start: "how to rent a car" [0] [0] worth noting that I personally get somewhat reasonable results for this, with a 3rd result from nerdwallet.com and a 4th from wikihow.com, both of which seem to answer the question in an unbiased way
- rickdeveloper 5y agoI think a lot of this is due to Google both owning search and the ads on the websites (AdSense). There’s an incentive for them to prioritize click farms (and other sites filled with their ads). I think in general there may be a correlation between the number of ads on a site and its usefulness to me, which is inverse to its usefulness to google. I’m curious what would happen if those products were split up into 2 separate companies.
- thomasmarcelis 5y agoI also can't help but wonder this. I'm certain people at google search want to provide the best quality search results and do this with integrity. But at some point in the business hierarchy you are at a level where people set objectives for both these departments ( search & ads ) and are trying to optimise for things like total revenue/profit.
- tonyedgecombe 5y agoYes. In fact if you wanted to cut out the SEO spam then delisting anything with Adsense would probably be a good start for a competitor.
- Jenk 5y ago> What would a paid version of Google Search results look like - where Google can just try to give me the best possible results and not be worried about generating revenue? God please no. YouTube premium shows what Google would do, i.e., they would further ruin the free experience by ramping up the amount of ads you see to "incentivize" the premium search.
- judge2020 5y agoPremium offerings like that are amazing simply for the fact that you can 'return' to the days where you obtained services by paying for them directly, not by looking at ads and paying with your mindshare. Google and YT aren't free services, and it's a miracle they continue to be accessible with ad blockers enabled.
- Jenk 5y agoOrthogonal to my point. Deliberately worsening the free experience after you introduced a premium service is a dark pattern for UX.
- NmAmDa 5y agoOne example, the website called gitmemory which crawls github data regularly and have better SEO than github that usually you will find results above original github links.
- snth 5y agoSeveral people mention DuckDuckGo in that Twitter thread. I use DuckDuckGo for my main search engine, and it's not obviously any better than Google regarding SEO spam.
- Kiro 5y agoI don't think they apply much spam fighting to the results they get from the underlying search index (Bing), but I could be wrong.
- deleted 5y ago[deleted]
- tester756 5y agoHow about Bing? Is it viable competition?
- ape4 5y agoOne approach would be to have moderators from the community who are allowed to make decisions about results.
- _xnmw 5y agoInterestingly, Google Maps doesn't suffer as much from the issues with Google Search. Maybe because it has those community-driven curation features that PG is talking about? Google Maps is fantastic at finding places to go to (and getting you there). Also, why hasn't Apple built a search engine yet? It baffles me that they chose to go head-to-head with Google on Maps, yet outsourced their search engine. I would've liked it the other way around: Google Maps and Apple Search.
- nostromo 5y agoApple makes 10+ billion a year from Google by setting It to the default search engine on iPhone. I agree they should make their own search engine. But currently they’re being paid a ton of money not to.
- techdragon 5y agoIt’s much harder to get SEO spam style content to maps given the geographic region limits involved… But it happens my favourite example is searching for suppliers of something basic, say structural aluminium extrusions, big and heavy and you ideally don’t want to ship it far so it’s an ideal thing to search for a local supplier of. In Australian results is basically a given that I’ll get results for my city because they will list it as a delivery area they supply to however when you actually try to find them on the map as a pin, nope it’s either not their or just is a sales office not a warehouse or workshop, so they have tricked the system into listing them as local to my area in a way that pollutes my maps search results. But this only works for certain industries. It’s much less common to see this kind of tactic if your searching for say a coffee shop because they sheer number of local results let’s Google be “hyper local” with these kinds of results.
- Nextgrid 5y agoGoogle Maps is spammed by fake locksmith or other trades that make it look like they're local but all route to the same boiler room (probably right next to the tech support or IRS scammers) from where they dispatch a crooked & most likely unlicensed tradesman that will do a poor job and overcharge you (destroying the lock so they can sell you an overpriced replacement instead of picking it, etc). For licensed trades the solution is to go to your official trade licensing body (for the UK it's the NICEIC for electricians and the Gas Safe Register for gas/HVAC technicians), for unlicensed ones it's more difficult. There are "review" sites that claim to provide good results but their business incentives & vulnerability to spam/fake reviews are unknown.
- noduerme 5y agoHey, smart people: It's called CURATION BY HUMANS.
- thr0wawayf00 5y agoI'm honestly not trying to take a potshot against PG or YC here, but it's kinda funny to see him saying this after I worked for a YC-backed startup years ago that built its core revenue streams around generating SEO spam, we just marketed it as something else. Just to be clear, I don't think PG or YC are responsible for all or even most SEO spam, but I know firsthand that they've profited from it through at least one of their incubated companies. I never considered the possibility that an incubator would support a specific product, then later on call for alternatives that would essentially freeze out the original product that they supported. I'm sure this very rarely happens, but it's interesting to see a real-world example in action.
- awillen 5y agoI don't think it's as counterintuitive as it sounds - just because they're playing the game doesn't mean that they think it's right. If your options are to not be successful at SEO or to do the SEO spam thing, I don't think it's necessarily wrong to do the latter - it's your job to make your startup successful, not to make a stand against the way Google does things. I view it as something like the rich folks who call for additional taxation of the rich. They're not going to just pay extra money that they don't have to under the current tax rules, both because it's not particularly fair and because one person paying extra taxes, even if they're very wealthy, isn't going to make a big impact. That doesn't mean they can't lobby to change the rules and be totally fine with it if everyone is paying additional taxes.
- vlovich123 5y agoDon't know how rarely it happens. After all, weapons dealers frequently arm both sides of a conflict. They don't really need to care who wins - they're just making twice the amount of money.
- coffeeroach 5y ago
- Veen 5y ago> Why not try writing a search engine specifically for some category dominated by SEO spam? Back in the olden days, there were lots of organizations that collated high quality content from the best writers. They nurtured expert writers and paid them well. They fact-checked the content and employed diligent editors and proofreaders so it was accurate and well-written. Over the years, they'd build a reputation for reliability and trustworthiness that kept people coming back for more. If you wanted to learn about fitness, or cars, or cooking, or science, you'd find a reputable author and publisher and buy their magazines or books. But then, in the early 2000s, the geniuses from SV "disrupted" the publishing industry and its financial model. They brought us a much better way to find content, the search engine. Because they were so much better than the old-fashioned publishers, search engines gobbled up the advertising money and became the dominant gateway to content. Publishers had to abandon expensive high-quality writing because rankings and eyeballs now mattered more than quality and trustworthiness. Instead of investing in writers, they invested in marketers and SEO specialists. The result: worthless content, writers banging out garbage for peanuts, and useless search engines. Two decades later, looking at the barren wasteland they had created, the SV geniuses thought: I know what we need, more search engines, but smaller ones that collate high-quality content from the best writers. There must be money in that, right?
- dehrmann 5y agoI think you're unfairly putting blame on Silicon Valley. Publishers were only able to produce high-quality content because, with no conversion metrics, advertisers were willing to overpay for placement. Tech undermined publishers' revenue, but what it revealed was that people don't actually want high-quality journalism, they want entertainment, and they're definitely not going to pay a premium for it. This was hidden behind publishers' business model.
- nemothekid 5y ago>Publishers were only able to produce high-quality content because, with no conversion metrics, advertisers were willing to overpay for placement. This implies that big budget advertisers (the CPGs, like Coke and P&G), are buying Google/FB because they have better conversion metrics. That isn't true today; only SMBs and gaming companies care about conversion metrics. There are interns in LA/NY probably collectively spending millions on FB for P&G and only reporting the number of likes back to their bosses. Google and FB has never meaningfully delivered on conversions past anything like app downloads. Tech undermined publisher's revenue because the internet cratered distribution costs. Advertising revenues for big media crashed because the eyeballs moved away, not because it was any less efficient.
- nanna 5y agoIf you were to start a search engine, what stack would you use?
- aghilmort 5y agoGoogle CSEs, Bing API, or our own YaCY instances are more or less what we do atm for various topic search engines as we bootstrap Breeze.
- aantix 5y agoTypesense looks easy to use. But with 3x memory needed for the indexes, the server costs probably aren't going to be bootstrap'able. Especially for a "small" crawl of a billion web pages, event at just 10k per page.
- marginalia_nu 5y agoEh, I run a 100M-index off consumer hardware in my living room. Very doable if you avoid bloated off the shelf solutions.
- aantix 5y agoWhat search software do you run? What sort of memory and space do you have on the single server? What's the average document size that you index? Genuinely curious on how doable a modern search engine is on modern hardware.
- marginalia_nu 5y agoI rum all custom software, I feel most off the shelf solutions aren't very resource effective. The server has 128 Gb RAM and the index currently fits on a single 1 Tb SSD + an Optane drive of 480 Gb. I find the average document to clock in at 7 Kb, in terms of raw HTML. In the index that's, dunno, probably less than 1 KB/doc.
- abakker 5y agoI have a version of fixing this that I would personally enjoy a lot. Leave google alone, let it crawl the web, prioritize what it wants to via algorithms. But, give me a version of that which ONLY surfaces results from discussion forums (including SO, Reddit, HN, etc). For most of the stuff where I am actively searching and not just looking stuff up, discussion forums of motivated, self-selected contributors have the stuff I need with the context I need. It used to be that blogs had answers, but that media has been categorically ruined by SEO. Now, one of the deficiencies here has been examples. Try this: "best miter saw". you will not find any websites that actually discuss the answer to this question, despite it being a product category with a lot of price variability and performance tradeoffs (weight, capacity, power, cord vs cordless, accuracy). Nearly any product reviews for large purchases follow the same pattern unless consumer reports has decided to dig deep (e.g. washing machines). How about guitar strings? Sandpaper? Printers? google's algorithm has allowed profit motivated websites to displace the commons to too great an extent.
- mjr00 5y agoMy current solution for this is to just tag `site:reddit.com` to the beginning of Google searches. A Google search for `site:reddit.com best miter saw` has a lot of relevant results. Marketers/SEO people are starting to infiltrate this as well, but since they can't control and SEO the content on Reddit nearly as much, this still works pretty well for now.
- abakker 5y agoOf course! this is a tip I got from HN years ago. it gets old to add Reddit, PracticalMachinist, Fine Woodworking, etc. I really want a proxy for user generated content where nobody got paid to write it. Fun story: I knew a person who worked for a home building website part time. She got paid to write stories on home renovations. She had never done _anything_ she wrote about. Mostly, she gathered up other blogspam and recycled and rewrote it without citation. Sometimes she went to forums, sometimes reddit, sometimes youtube. But, the universal part of it was that she had to produce 2x pieces of content per week endlessly. Just for a local LA builder. Most of the content wasn't "wrong" but, it also wasn't exactly incisive and didn't include any details that would have been useful. Instead it was just filler. The worst part is that it consistently improved that company's search ranking. Content farming needs to die.
- aronpye 5y agoA lot of the spam results just seem to be copy pasted content. I wonder how difficult it is to compare the main body of text in search results, then say if it is over a 95% match with another site (I.e. it has been copy-pasted), demote it in the search results. If a site generates too many of these demotions then it gets blacklisted from the index.
- nikanj 5y agoHow would you avoid throwing the original site out with the bathwater?
- aronpye 5y agoMaybe try and time stamp the page, presumably the earliest page is the original source. Could also combine it with a site reputation rating or something similar.
- neoneye2 5y agoI have experimented using LSH (Locality Sensitive Hashing) for identifying similar documents, among 50k documents in total. My LSH implementation is here: https://github.com/loda-lang/loda-rust/blob/develop/script/task_instruction_similarity_lsh.rb https://github.com/loda-lang/loda-rust/blob/develop/script/t... Example of the 100 most similar documents: https://github.com/neoneye/loda-identify-similar-programs/blob/main/oeis/014/A014017_similarity_lsh.csv https://github.com/neoneye/loda-identify-similar-programs/bl... There can be false positives, so after LSH then do a more in-depth comparison.
- yuliyp 5y agoWas this linked for the irony of everyone spamming replies advertising their startups which don't solve the problem but kinda-sorta do, resulting in something hard to read and understand?
- floatingatoll 5y agoHe’s just describing Webrings, except in a reactive tense (“filter out spam sites”) rather than a proactive tense (“associate your site with other worthwhile sites”). Google’s ranking algorithm only works when someone is proactively curating, and only SEO spammers do so these days. Reactive curation is not a viable way to manage information. The simplest way to compete with Google is to create a DIY Webrings site that disallows harvesting of data by Google. Charge curators to create a webring, and let curators select three hashtags and a description that represent their list of fifty or fewer sites. Use the revenue to pay a human to curate the list of hashtags, and let users tip a webring curator in gratitude with an Apple Pay button. This is how to make a million dollars, Pinboard-style, out of the ashes of the original curated Yahoo idea and the information structures of hashtagging. It doesn’t work if you allow free-for-all infinite-sized lists, it doesn’t work if you allow free-for-all hashtags, but with clear limits and moderation of tags (instead of webrings), it would thrive. By moderating tags, users can keep the webring they paid for, and SEO rings will be stick out for having no shared network with any other rings, which allows for easier detection and culling of malicious non-participatory actors. Plus, with the curation networks in place, it becomes possible to bubble up rings that have unusual content for positive human moderation activity. I tried to find some good podcast lists yesterday and each site I visited had a really interesting cross-section, but there were so many duplicates. I wish the ring site existed, so that it could remember what it had shown me already, and I could say “show me rings that intersect with this podcast and have something new I haven’t seen before”. That’s where the theory of pagerank and the practice of curation and the capabilities of search align, and given that moderation of hashtags scales very cheaply, is a billion dollar opportunity that Google and Amazon cannot compete with if handled properly. It’s not about trying to get a cut of every visit’s revenue potential. It’s about giving human beings a directory that respects their time and remembers what they’ve seen.
- jart 5y agoIf I understand correctly, you're saying you'd create an exclusive webring, but the rule to joining is that you have to Disallow: Google, Bing, etc. in your robots.txt file. That sounds outrageous, but speaking as a content creator, I wouldn't be giving up much. My blog gets about 3% of its traffic from search engines. I have no idea who these visitors are or what they searched for, since browsers no longer send referral links. If a webring offered me the benefit of positive regularly engaged community, then having my blog part ways with search engines would be a no brainer. That is after all the Facebook model, except tailored for the open web. Believe me when I say that we bloggers are waiting to be rescued.
- krono 5y agoJust a minute ago, I made a small typo in a non-obscure programming-related search term. Showing results for searchterm No results found for searchterm Followed by an unending list of random celebrities I don't know nor care about, businesses I've never been that sell items I have absolutely no use of, and random foreign news articles. Failure to recognise the typo is unexpected but forgivable. But then, rather than helping me with my search, they attempt to distract and lead me away from it - using triggers that you'd think they should have known wouldn't work. I really don't understand how this is even possible, and it's not a rare occurrence.
- cassianoleal 5y agoI'm not very well versed in SEO but isn't this just good old Goodhart's Law? Come up with criteria to determine which websites are "better quality". Measure them, rank them, put the ones that fit the criteria best at the top. On the other side, there's the people promoting their websites. Do what you can to get as close to Google's ideal as possible through whatever means. Profit. At this point the criteria becomes useless for any real quality analysis.
- ttiurani 5y agoIs there a search engine for programming? One that not only searches stackoverflow, github, relevant subreddits and the other big sites, but also finds programming articles in personal blogs? That would be valuable to me.
- bretpiatt 5y agoThis is already happening for a bunch of verticals: Travel - Expedia, Hotels.com, Kayak, etc. Consumer Goods - Amazon, WalMart, EBay, Etsy, etc. Automobile Purchase - Cars.com, Autotrader, etc. Career/Job - Indeed, LinkedIn, etc. As Google continues to lose search volume on these big revenue categories it is going to make spam much more difficult as they are working to sort out long tail spam. Way harder.
- truculent 5y agoExcept almost all of these aren't search engines, but their own walled gardens. You can't search for an item on Ebay and find a link to the product's Amazon page, for example.
- jeffbee 5y agoWow so that's completely opposite of my feelings on this topic. I would never, ever use Expedia for travel search over Google flights/hotels. Google travel is the meta-search engine for this vertical. Expedia etc. are all-in on spam and scams, trying at every opportunity to take an extra dollar from you. The same with Amazon. You'd think that with all the armchair search quality experts cropping up lately there might be more vocal complaints about the fact that Amazon's own search can't find basic consumer products sold by Amazon itself. If I want to find stuff on Amazon, I search Google for it.
- bradyo 5y agoYeah the comments in this thread are baffling (or they didn't read the 100 character tweet lol). The tweet is just describing domain specific database-like websites. Have people not heard of allrecipes.com? Yummly? Or one of the other thousands of recipe db sites? No blogspam, just structured recipe search. You can even search by ingredient! Mayo clinic, Harvard health, and pubmeb do a great job with health info. IMDb for movies, Goodreads for books, *gearlab.com for reviews, booking.com for accomodations. I think the biggest threat to Google isn't a better general search engine, it's user behavior switching to more domain-specific websites as the top of the funnel. E.g. people going directly to Amazon to search for products instead of first searching Google. To some extent, Google has figured this out, which is why they now have a dedicated flight search, hotel search, product search (Google shopping still exists and it's pretty good!), etc.
- ben7799 5y agoIt'd be really interesting if Google allowed upvote/downvote on search results... but it'd be super hard to imagine them every taking the votes into account much versus ad revenue. And the upvote/downvote would be very tricky to implement in a way that the SEO crowd couldn't just game it horribly.
- jeffbee 5y agoClicking a result is essentially an upvote. Immediately returning to the results page is essentially a downvote. You can't really crowdsource this stuff, because the problem of brigading and other forms of abuse is way too high. Just imagine what the crowdsourced results for "trump won" or "trump lost" would look like, or hydroxychloroquine, or ivermectin, or to go with some older cults of personality, Hitler or Ataturk.
- nabla9 5y agoWhen the search engine is funded by ads, there is incentive to produce results that people who click ads like.
- bombcar 5y agoJust give users the ability to blacklist domains when searching; pretty soon you'll have a decent list of what users consider worthless. And pintrest would die.
- slig 5y ago>And pintrest would die. They're friends with the king, so don't hold you breath.
- krono 5y agouBlock Origin static filters to the rescue! Block results from specific domains on Google or DDG: google.*##.g:has(a[href*="thetopsites.com"]) duckduckgo.*##.results > div:has(a[href*="thetopsites.com"]) And it's even possible to target element content with regex with the `:has-text(/regex/)` selector. google.*##.g:has(*:has-text(/bye topic of noninterest/i)) duckduckgo.*##.results > div:has(*:has-text(/bye topic of noninterest/i)) Bonus content: Ever tried getting rid of Medium's obnoxious cookie notification? Just nuke it from orbit: *##body>div:has(div:has-text(/To make Medium work.*Privacy Policy.*Cookie Policy/i))
- li2uR3ce 5y agoGoogle used to have such a feature. It would be nice if Google would ask you the simple question: "Did you find what you're looking for?" Instead they rely on the assumption that users only stop looking when they've found what they're looking for. These days, there's a reasonably high chance that I quit looking because I gave up in futility--not because I found what I was looking for. It's also the case that there's no way to train Google not to omit search terms or generalize them to the point of uselessness. I really wish abusive SEO were the only problem but it's far from the case. Search results being crappy is a cumulative effect. You could solve SEO spam and I'll still not be able to find a USB SuperSpeed cable because it gets generalized to "usb cable" and there are a gazillion more charging cables than there are SuperSpeed cables. Used to be that you could quote things to indicate that you really meant it. That's fuzzy now too. Every time we figure out how to circumvent the bad results, features are removed.
- yumraj 5y agoIf I remember correctly, it used to be called About.com - with categorized and human curated links. It was big during the dot com days, but withered after Google. Interestingly, I do think that that model may need to be revisited. Edit: I feel that Reddit is filling some of this need, at least for things like Vaccuum and Espresso machines with dedicated spaces.
- celestialcheese 5y agoIt's funny you mentioned About.com. Or as it's known today DotDash AKA IAC. It's easily one of the largest SEO players out there, and they've been on a buying spree recently with their purchase of Meredith. The quality of the content has gotten better, but it's still a monster of an SEO optimized content machine.
- yumraj 5y agoYes, I know nothing about what they’re doing today. In fact my pihole blocks them, must be in some list. So I’m pretty sure they’re crap today. I just remember the concept from way back.
- jrockway 5y agoTo some extent, I worry that the problem with search engines is that there isn't any data worth returning. Yesterday's thread talked a lot about reviews. Writing a review is hard work that requires deep domain expertise, experience with similar products, and months of testing. If you want a review for something that came out today, there is no way that work could have been done, so there simply isn't anything to find. Instead you'll get a list of "Best TVs 2021" or whatever, with some blurb and an affiliate link, not an actual review. That's what people can make for free with a day's notice, so if you write a search engine that discards those sites, that's fine, you'll just return "no results" for every interesting query. I guess what I'm saying is that if you want better reviews, you probably want to start writing reviews and figuring out how to sell them for money. Many have tried, few have succeeded. But there probably isn't some Javascript that will fix this problem.
- snth 5y ago> If you want a review for something that came out today, there is no way that work could have been done, so there simply isn't anything to find. That's not strictly true, given that reviewers are often sent pre-release versions of things in order to do that work before release day.
- loceng 5y agoNot sure why you're being downvoted, as you're correct - however to point out there seems to be a trend where reviewers are only given pre-release versions if they practically always give favourable reviews to the products they list, especially if they're provided the product for free; there doesn't have to be an express relationship or contract between a reviewer and a company either, it's the reverse of how Bill Gates apparently has given $200 million+ to different news channels/media organizations - and so they're less likely going to as freely share negative news about him or perhaps his organizations, so then ; this makes me think, similarly to how stocks being sold by CEOs (etc) must be pre-planned to avoid shenanigans like market manipulation, that anyone giving large sums of money to any media/journalism organization must divide the amount up over 20-40+ years, so that organization at least has a runway and not dependant on larger "dopamine hits" at shorter intervals.
- notananthem 5y agoWe still need a search engine that actually blacklists everything serving ads. Google beat altavista, now we need to beat google. I mean no mincing about- recipe sites that are ads are blocked. Results with pixel tracker etc are blocked. Hell, results that are paywalled are blocked because they're useless.
- throwaway14356 5y agoah, so everyone wanted to move from carefully crafted personal websites where every detail counts and low effort publications are harshly punished to platforms with guranteed readership and now we have a curration problem? Someone (who probably doesnt have a website) said that comment moderation on your own website is to much work. Perhaps the whole internet is to much work? But i like the spam search engine by and for spammers as a way of finding the latest and greatest affiliate marketing and blockchain swindle.
- jeffbee 5y agoAll of the amateur search quality experts forget to mention the regulatory environment. Obviously, Google could nuke Pinterest from orbit, dramatically improving image search results. Clearly, Google could effectively take down Statista, technically. But various Eurocracies have shown an extreme willingness to take the side of Yelp, Pinterest, and whatever other spam/scam mills are able to form a shadow alliance with Microsoft's astroturf campaigns like "fairsearch" and whatever.
- wakiza33 5y agobig part of it. and the scrutiny is only going to get more intense
- notreallyserio 5y agoIf this is a real concern they can side step it by simply allowing users to block specific domains, like they have in the past.
- Covzire 5y agoCould this issue be related to Gmail's spam filtering? For approximately 2 years now it's been downright porous, I'm getting on average 1 obvious spam message in my inbox that is something like: c0nGrats-You_HaVe_Won_ThE_Pr1ze! ..Or some silly variation of this that takes literally 0.1 ms for a human to discern that it's spam. Yet something happened to Gmail's spam algorithm in the last couple years that has been consistently letting these through. To be fair, it does catch most spam but it's only batting something like 75% and the spam it does catch is often times much less obvious to human eyes than the stuff it lets through.
- liveoneggs 5y agodo google search engineers use ad blockers?
- dang 5y agoThis was in response to mwseibel's thread, which had a big discussion yesterday: Google no longer producing high quality search results in significant categories - https://news.ycombinator.com/item?id=29772136 https://news.ycombinator.com/item?id=29772136 - Jan 2022 (1167 comments, spread over multiple pages - note the "X more comments" links at the bottom)
- wenbin 5y agoFYI - Google hires 10,000+ search result raters [1], who are contractors, to evaluate search result quality. In an ideal world, you build a thing, and it's done. It runs automatically and prints out money. In reality, you still need human labors to do manual tasks, even in tech industry. [1] https://www.searchenginejournal.com/google-eat/quality-raters-guidelines/ https://www.searchenginejournal.com/google-eat/quality-rater...
- Ros2 5y agoThanks for posting this. An acquaintance of mine did this job 6 years ago and I wasn't sure if it still existed. Crowd-sourced humans are making Google appear more intelligent than they actually are. I always envisioned that spam efforts would just immediately set off an alarm that would be handled by a bot to blacklist you without a human even knowing your site existed, but there still seem to be at least a few ways to game Google's search ratings.
- freediver 5y agoGoogle’s job is to serve its customers, and it does that really, really well. The problems being discussed today (and yesterday in the similar thread) come from the fact that for Google user != customer. When you have incentives that are misaligned like this, you can only go so far! We seem to have reached that point with Google, where there is not much more that can be done on the search experience front without jeopardizing customer experience (ad revenue). Disclosure: I’m working on a paid search engine to solve this problem on a fundamental level, by aligning the incentives and making user also the customer so we can best serve them and their needs. It is called Kagi and is currently in closed beta accepting beta-testers. https://kagi.com https://kagi.com
- netcan 5y agoI'm almost certain that pg gave "compete with Google by competing in some niche" advice 10+ years ago. In any case, I'm not sure that competing in search is a very attractive notion. AdWords is the only meaningfully profitable search and business. Even if you steal 10% of Google's market, that absolutely doesn't translate into 10% of the revenue. That said, recipes. Someone make a search engine where the top results don't start with 500 words on the history & etymology of butter, because that's what Google want.
- colordrops 5y agoIt's usually more than 500 words lol
- deltarholamda 5y agoThe quote-Tweeted thread mentioned recipes as one of the things that has been SEObliterated. It's a great example of the problem, and also a great example of the problems any solution will encounter. Recipes have become a bellwether Internet problem. In the past, your great-grandmother had a card file with a bunch of 3x5 index cards with the ingredients and instructions on how to make everything, and they pretty much all fit on one side. There was a great deal of domain knowledge required (e.g. "whip to stiff peaks"), but these things reveled in their terseness. Internet recipes all begin with 9 paragraphs of the author's first time encoutering the dish in a Moroccan bazaar in 1997, and the life story of the chef. There are two embedded 10-minute videos of the lifecycle of the vanilla bean. And then you get to the ingredient list. Then two more 10-minute videos, then instructions. The drive to make recipes full-contact Internet content has changed what it means to be a recipe. This is similar to how cooking shows evolved from Julia Childs working on a sound stage to a carnival barker presentation with vivid personalities dominating the scene. I'm not sure there is any technological solution to a problem that has fundamentally changed what it means to be a recipe, short of establishing a new informational silo in the form of a new Web site devoted to recipes only. You could encourage an RSS-like format for recipes, but that requires buy-in from places that profit from the new evolution. This new status quo may be good or bad--you can make the argument either way--but it is what it is. A cultural change is required more than tweaking algorithms. (Unless tweaking algorithms can be foundational to cultural change, in which case we really, really, really need to take a hard look at the corporate behemoths and their algorithms, and sooner better than later.)
- foobarian 5y agoThe recipe problem is mostly because actual recipes are not copyrightable. See e.g. https://www.plagiarismtoday.com/2015/03/24/recipes-copyright-and-plagiarism/ https://www.plagiarismtoday.com/2015/03/24/recipes-copyright...
- leobg 5y agoNeither are ideas. Hence article spinning and book summaries. We need a semantic dupe filter: If it doesn’t add new facts or new ideas, treat it as an identical copy.
- cpeterso 5y agoAmazon’s search results and scammy third party sellers are a similar trust problem. When possible I try to purchase directly from the product manufacturer’s website. Similarly, I don’t search Google for product reviews, I go directly to trustworthy review websites.
- 1024core 5y ago... and the moment you gain some traction, the SEO monster will train it's eye on you like Sauron; and without a billion dollar budget, you will be toast.
- deadalus 5y agoI also consider Paywalls to be spam. Clicking on a link and finding out that it is paywalled, is a massive waste of time.
- imranhou 5y agoAgree that is annoying, but if you start excluding such results then how does one find that type of content?
- deadalus 5y agoYou clearly label paywalled content with a symbol or image.
- narmiouh 5y agoI think the only issue there might be that google might be unaware that it is a paywalled content due to how many sites allow crawlers access to content but not to users (based on crawler ip ranges). Agree such a flag would save time when available or even a search filter option to skip those results.
- RichardHeart 5y agoHis suggestion basically is to become DMOZ.org If you are old enough to remember it.
- PaulHoule 5y agoFor medical search the answer is pubmed. Not only is the collection of documents clean (of low-grade scammers, pharma companies have to pay big $ to play) but the NIH has done a large amount of search quality and ontology work -- the system knows "Tylenol" is synonymous with "Paracetamol", "Acetaminophen", etc.
- titzer 5y ago> the system knows "Tylenol" is synonymous with "Paracetamol", "Acetaminophen", etc. This is the exactly the kind of thing that Google cannot fathom manually doing. As if entering facts into a computer were morally wrong somehow. They'd much rather launch the equivalent of a shell script that harnesses face-melting amounts of computational power, processing literally trillions of webpages in bulk, signal and noise together, junk, spam, and misdirection alike, to learn bad associations and then serve them up with no human review and then put the full force of their reputation behind a results page that apparently people never check and certainly can't correct because of the inscrutability of a machine-learned model that has few to no levers to adjust.
- PaulHoule 5y agoGoogle has long lied about what they do. I had a chance to debrief people who had left their relevance team and they told me things that were outright contradictory to what rank-and-file Google employees have told me. (What they told me did make sense in terms of my experience as an IR system developer, SEO publisher, etc.) Microsoft bought a company called PowerSet that had extracted a large database of entities and relationships from Wikipedia and used the technology to make the "Bing" search engine. Earlier Microsoft engines were a joke, but Bing was so good that Google saw it as a threat so they bought Freebase to get a similar kind of database, then they killed it to incorporate it into the "Google Knowledge Graph". For all of their hating on semantics note that they hired R. V. Guha as their chief scientist, who worked with Doug Lenat on the notorious https://en.wikipedia.org/wiki/Cyc https://en.wikipedia.org/wiki/Cyc
- leobg 5y agoIsn’t that kind of easy to do? I mean, you can do that kind of thing as a one-man-show on consumer hardware using GloVe or fastText.
- rhtgrg 5y agoI think pg is missing something important here. The reason Google was able to beat Yahoo, Altavista, Ask, etc. was not just because they had a better formula — it was also because they started in the era where 'search' was still seen as secondary to 'portals' by the big guys. Had these companies known how important search is to the internet back then, they would've copied Google's secret sauce and crushed it long before it could suck up their traffic. This isn't going to happen again. Google isn't going to sit around twiddling its thumbs while a competitor develops a better algorithm. You have to attack the problem from a different angle entirely (make something that looks nothing like a search engine), I don't think a niche market is going to be enough. Perhaps you just want to make something that scares Google into acquiring you, rather than actually bettering the situation. If that's the case, I implore you to think of doing better ways to spend your life.
- titzer 5y ago> I don't think a niche market is going to be enough. To displace a giant gobbling 180 billion dollars a year? Yeah, no kidding. But nobody is asking for that. They are just asking for decent search results.
- deleted 5y ago[deleted]
- jart 5y agoAltavista was the only search engine in the same league as Google. So Google hired the guy who built it. Altavista infrastructure couldn't scale beyond a single server, because that's how DEC was, so it was a smart move for him.
- imranhou 5y agoI believe google tracks click throughs from search results pages, which should provide in theory plenty of insight into what links aren't really working for specific keywords and what are... thus helping improve or reduce rankings of SEO laden sites. Wonder if someone can throw light on to why this isn't effective.
- deleted 5y ago[deleted]
- streamofdigits 5y agoEventually search will become a decentralized activity (No, not a web3/crypto/coin type decentralization, I am talking about the useful type). Is there any particular reason why internet search has to have a distorting gatekeeper to the global commons (that pretends playing Maxwell's demon). For chrissake, the stuff being indexed is public.
- narmiouh 5y agoNot all that google indexes is public, primarily paywalls/loginwalls allow google IP's to crawl information unhindered but as users you are not, so a new search engine will have to get to a scale and popular enough for others to open up. Quick example: Google can index many news sites, or LinkedIn profiles for example that a regular user with no account cannot.
- streamofdigits 5y agoThats true, but probably something that can be tackled later and in any case it would not be a show-stopper for creating a valuable alternative (There are similar thorny related issues around IP e.g. for news sites)
- mrkramer 5y ago>Eventually search will become a decentralized activity (No, not a web3/crypto/coin type decentralization, I am talking about the useful type). People care about UX not about technology remember that unless people are willing to sacrifice good UX in order to have greater security and privacy. These things are tricky and there is no right formula.
- leoc 5y agohttps://twitter.com/mwseibel/status/1477707884632834049 https://twitter.com/mwseibel/status/1477707884632834049 > I’m pretty sure the engineers responsible for Google Search aren’t happy about the quality of results either. I’m wondering if this isn’t really a tech problem but the influence of some suit responsible for quarterly ad revenue increases. Please no more of this. Two men, Page and Brin, together have basically unfettered control over Google.* If Google does something bad then, unless it's genuinely something small enough that those two could not be expected to hear about it, it's happening with—at the very least—their acquiescence. And low overall search quality is not something that some "suit" is successfully hiding from Good Czar Larry. They could fire the "suit", or command him or her to make other decisions. This is—again, at the very least—something that they have chosen not to do. The responsiblity lies with them. * There is the risk of lawsuits from the minority shareholders, I assume. But IIUC this is not realistically that big a restraint on what shareholders with a majority of votes can do. However IANAL.
- usrusr 5y ago“You might need to do a lot of manual spam fighting initially“ How would this be limited to "initially"? Wouldn't it be a lot, initially, and then only get worse?
- hnbad 5y agoI guess Paul's definition of "beating Google" is "creating a startup without clear revenue path aiming to be acquired by Google or a competitor" as I can't think of any meaningful way a niche search engine would provide a good enough value proposition against existing Google competitors or embeddable search engines (as well as SaaS like Algolia).
- nojito 5y agoThe issue is and will always be monetizing. Anyone competing with Google will need to have a robust monetizing strategy to survive.
- djoldman 5y agoGoogle knows how to surface relevant results and they choose not to because they aren't optimizing for relevant results, they're optimizing for revenue or profit within some constraints (don't lose too many users, privacy, avoid actually terrible or completely irrelevant results). All the various suggestions in this thread plus far more complex and insightful solutions are known to Google. Most of it boils down to using automated user feedback to improve or measure search result relevancy. Google doesn't need to solicit user upvotes / downvotes to improve rankings. They can monitor user clicks on results in addition to analytics on the sites the users visit to determine which sites are relevant to which searches. Google doesn't optimize for search relevancy.
- lubesGordi 5y agoLikewise Youtube doesn't optimize for relevant results, only engagement to maximize ad exposure. The side effect of this is polarizing content gets returned more than relevant content (polarizing content being more engaging than relevant content apparently).
- deleted 5y ago[deleted]
- AtNightWeCode 5y agoI think it will be hard to create a great a search engine while the web works as it does today. Maybe there could be like a sitemap but for text content that has the content structured, indexed, and signed by a trusted party in a way that makes it easy to analyze for plagiarism and so on.
- EamonnMR 5y agoBeing able to flag Fandom, Quora, and Pinterest results would bring me great joy.
- PaulHoule 5y agoI have wondered about this. When I run web sites I frequently look at the log and find a large fraction of the traffic is from search engines. This is a problem because it costs me money to serve that traffic. It might not be initially obvious but it costs more than serving real users because the search engines will scan everything and break the cache. Google sends a significant amount of traffic. Bing sends a detectable amount of traffic. Baidu's crawler might be more active than the two of those together but I never get hits from Baidu. Other crawlers deliver me trouble instead of value: even if I'm not interested in hosting pirate or plagiarized content, a crawler that is looking for trouble is only going to bring me trouble. I hate doing it but I turn off crawlers other than Google and Bing both at the robots.txt and web server level because I just can't afford to serve Baidu queries. I'd like to sign an exclusivity contract with a search engine such that they get exclusive access to crawl it and in turn I get a privileged position in search results. This would give the search engine and myself an incentive to deliver end-to-end quality results.
- beefield 5y agoOkay, given that we have pretty successful examples of wikipedia as a general crowdsourced information storage and stackoverflow as a specialized domain crowdsourced Q&A site, would it be impossible to build a crowdsourced search engine? Not even scraping the web, but I would just type my search term, if that is already searched and results voted, I would see those. If it wasa completely new search term, I would get no immediate results, but my search would be displayed in "new searches page", which some voluntary people would be following and trying to add relevant results.
- blunte 5y agoGoogle search results are garbage, at least from a developer's perspective. Most of the results are poorly formatted content "gathered" from stackoverflow, github, quora, etc. And from a "person who wants to see an image" perspective, Google is purely a gateway to Pinterest or Gettyimages.
- nobbis 5y ago10 years ago, the original engineer of Google's search engine told me what he now wanted was asynchronous, human-powered search with curated results, e.g. a Google-like interface, but queries cost $5 and take 15 minutes. Money's no object for him, so he wanted to outsource the filtering, ranking, and interpreting of results. Would be even more useful today (albeit a tiny TAM.)
- mitchtbaum 5y agoHow will these search engines interoperate?
- lubesGordi 5y agoHow about a platform for curation. Curators who know a subject well can link to content that looks good to them. Search goes through the curators, people can favorite certain curators. Lots of people like to curate. This is a better idea than trying to go after spam.
- li2uR3ce 5y ago301 Battlefield moved: curator spam.
- pictur 5y agoPeople don't want to search anymore. they want to see well-categorized data. For example, instead of searching for cheap vacuum cleaners, I think they want a site that lists vendors that sell cheap vacuums.
- baby-yoda 5y agosearch ads responsible for the rise of google search[1], content ads (seo spam) responsible for google search's fall? my guess is the rate of spam content production far outpaces the rate of original content creation. so the power law concentrates even further in the tiny percentage of OC and a moat forms around them (highest ad $, highest authority/authenticity). where do we end up 5 years from now? further consolidation and the continued return to aol style portals (telco/media giants and fast-lane to own content?) pay-to-access silos dominating the internet? [1] oversimplifying a bit of course, there was a novel ranking method that was more than accurate enough, and it scaled, which allowed for the search ad business to go gangbusters.
- canyonero 5y agoI've been troubled by the just plain awful results being delivered by Google search over the last few years. I think these are just plain hard problems to solve and that Google is not incentivized to solve. Google wants you to click on ads at the end of the day, full-stop. Often times I find myself searching for "best ($product|$thing_to_do)" which I think many other people do as well because we all want the best. Other times I'm looking for a music or a book recommendation with some depth. This of course nearly always leads to SEOd trash. There is no relevance nor is there trust. So, I like others to use keywords like "reddit" or "forum" to get to real humans who I trust and intentions are not to sell via affiliate links. These issues often lead to the need in finding trust in real human-centered recommendations that stem from real human interests and needs. I've never found an algorithmic solution to this problem. This is why I think college radio stations or those south-of-the-dial end up being so, so much better. And why beer recommendations from your local brew-shop owner are better than anything you can find on the net. I think building search vertical that are hand-curated would be very interesting to see. But I also think we need to build more communities which allow recommendations to be shared without an incentive to get hits via search and aren't paid for by large corporations and where community impact/quality _is_ incentivized. I do worry that those days may be gone and there are just not may be enough folks (not in tech) willing to spend so much time online and contributing to niche communities. A lot of folks spend much of their time in walled-gardens like Facebook, Instagram or Twitter, so it'll be challenging to be sure.
- ijidak 5y agoIn some ways paid search disincentives Google from delivering quality organic results. The larger the gap between paid results vs organic results, the more users click the paid results. Not sure how to solve this problem.
- topicseed 5y agoBut paid results do not always, if ever, answer the search query better in any shape of form. So this would end up with displeased users and bounce backs.
- 5y ago
- heisenbit 5y agoAffiliate links are environmental toxic waste and it would be only logical to tax such affiliate payments to fund cleanup and mitigation efforts.
- cblconfederate 5y agoWho would write good or even decent content for free?
- pkamb 5y agoA search engine that only indexed Reddit, Stack Exchange, Wikipedia, and a small number of other "good" sites would get 80% of the way there. No, DDG bang operators don't let you do this. I want an SERP, not a shortcut to a single site's on-site search.
- marban 5y agoFor business news, I do this with https://yup.is https://yup.is
- legohead 5y agoIt would work until you got big enough, then you'd end up following the same path as Google, as that's where the money is.
- pwdisswordfish9 5y ago> Lots of people want to be amateur police. And boy would Google find it hard to follow you down that road. Kinda like they tried with YouTube Heroes? But then, who’s to say you won’t get the same kind of backlash?
- mrkramer 5y agoGood idea Paul I had similar one but no way you would do it manually. Machine learning algorithms need to detect spam not people because that way search engine can't scale. If people were marking what's good content and what's not such search engine would be reduced to content curation not organic search and discovery.
- jefftk 5y agoI'm seeing a lot of comments along the lines of, "Google shows ads on the SEO-gamed sites that show up in results, so their incentive is to give spammy results". But wouldn't this predict that results would be much better on Bing and other search engines that don't have much presence in the "put ads on random sites" market? (Disclosure: I work on ads at Google, speaking only for myself)
- going_ham 5y ago> But wouldn't this predict that results would be much better on Bing and other search engines that don't have much presence in the "put ads on random sites" market? Honestly what I think is every search engine sucks these days, and Google manages to suck a little bit less. The reason is because how easy it has been to publish low quality content. It's rare to find high quality contents. The issue with search engine is that they don't show these rare contents. These aren't recommended by default. These are hidden. What happened is the recommendation system is broken! If there weren't any neural networks making decision, it would have been different issue. But with modern search engines deploying recommendation system, I think it is all about rich gets richer scheme. You can't recommended new or fresh but quality content because it was never visited! So, when the entire backend is relying on user data, the system is being fed crappy data because users don't care and those who do are few in number. As long as system is making revenue, it will be this way. Most people never care at all and would never be bothered because they only care for simple queries. If anyone deviates from the norm, Google search results are pretty bad like every other search engines.
- deleted 5y ago[deleted]
- Marazan 5y agoGoogle's ranking alhorithm shaoes the web. And the web now looks like a 1500-2000 word listicle with 3 images becasue that is what thr ranking algorithm favours. If you find the info you need and leave quickly that actually down ranks the page. That is is idiotic. Pages that give you what you want quickly are punished!
- Quenhus 5y agoFor developers, you can remove some spam websites from Google and other search engines, with these uBlock filters: https://github.com/quenhus/uBlock-Origin-dev-filter https://github.com/quenhus/uBlock-Origin-dev-filter
- anovikov 5y agoSadly the only way of fixing it is making search results unattractive for cracking. Ranking of a page in search results is a metric, and every metric is a hackable metric. Only way they won't be hacked is if there's no incentive to. Sure a search engine that specialises on narrow area of knowledge without much money in it, can be very relevant and bullshit-free. But there's no way to make it work for the general web search. People hack things. If they didn't we'd have Communism built by now (yes the "good" - classless, stateless one).
- deleted 5y ago[deleted]
- anfilt 5y agoAn other thing does not help is how some sites gate content from being scraped. Also forums are not as popular today again reducing the amount of indexable content. Think about some sites have migrated from using forums to something like discord.
- mirekrusin 5y agoThe whole thing seems "simple" to me – graph of identities with url vetting/liking/approve-this-message-like actions, you don't need anything else. Reputation, non-fakeness etc. can be derived from it for anybody - you just list identities you trust/follow (with weights?) and anything you look at can be scored. Virtual identities can also be created, ie. identity listing all links mentioned on HN (with positive sentiment only?), links from wikipedia etc. so people can follow those to create their reality graphs. The interesting part is that it doesn't claim universal truthness - depending on who you follow your results will be skewed towards their opinion of the world. Ie. if you follow MIT, Wikipedia and E. Musk you'll see different view of truthness than somebody following FOX News and Flat Earth Society for example. It could be interesting to focus on "dislike" marking (only?) as it may be much more lightweight to approach it from blacklisting side.
- mrlanderson69 5y agoIf anyone has the skills to work on something like this please email me (email address in my profile.) I can show you a demo. Just to show I am not screwing around: if you don't like the demo I will pay you $500.
- anyfactor 5y agoI have seen several google alternative search engine projects being posted in HN every other week. You have your privacy focused open source google alternative search engine for "insert niche here" with big hopes of disruption. I will give you my two cents. I have used duckduckgo, bing, searx etc. for extended periods of time and hated every one of those things. The problem is that what you search seem to be essentially the gateway to wild west of internet. I understand the proposition of spam control in search engines, but atleast to me I think the early days of google without DMCA and copyright bans made google the best. I fear SEO spam control will only bring the worst of the moderated internet. It will not be the first time big tech tried to douse a gasoline fire with more gasoline because they taught the more fire meant the previous fire will get suffocated by the lack of oxygen(?). Rather than using "AI" as a crutch to solve SEO as a problem, I want to see an option that is true to 2005 era google.
- lifeisstillgood 5y agoGoogle is not important because it has all the information - it's important because it has hardly any. A major complaint is that there used to be good free reviews of commercial products that could be easily found. That is not "all the information". Information about the current round of commerically advertised products is something like 5-10% of all commerce (or less). And we are entering a world where "all the information" is what we do all day, what we say, how we react to different stimuli. That is the real review sites - why do people take this train and not that, why is that park safe and this one full of muggings. We need to solve the Google problem not because we want blogging like it's 2009 but because epidemiology is about to open humanity's eyes. And it's going to hurt if we don't make it free and open.
- deleted 5y ago[deleted]
- waynesonfire 5y agowhat a great idea.
- wslh 5y agoIn 2013 I elaborated about this topic: http://blog.databigbang.com/letters-from-the-future-challenging-googles-search-engine/ http://blog.databigbang.com/letters-from-the-future-challeng... I would add that in 2021 we can easily do Natural Language Understading (NLU) and Natural Language Generation (NLG) and can build zillions of web pages that don't follow the original page ranking concept of Google. Probably important sites share less low rank pages and there are many more link rings and clusters. More decentralized blogs seems a thing of the past (expecting to be rebooted in the future).
- aaron695 5y ago
- donio 5y agoWhat I am looking for is control over the results. Personalized blacklists and lists of sites to be (de)prioritized and also the ability to subscribe to community curated versions of the same. And to be clear I want to be able to control these myself, not algorithm trying to guess my preferences. No guessing, just do what I tell you to. Multiple search profiles with different priorities would be nice too. I would like the search algorithm to be transparent, I should be able to tell why I got a certain result and how I can avoid such results in the future.
- oefnak 5y agoYou can use ublock origin to blacklist domains from your search results. I do this for example with codegrepper and other sites who just copy paste Stack Overflow comments in a less readable format.
- donio 5y agoThanks. I've been using other browser extensions for this, doing it from ublock origin is a good idea, will simplify things.
- 1vuio0pswjnm7 5y agoRanking search results on popularity is flawed. It may improve search engine performance and the effectiveness of online advertising but it penalises users who can think critically and independently. There is an underserved market that has been left behind by Google and its "competitors". PageRank seemed to borrow from the concept of citation count. The idea that "importance" could be measured by the number of times a webpage, like a paper published in a peer-reviewed academic journal, was referenced by other webpages, like other papers published in peer-reviewed academic journals. The initial name for the project before "Google" was "Backrub", referring to the reliance on "backlinks" to quantify importance. An index of a commercially-oriented www full of sites supported by online advertising is nothing like Web of Science or some other database collection that allows ranking by citation count. The www has no peer-review and no limits on commercial activity. Google succeeded in creating something highly profitable and sometimes useful, but the founders never delivered on their original promise. That was a search engine in the academic realm, where the technical details were public, and one that would be free from the influence of advertising.^1 Instead the project was turned into an online advertising business. A 180-degree pivot. The moral/ethical debate went from the question of being advertising-supported to the question of invading the personal privacy of users, for the benefit of advertising. Whatever ideas the founders held in 1998 regarding the influence of advertising on web search were overtaken by the lure of pure financial success. Once oppposed to idea of using cookies for advertising purposes, the founders were persuaded to purchase DoubleClick, a company with a terrible privacy record that uses cookies and purchasing data to profile users as ad targets,^2 for almost double what they paid for YouTube. Not sure what if any moral/ethical debate remains today. While the company is being sued simultaneously by hundreds of plaintiffs, including the US government, one of the founders is "hiding out" on a small island in the South Pacific. Whatever motivations he had to make an open, academic search engine free from the influence of advertising, they seem to be gone. In sum, the world still needs a decent web search engine free from the influence of online advertising. 1. https://infolab.stanford.edu/~backrub/google.html https://infolab.stanford.edu/~backrub/google.html Excerpts: "Up until now most search engine development has gone on at companies with little publication of technical details. This causes search engine technology to remain largely a black art and to be advertising oriented (see Appendix A). With Google, we have a strong goal to push more development and understanding into the academic realm. Appendix A: Advertising and Mixed Motives Currently, the predominant business model for commercial search engines is advertising. The goals of the advertising business model do not always correspond to providing quality search to users. For this type of reason and historical experience with other media [Bagdikian 83], we expect that advertising funded search engines will be inherently biased towards the advertisers and away from the needs of the consumers. Furthermore, advertising income often provides an incentive to provide poor quality search results. [T]here will always be money from advertisers who want a customer to switch products, or have something that is genuinely new. But we believe the issue of advertising causes enough mixed incentives that it is crucial to have a competitive search engine that is transparent and in the academic realm." 2. https://www.nytimes.com/2000/02/17/technology/us-investigating-doubleclick-over-privacy-concerns.html https://www.nytimes.com/2000/02/17/technology/us-investigati... https://slate.com/technology/2005/11/why-web-surfers-love-to-hate-cookies.html https://slate.com/technology/2005/11/why-web-surfers-love-to...
- rel2thr 5y ago> You might need to do a lot of manual spam fighting initially This is why I am very hyped on Brave search's goggles feature , it will let you share exclude / include site lists to use w/ the search engine. Hopefully it will empower these niche communities to curate a list of non-spam sites ( like the ad-blockers do with ads today )
- lucasyvas 5y agoI don't think this problem can be solved by another search engine as they currently exist. The problem must be solved by a new kind of search engine that exclusively searches Internet communities (HN, Reddit, etc.). The content must be community produced, since all other forms of writing have monetary incentive. You will find the best results in communities of enthusiasts with respected moderation teams. So, a curated strategy where the users can UP/DOWN sources they trust for particular topics. Relevancy and user ranking of answers determines score. Of course, the user can search any sources they like, but a voting system would control the defaults. I think this avoids slanted results as well, because topics are objective. The subjectivity will be in the comments, where they belong. That's in contrast to today where SEO scams can determine how high up results are. So, to game my proposed search engine, you need to infiltrate the users. I believe this is harder to do across numerous sources compared to the current system, which is game the algorithm.
- quickthrower2 5y agoGoogle is so spammy I now instinctively use other search methods at times. Which is very interesting, because doing so is high friction. But it's so spammy out there that pain(spam) > pain(friction). It is not to be "un-Google" but because I get better results. For example searching in a good subreddit can be more fruitful, giving answers from genuine people in moderated parts of the internet. If you get crap then try another subreddit - some mods are better than others. Is this a business opportunity - I think so, although I have no idea how you would go about it. Maybe a decent search engine for programmers would be a good start! E.g. "Exception Message XYZ" + site with decent answers.
- 1vuio0pswjnm7 5y agoRanking search results on popularity is flawed. It may improve search engine performance and the effectiveness of online advertising but it penalises users who can think critically and independently. There is an underserved market that has been left behind by Google and its "competitors". PageRank seemed to borrow from the concept of citation count. The idea that "importance" could be measured by the number of times a webpage, like a paper published in a peer-reviewed academic journal, was referenced by other webpages, like other papers published in peer-reviewed academic journals. The initial name for the project before "Google" was "Backrub", referring to the reliance on "backlinks" to quantify importance. An index of a commercially-oriented www full of sites supported by online advertising is nothing like Web of Science or some other database collection that allows ranking by citation count. The www has no peer-review and no limits on commercial activity. Google succeeded in creating something highly profitable and sometimes useful, but the founders never delivered on their original promise. That was a search engine in the academic realm, where the technical details were public, and one that would be free from the influence of advertising.^1 Instead the project was turned into an online advertising business. A 180-degree pivot. The moral/ethical debate went from the question of being advertising-supported to the question of invading the personal privacy of users, for the benefit of advertising. Whatever ideals the founders held in 1998 were overtaken by the lure of pure financial success. Once oppposed to idea of using cookies for advertising purposes, the founders were persuaded to purchase DoubleClick, ground zero for the explosion of online ads, for $3.1 bilion. Not sure what if any moral/ethical debate remains today. While the company is being sued simultaneously by hundreds of plaintiffs, including the US government, one of the founders is "hiding out" on a small island in the South Pacific. Whatever motivations he had to make an open, academic search engine free from the influence of advertising, they seem to be gone. In sum, the world still needs a decent web search engine free from the influence of online advertising. 1. https://infolab.stanford.edu/~backrub/google.html https://infolab.stanford.edu/~backrub/google.html Excerpts: "Up until now most search engine development has gone on at companies with little publication of technical details. This causes search engine technology to remain largely a black art and to be advertising oriented (see Appendix A). With Google, we have a strong goal to push more development and understanding into the academic realm. Appendix A: Advertising and Mixed Motives Currently, the predominant business model for commercial search engines is advertising. The goals of the advertising business model do not always correspond to providing quality search to users. For this type of reason and historical experience with other media [Bagdikian 83], we expect that advertising funded search engines will be inherently biased towards the advertisers and away from the needs of the consumers. Furthermore, advertising income often provides an incentive to provide poor quality search results. [T]here will always be money from advertisers who want a customer to switch products, or have something that is genuinely new. But we believe the issue of advertising causes enough mixed incentives that it is crucial to have a competitive search engine that is transparent and in the academic realm."
- jliptzin 5y agoImprove your search engine results with this one weird trick! Just block any domain containing the word pinterest
- ChuckMcM 5y agoAs I pointed out to paulg yesterday this was exactly the business model / concept that Blekko was created to address. The idea being that one could use "slashtags" to curate web sites that were "good" on a topic (not spammy) and pull results from that rather than the general web. Guess what? It works great! Also, it doesn't make enough money to support the company using advertising. For a couple of years, Blekko ran a "3 card monte" game where we white listed the results from Google, Bing, and our own index. For every "contested" query, Blekko consistently beat the others by a significant margin. If the query wasn't contested, Bing and Google did about the same, and if the query was obscure, typically Google did better than Bing or Blekko. What is a "contested" query? That is one where there is a lot of money on the line. My favorite one was "best credit card" (which is search engine shorthand for "What is the best credit card?" because the stop words "What", "is", and "the" are removed). Why is it contested? Because if you put an advertisement into the results of that query, and the person making it clicked on that link and signed up for a credit card, you could be paid $50 or more. For a single click. Other queries that advertisers would pay well for getting the traffic of the user were, car dealerships, hotel chains, jewelry retailers, and university "referral" services (like the one that was busted for getting people into Ivy League schools by faking academic records). Extremely few people click on an ad put onto a page of search results for the query "what is shoe rubber made of?"[1]. However it is required to serve queries like that so that people will come back when they are looking to spend money on something. So using the same exact idea that Paul proposed Blekko built an English language index which allowed you to curate the crap out of your search results and return much better data. The "value" of that was not considered to be high enough to insist on people logging in to use the engine. Knowing an id for the person making the query allowed for user specific blacklists of spammers (so if for example you never wanted to see a Pinterest link in your results you could make that happen). Without sufficient traffic, using the feedback loop "of these documents, which one was clicked as the 'best' answer?" type algorithms for ranking fail to converge rapidly enough for decent ranking. Without a credible threat that if your site is not included in the index, your traffic will be greatly reduced, it is difficult to negotiate with web sites to permit crawling, rather than deny your crawls with the robots.txt file. Blekko's best customers and most ardent fans? Reference Librarians. Yup, people who needed web search to do their jobs, not to find the movie times for the latest feature. Blekko never did try to create a subscription service, but I think such a service that is somewhere between free and the $$$ of LexisNexis has a shot, at least as a lifestyle business. You still need to get rights to the data and that gets harder and harder. [1] Okay, bots do, but humans don't
- calltrak 5y agoHere is a handy list of alternatives to google search https://fabform.io/a/alternative-search-engines https://fabform.io/a/alternative-search-engines
- anderspitman 5y agoI for one am optimistic what a "post-search" world might look like. Maybe a lot like the early web. I don't think affiliate links themselves are necessarily the problem. I'll gladly use a link from a high-quality reviewer to give them a little kickback. SEO seems to be the issue. Maybe we end up with trusted brands for reviewing specific things. For example, I trust outdoorgearlab.com for pretty much anything camping related, and no purchase comes to mind that I've regretted yet.
- MarkMc 5y agoI just switched to Duck Duck Go yesterday and was not impressed. When I searched "define hot take" Google gives me a canonical, prominent definition with bold-font title and button to hear the pronunciation. DDG gives me multiple definitions in a normal search results page with none prominent, and no way to hear the pronunciation. I'll be switching back to Google
- epolanski 5y agoKey difference you seem to miss: Google doesn't want you to leave the search engine (unless it's ads), so whether you look for a translation, sport results, weather or definitions they hoard content and shove it on their page. I agree it's often convenient, but realize Google does more than provide links.
- Fede_V 5y agoI think the complaints about SEO spam are valid - but - I think msweibel and pg misdiagnose the challenge. The challenge is that you are dealing with an adversarial system, and, the better your search engine is, the more widely used it becomes, the more valuable it is for your adversaries to find ways to game your rankings. Any new niche search engine will go through a small window of time where they have the luxury that none of the sites they are indexing are spending all their effort trying to reverse engineer your signals, and optimize against them. I'm incredibly skeptical that they can remain useful once people all the SEO efforts of various marketers start to be turned against them.
- ajmurmann 5y agoThe main motivation to SEO crappy content seems to be ads and affiliates links. What if you take the motivation away to SEO crappy content by deprioritizing sides that contain ads or affiliate links? Of course Google would never do this, but someone else could.
- inetknght 5y ago> The main motivation to SEO crappy content seems to be ads and affiliates links. Maybe the main observed motivation. But I'd argue that a lot of that is just a fraudulently-profitable front to much more devious problems.
- Drew_ 5y agoThis doesn't work because all of the "good" content has ads and affiliate links too.
- stickyricky 5y agoWhat are the pros and cons of a user generated tagging system? If you have a community of dedicated individuals who maintain a group of tags, searching those tags should yield high quality results.
- konaraddi 5y agoA problem is that good SEO doesn’t meant good quality. And assessing quality is hard, so people lean on other people to assess quality (either by appending something like “Reddit” to their search queries or asking friends irl or on twitter/discord). I wrote a bit more about search engines competition and problems/opportunities here - How Alternative Search Engines Can Win Users https://konaraddi.com/writing/2021/2021-08-05-on-search-competition/ https://konaraddi.com/writing/2021/2021-08-05-on-search-comp...
- hammyhavoc 5y agoWhy not Searx or YaCy?
- mrlanderson69 5y agoWe are working on exactly this problem. IF anyone wants to see a demo please email me.
- mrlanderson69 5y ago
- mrlanderson69 5y ago
- zavkz 5y agoOh so it's not just me... Most recently I was trying to find a way to reset a printer and also fix a certain error code. I search on google and it's filled with irrelevant content, unrelated to the model number I just put in, scam websites, ink sellers, even though I used correct filters such as + sign and the syntax, which is a joke if you think about it. A billion dollar company and this is the best they do. The advanced search is burried, the syntax they have is explained by third parties or burried somewhere in their options. I jumped onto youtube, I put in the model number, same thing pretty much. I get unrelated videos mixed in with the model number I put in. I'm pretty sure some videos even though the model number is the same is not being shown. Ironically there's a video explaining a certain solution and warning people not to fall victim of another video scamming people, which the dislike has been removed, comments obviously deleted so some people may be calling and getting scammed.
- mtnGoat 5y agoBig money in SEO, I had an acquaintance all the way back in the early 2000s that had tens of thousands of domains that he ran experiments on to reverse engineer how the algorithm worked. He also had tens of thousands he ran SEO/link networks on. He made a lot of money for a long time by being front page for a lot of terms. Same thing is happening today, there are just more of these actors doing it. They just game the algorithm for terms related to products. Notice you still get decent serps in Google for terms that don’t relate to something that can be sold using an affiliate link. Fairly easy problem to fix, but Google would have to hire a black hat to help solve it. But the good ones ain’t gonna work there.