4 ms·
One reason it might deteriorated is that goog is constantly battling ppl 'optimizing' their content for Google while competitors likely see less than 1/1000 of
by RandyRanderson 5y ago
One reason it might deteriorated is that goog is constantly battling ppl 'optimizing' their content for Google while competitors likely see less than 1/1000 of this.
- debesyla 5y agoThat's a good point! A lot of online content is created for search clicks only, not quality. For example, I often search Google in lithuanian language and I keep getting whole lot of auto-translated blog/article farms as a result. Those translations almost always are shoddy, but these pages keep popping up in the first Google results because a) there is not a lot of native (language) content to push out these clickbaits and b) those clickbaits are on domains that are also, well, in the first results of other lesser languages... I think these auto-translated pages started to pop up two or three years ago...
- colordrops 5y agoThat would make sense if so many easy to detect low hanging fruit like blog spam, scraped stack overflow pirates, and listicles didn't make it to the top. Those are easy for google to de-rank and yet they don't.
- IndexPointer 5y agoThere are billions of dollars on stake from both sides, search engines and spammers, in an endless arms race that has been going on for more than 20 years. Trust me, it's beyond naive to say fighting webspam is a low hanging fruit problem.
- colordrops 5y agoWhy should I trust you? I trust my own eyes. I regularly see spam sites that get to the top of results that are seen by many people for months. These could be filtered with a one line change.
- caconym_ 5y agoWhat's the source for it being easy? It seems easy, from the perspective of a human looking at results, but I'm not sure how much that's worth given the scale and complexity of the problem.
- colordrops 5y agoCaptchas are way harder to solve than it is to detect these sorts of poor results. Google should have absolutely no problem building a classifier that could scale to solve the problem. But as you said, it's not worth it to them, but for the reasons of losing revenue rather than scale and complexity.
- Kavelach 5y agoJust block the domain. At first, you can block manually, but we know Google doesn't like doing things that way. Fortunately, they have a lot of heuristics to find sites like that; usually the content is just copied from another source. And since they scrape the web all the time, they should know which content has appeared first where. But the issue isn't that they can't; the issue is that they don't want that. Why the sites with copied content exists? To earn money through ads. What earns Google money? Ads!
- IndexPointer 5y agoAny simple heuristic has false positives, meaning they'll end up taking down legitimate sites that had repeated content for a good reason. Say, for example two sites quoting text from the us constitution. The second one to be crawled would be considered to be spam copying the first one and removed from web results. Then you'll get comments on hacker news complaining that Google is censoring it for political reasons. And any simple heuristic is quickly reverse engineered by SEOs, who will find a way to mask it as legitimate. tl;dr it's a hard problem.
- Kavelach 5y agoThey could use the heuristics to build a list of domains to block and then have someone review it. After doing it for a long time, they could build a neural model on top of that, and automate it. As I have said, the reason they don't do it is not because they don't have the skills and know-how.