3 ms·
This is fascinating in a potentially terrible way. In summary: when an image is posted, a reverse-image search is done and top results are scraped then added to
by alpha_squared 6y ago
This is fascinating in a potentially terrible way. In summary: when an image is posted, a reverse-image search is done and top results are scraped then added to a "More like this" section for the post. This ranks highly because it's exactly what Google already associates with that image, except now all in one page instead of across multiple pages.
Applying this same method to other content (blog posts/pages and video posts) would, presumably, also work. The terrible part of all this, is that it would create more junk posts ranking higher and stratifying the word association of that content with the respective search terms. New content could end up having zero chance of ever ranking highly because another factor of result ranking is content age (older content is weighted heavier).
Am I understanding all that correctly?
- nullc 6y ago> Applying this same method to other content (blog posts/pages and video posts) would, presumably, also work. I believe reddit started recently doing something like this for text and as a result has made google searching for reddit posts essentially useless. Basically when you view any post/thread while logged out there are a bunch of other threads shown on the page that have related text, with all the text of their posts included collapsed in the HTML. The result is that when you search for the context of any reddit post you get hundreds of results which on vaguely similar topics which don't contain the post that you're looking for (unless you log out and count hidden text collapsed under other threads).
- specialist 6y ago"...all the text of [the related] posts included collapsed in the HTML" Huh. Facepalm. Is there any way to exclude portions of content from indexing? Like maybe inlining robots directed pragmas: <span data-robots="noindex"> blah blah blah </span> Or using meta tags to spec exclusions: <meta name="robots" content="noindex:.related-posts" /> FWIW: https://en.wikipedia.org/wiki/Robots_exclusion_standard#Meta_tags_and_headers https://en.wikipedia.org/wiki/Robots_exclusion_standard#Meta...
- nullc 6y agoI assumed that the practice was intentional to make reddit be more heavily represented in google results. Reddit staff doesn't actually use the site that much (or at least that was my impression a couple years ago after having a meeting in the office and finding that I knew a lot more about the meme-art that users sent them than their staff did)-- so it wouldn't be shocking that they'd be indifferent to making google search unusable for users.