9 ms·
I know it’s easy to throw stones from the outside, but Google’s results are so compromised it seems like it’s a good time to get back in. As one just example,
by boardwaalk 4y ago
I know it’s easy to throw stones from the outside, but Google’s results are so compromised it seems like it’s a good time to get back in.
As one just example, I searched for a unique error message in code that exists on GitHub, is in a fairly popular repo, and is not new and Google just could not find it. That seems like a very basic failure.
- saddd 4y agoI have the feeling that whatever you're talking about is explicitly not crawlable.
- celdon25 4y agoThat doesn’t change anything regarding the actual point of the comment.
- astrange 4y agoYour idea for a search competitor is to ignore robots.txt?
- VWWHFSfQ 4y agoor an advertising competitor that ignores DNT! oh wait
- berkle4455 4y agoIt's a just a text file.
- chihuahua 4y agoThis seems perfectly logical to some people: Google ignores robots.txt: "Google is evil! They're trespassing on my webserver!" Google follows robots.txt: "Google search results suck! They're not indexing GitHub!"
- celdon25 4y agoYes, clearly that's the best possible interpretation of what I said. /s
- saurik 4y agoA public git repository is definitely crawlable. Google seems to have given up actively going out of their way to index things that are hard to crawl as they got so big and important it was easier to just tell people "thou must do X or we won't index you and you want to be indexed", but increasingly the content I want to find is in weird little silos.
- simonw 4y agoYeah, the GitHub robots.txt is surprisingly restrictive: https://github.com/robots.txt https://github.com/robots.txt User-agent: * Disallow: /*/pulse Disallow: /*/tree/ That "/*/tree" rule means that search engine crawlers are allowed to hit the README file of a repo but effectively NONE of the other files in it. Which means that if you keep your project documentation on GitHub in a docs/ folder it won't be indexed! You need to publish it to a separate site via GitHub Pages, or use https://readthedocs.org/ https://readthedocs.org/ (Side note: I just noticed https://github.com/ekansa/Open-Context-Data https://github.com/ekansa/Open-Context-Data is explicitly listed in the robots.txt for GitHub - the only repo that gets a mention like that. I'd love to know the story behind that!)
- burkaman 4y agoThat repo apparently used to be the largest on GitHub: https://news.ycombinator.com/item?id=5912922 https://news.ycombinator.com/item?id=5912922. I bet Google was repeatedly scraping the entire thing and putting too much strain on their servers at the time it was added. It's been 10 years, what are the odds nobody at GitHub today remembers why it was added? Also, very relatable to see a decade old "I'll update this shortly" comment that was never updated. We all have a few of those.
- snowycat 4y agoIt appears that the creator of the repo actually confirmed this: https://twitter.com/ekansa/status/1137052076062650368 https://twitter.com/ekansa/status/1137052076062650368
- knute 4y ago/*/tree is only for directory listings. File contents will be under a /blob/ path, e.g. https://github.com/facebook/react/blob/main/AUTHORS https://github.com/facebook/react/blob/main/AUTHORS, and should be, AFAIK, indexable. (mandatory disclaimer: I'm a GitHub employee, not speaking on behalf of the company)
- staplung 4y ago
- sebosp 4y agoCurious, if I had the list of repos, is there anything that forbids me from `while read url; do git clone $url data;./train data; rm -rf ./data; done`. Besides licensing, ie ratelimit/throttle, similar question, the search for code across all repos provided by github ui gets throttled pretty fast, what do people do? (not suggestion in a hundred(?) years to do the while loop for this tho ;))
- gonzo41 4y agoDo you think google cares that it's loosing it's edge? How do they not know it's getting worse.
- userbinator 4y agoYes, I remember several years ago --- more like 8 now(!) --- easily finding results in GitHub repos whenever I've needed to look up error codes and such. Now even site:github.com doesn't (and if you try too hard, you get the hellban for a while). Another extremely noticeable degradation is in finding part numbers, IC markings, service manuals (NOT the useless user manual), schematics, and the like. Anything that proponents of right-to-repair would be extremely interested in, to the extent that I wonder if there's been some sort of conscious effort being made by certain interests to eliminate or limit such information. Then there's the niche-but-legal adult content. I won't go into too much detail about that, but suffice to say it used to be far easier to find. It's been 5 years since this notorious item here, and I've only seen Google get worse: https://news.ycombinator.com/item?id=16153840 https://news.ycombinator.com/item?id=16153840
- bloodyplonker22 4y agoNot only have they become compromised from a technical standpoint, for some searches in particular, the results have been modified to be heavily politically biased and woke.
- selimthegrim 4y agoDon’t downvote them until you check out Frank Zappa’s discography.
- nus07 4y agoSundar Pichai has so Mckinsified and MBAfied Google that at this point Google search seems like an A/B test to deliver the best targeted ad . Probably better of using any other search including Yahoo .
- goldforever 4y ago[dead]
- Denzel 4y agoWould you mind providing details like the search query and link to the page you expect to be found? To test your hypothesis, I did a basic search for exact matches on "we do not synchronize on the update of the broker node" and Google returned 2 search results in 240ms: - https://github.com/a0x8o/kafka/blob/master/core/src/main/scala/kafka/coordinator/transaction/TransactionMarkerChannelManager.scala https://github.com/a0x8o/kafka/blob/master/core/src/main/sca... - https://jar-download.com/artifacts/org.apache.kafka/kafka_2.12/2.3.1/source-code/kafka/coordinator/transaction/TransactionMarkerChannelManager.scala https://jar-download.com/artifacts/org.apache.kafka/kafka_2.... Which contain exactly the source code from GitHub that I was looking for. You'll notice that the first result is actually a0x80's fork of apache/kafka. Google states that some entries very similar to the 2 already displayed were omitted, and I'm able to remove that filter. With that filter removed, I can see the same document indexed from apache/kafka on GitHub. There's nothing I can do or promise directly, but I can assure you that Google takes the quality of our search results very seriously. If you believe we're not delivering quality results, I strongly encourage you to click that "Send Feedback" link at the bottom of your results so that our teams can act upon your feedback. Disclosure: I work on Search at Google. Disclaimer: The words, views, and opinions expressed in this post are my own. They are not representative nor do they represent my employer in any capacity.
- sefrost 4y agoHow often do people use the send feedback button? How many of the reports are looked at?
- Denzel 4y agoI’m not authorized to disclose data that’s not public knowledge. What I can say is that we have a feedback process in place for Google Search that we use to improve our product. You can send feedback and check the box to allow our teams to contact you if you’re interested in a follow up. Of course, given our scale, we’re not able to follow up on every bit of feedback but that doesn’t mean we don’t review or act upon that feedback in some way.
- 4y ago
- bane 4y agoI was just searching for an old friend of mine who's last name happens to be a substring of another common last name. I tried everything, quotes, + signs, - signs, middle initials, middle names, cities we lived in together, etc. Every single returned link after the first 3 had the superstring version of the name and not the correct name. It turns out that this returns endless results for a fairly well known singer, not my friend. So now did I not get the results I was looking for, I got tons of results that were objectively wrong. Then suddenly, about 6 pages into those results, I started getting ones for the correct last name, but now the first name is a mess. This happened on Google, DDG, Baidu, Sogou, Haosou, Dogpile, the current Yahoo search, Bing, and to some extent on Yandex. Naver was worse, Daum totally worthless with incorrect results. Utterly worthless. The thing is, my friend's name is surprisingly fairly unique, there's probably less than 20 people in the world with that specific name. It's like the search engine's desire to fill the screen with worthless garbage results has overpowered the need to supply the 2 or 3 that are actually correct, even if the quantity is a little disappointing.
- behringer 4y agoTry neeva?
- bane 4y agoNope, same shitty results. I've basically decided that my friend's name is going to be my search engine quality test from now on, the results are so spectacularly terrible. All I want is something the crawl the web, suppress SEO spam, and let me "search" on things exactly as I've quoted them. Like we used to in 1998.
- zbrozek 4y agoI've found that with Google I have to use verbatim mode to get it to do anything vaguely sensible. Synonymization is the worst thing to ever happen to search, and it keeps getting more and more aggressive.
- deleted 4y ago[deleted]
- MuffinFlavored 4y ago> but Google’s results are so compromised I read this constantly here (echo chamber) and I can't help but feel it's a little biased/overdramatic.
- dimator 4y agoHonestly, I'd be all for using ddg exclusively. but I find myself doing !g (their google redirect operator) when I don't find what I want on DDG, and it's almost always the top result on Google. And this happens daily.
- karmakaze 4y agoSeems like the worst time, unless they're doing so with ChatGPT and the like. What regular search lacks is context and a natural way of refining queries by adding context that doesn't always work well with keywords.
- pjmlp 4y agoRight now searching on Google is way worse than Yahoo, Altavista or Ask Jeeves. Now only we get tons of ads back as first results, Google keeps rewriting the queries for whatever "helpful" nonsense.
- asddubs 4y agoand it just straight up ignores keywords even when there's matches containing all of them. google has become so much worse, and yes part of it is that there's a ton of spam, which is also a problem, but it has also gotten worse in other respects too
- fIREpOK 4y ago> As one just example, I searched for a unique error message in code that exists on GitHub, is in a fairly popular repo, and is not new and Google just could not find it. That seems like a very basic failure. I have recently almost completely stopped using Google's search engine due to the fact that I am very often offered zero search results for simple queries (usually involving quotes though) .. It's so bad I can't even believe it. Note: I've been a Google search since it started... Gmail since Beta, etc... At one point, I thought that maybe they started punishing ad-block users excessively.
- papito 4y agoSeems like the giants that were nearly synonymous with "Internet" - Google and Amazon, are rapidly deteriorating and creating a massive market opportunity. Pure speculation, but innovative companies at first, they started over-hiring and bloating, using questionable interviewing techniques (puzzles, Leetcode), taking on thousands of employees who were just there to game the system, coast, and collect the check. It just looks like they stopped caring.
- MagicMoonlight 4y agoMost tech companies have ruined their products now. They’ll have 10,000 engineers and 15 iterations of the UI but you try and buy a hard drive and it’s a box with an SD card taped inside. It’s time for competitors to start wiping them out.
- pharmakom 4y agoWhy do they even bother with the SD card?
- trilbyglens 4y agoMy guess is that there is now so much ml Blackbox shit going on in the search algo that no one can reasonably tell you why it returns what it does.