4 ms·
I have the feeling that whatever you're talking about is explicitly not crawlable.
by saddd 4y ago
I have the feeling that whatever you're talking about is explicitly not crawlable.
- celdon25 4y agoThat doesn’t change anything regarding the actual point of the comment.
- astrange 4y agoYour idea for a search competitor is to ignore robots.txt?
- VWWHFSfQ 4y agoor an advertising competitor that ignores DNT! oh wait
- berkle4455 4y agoIt's a just a text file.
- chihuahua 4y agoThis seems perfectly logical to some people: Google ignores robots.txt: "Google is evil! They're trespassing on my webserver!" Google follows robots.txt: "Google search results suck! They're not indexing GitHub!"
- celdon25 4y agoYes, clearly that's the best possible interpretation of what I said. /s
- saurik 4y agoA public git repository is definitely crawlable. Google seems to have given up actively going out of their way to index things that are hard to crawl as they got so big and important it was easier to just tell people "thou must do X or we won't index you and you want to be indexed", but increasingly the content I want to find is in weird little silos.
- simonw 4y agoYeah, the GitHub robots.txt is surprisingly restrictive: https://github.com/robots.txt https://github.com/robots.txt User-agent: * Disallow: /*/pulse Disallow: /*/tree/ That "/*/tree" rule means that search engine crawlers are allowed to hit the README file of a repo but effectively NONE of the other files in it. Which means that if you keep your project documentation on GitHub in a docs/ folder it won't be indexed! You need to publish it to a separate site via GitHub Pages, or use https://readthedocs.org/ https://readthedocs.org/ (Side note: I just noticed https://github.com/ekansa/Open-Context-Data https://github.com/ekansa/Open-Context-Data is explicitly listed in the robots.txt for GitHub - the only repo that gets a mention like that. I'd love to know the story behind that!)
- burkaman 4y agoThat repo apparently used to be the largest on GitHub: https://news.ycombinator.com/item?id=5912922 https://news.ycombinator.com/item?id=5912922. I bet Google was repeatedly scraping the entire thing and putting too much strain on their servers at the time it was added. It's been 10 years, what are the odds nobody at GitHub today remembers why it was added? Also, very relatable to see a decade old "I'll update this shortly" comment that was never updated. We all have a few of those.
- snowycat 4y agoIt appears that the creator of the repo actually confirmed this: https://twitter.com/ekansa/status/1137052076062650368 https://twitter.com/ekansa/status/1137052076062650368
- knute 4y ago/*/tree is only for directory listings. File contents will be under a /blob/ path, e.g. https://github.com/facebook/react/blob/main/AUTHORS https://github.com/facebook/react/blob/main/AUTHORS, and should be, AFAIK, indexable. (mandatory disclaimer: I'm a GitHub employee, not speaking on behalf of the company)
- staplung 4y ago
- sebosp 4y agoCurious, if I had the list of repos, is there anything that forbids me from `while read url; do git clone $url data;./train data; rm -rf ./data; done`. Besides licensing, ie ratelimit/throttle, similar question, the search for code across all repos provided by github ui gets throttled pretty fast, what do people do? (not suggestion in a hundred(?) years to do the while loop for this tho ;))