Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

In case any github people are reading this: you also have an annoying approach to web crawling "robots". Your /robots.txt is based on a white-list of user agents with a human readable comment telling the robot where to request to be whitelisted. Using robots.txt to guide whitelisted robots (like Google and Bing) is against the spirit of the convention. This practice encourages robot authors to ignore the robots.txt and will eventually reduce the utility of the whole convention. Please stop doing this!


Robots.txt is a suicide note.

http://www.archiveteam.org/index.php?title=Robots.txt

My personal server returns a 410 to robots.txt requests.


I have no clue as to why the author of that shit is as angry as he is, but I have zero interest in his opinion until such time as he learns to show me the issues, and not just blindly assume that anybody who is not as enlightened as him is a blind idiot.


Okay.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: