11 ms·
Robots.txt for the NYT has a specific exclusion for an 1996 news article
- fennecfoxen 6y agowell, let's see what it once said at that URL... oh dear. "A 21-year-old woman said Tarik Sansal raped her in the apartment's bathroom before 2 A.M. yesterday, the police said. Mr. Sansal was charged with first-degree forcible rape. The incident occurred in the apartment of Englebert Theuermann, a first secretary of the Austrian Mission to the United Nations..." "Addendum: March 18, 2008, Tuesday. The Times did not report on subsequent developments. Tarik Sansal has a certificate of disposition by the Criminal Court of the City of New York, saying that the case against him was dismissed and sealed on June 5, 1997."
- mjevans 6y agoThe incorrect reporting and update is newsworthy; but the article as it existed is problematic to have as a primary source on a search result. There should be some way of annotating the contents are deprecated and referencing another hyperlink which explains the currently known facts. This link might be to a court case for an authoritative result of non-conviction / assumed innocence.
- williamscales 6y agoWhat is incorrect about the reporting? And why should the contents be considered deprecated? A sealed court case is not linkable.
- fennecfoxen 6y agoThe existence of a sealed court case doesn't mean that there's been justice, either, particularly where international diplomats have been involved.
- colinmhayes 6y agoThe reporting isn't incorrect though? They shoudln'tve revealed his name, but he did have an accusation against him.
- mjevans 6y agoIt's incorrect in the sense of the summary or what someone who only saw the initial version might draw as a conclusion. Further the update mentions this initial version was _historically_ (Edit: from sometime after publication in 1996 and not updated until 2008!!!) left online and indexed by search engines well past when followups should have updated it with conclusions about that case. Hence it is 'incorrect' to present this article in a cursory search result, but it is correct to produce it as a link in relation to more specific searches, or better, to instead of this page, any designated URL link when a hit would point to this page.
- orangeoxidation 6y agoPublishing the names of suspects and accused is just a bad practice. The court of public opinion knows no presumption of innocence and can ruin the life of someone acquitted easily. If they have to name the name they should wait until after conviction.
- kristofferR 6y agoThat's how it's done in Norway and a bunch of other European countries, though the media usually don't even name convicted people unless there's major reasons/interest. Also interesting, and quite sad: https://www.youtube.com/watch?v=UQSPBQjaWhI https://www.youtube.com/watch?v=UQSPBQjaWhI
- Scoundreller 6y agoThe way it’s been explained to me is that if a journalist is going to name an accused, they should follow the case and publish the outcome as well. For minor things, that’s not worth their time so they won’t name them. Not a 100% practice, but I guess that’s what they teach in school.
- wyldfire 6y agoI dunno, it's tricky. Should it really be unlawful to report on the suspects identity if the government doesn't reveal the name of a suspect but the journalist(s) are able to figure it out by interviewing witnesses? Also tricky needle to thread habeas corpus while trying to protect the accused.
- elzbardico 6y agoThe public opinion is a court that almost never acquits and that only serves life-term sentences
- Veen 6y agoYes, but secret trials don’t have a great history either. At its best, the media can prevent miscarriages of justice and make sure the government follows proper procedure.
- pseudalopex 6y agoIt was about a rape case. The Internet Archive shows the article was updated in 2008 to say the case was dismissed and sealed in 1997.
- nieve 6y agoDiplomatic immunity extending to the point of sealing the case and suppressing reporting?
- gostsamo 6y agoNot necessarily. The crime scene was a diplomats residence, but the ddg results for the name are for some kind of New York businessman who might have been a guest. The dismissal of the case might have been an out-of-court agreement with lots of money involved.
- bellyfullofbac 6y agoWell, he's about to wonder why everyone's suddenly visiting his LinkedIn profile. Also his name plus his alleged crime gets an AP News article as the top hit.
- throwbigdata 6y agohttps://en.m.wikipedia.org/wiki/Apophasis https://en.m.wikipedia.org/wiki/Apophasis
- zaroth 6y agoSo much for the right to be forgotten. And this is a dismissed case.
- williamscales 6y agoThere is no right to be forgotten in the US.
- adventured 6y agoFortunately. I'm glad to be able to read the NYT story courtesy of the Way Back Machine and attempt to learn more about the context.
- Dylan16807 6y agoYou get to read this article, and a bunch of people have false accusations about them stuck at the top of google forever.
- FeepingCreature 6y agoWithout making a value judgment either way: the right to be informed is inseparable from the risk of being misinformed.
- Dylan16807 6y agoWhat does a "right to be informed" mean, exactly? Actions based on right to be forgotten don't affect what you're allowed to know or what you tell individuals, they affect what you're allowed to publish to the entire world. Much like publishing a photo of someone; sometimes you need permission.
- michaelmrose 6y agoThis is a poor analogy. There is no particular reason to assign someone priveleges over the public facts of their life for example convictions trials misdeeds. In fact by definition no evil doer would ever give such permission. Would you ask a rapist or their victim for permission to discuss the crime? RTBF isn't ownership of information about self it is a privelege for villains to censor their victims and the general public.
- williamscales 6y agoHere is the article: http://web.archive.org/web/20091128124216/https://www.nytimes.com/1996/06/17/nyregion/guest-at-diplomat-s-party-accused-of-rape.html http://web.archive.org/web/20091128124216/https://www.nytime...
- nabla9 6y ago... has a certificate of disposition by the Criminal Court of the City of New York, saying that the case against him was dismissed and sealed on June 5, 1997.
- solosoyokaze 6y agoWhy was the case sealed?
- arbitrage 6y agoprobably diplomatic immunity/coverup. bigger fish to fry, so not worth prosecuting a rapist. that's pretty messed up.
- solosoyokaze 6y agoThat’s how I read it too. Having the case sealed seems like an invitation to actually do some investigation. Instead the NYT played along and tried to bury the story (albeit in a technically ignorant way).
- causalmodels 6y agoI am not a lawyer but I am fairly certain that New York State automatically seals criminal cases which are dismissed due to a positive finding for the defendant. Source: https://www.nycourts.gov/courthelp/criminal/sealedGoodResult.shtml https://www.nycourts.gov/courthelp/criminal/sealedGoodResult... edit: a word
- gostsamo 6y agoI've added a special exclusion to a robot.txt file for a specific article. It was some years ago while in college. The article in question was about the presentation of an assisting professor who had some kind of misunderstanding with the campus newspaper and therefore the article wasn't especially positive in tone. Couple of years later I was the sys admin of the newspaper website and a letter arrived in my university email. The professor had found that I'm responsible for the website and had sent me a tearful story about how this article is ruining her life, because it is the top Google result for her name, and how she had spent thousands of dollars on scammers who had promised to change that, and she was asking me to remove the piece. Long story short, I forwarded the case to the newspaper editor at the time and she agreed to let me add a line to the robot.txt. Edit: newspaper -> newspaper editor
- gostsamo 6y agoOut of curiosity, I checked the student's newspaper website. It turns out that they've made redesign after my time and they've removed the robots.txt file. However, googling the name of the professor in question returns much more recent results and the article is hard to find. It turns out that Google's algorithm buries some stories over time. One more decade and noone will be able to find this part of her history unless they know what they are looking for.
- Scoundreller 6y agoI find that Google discriminates against old content just because. Old, to the point, websites hand-written in notepad.exe are rarely in the top 5-10, even when they have precisely the answer you’re looking for.
- nikanj 6y agoIt's incredibly hard to find any information that wasn't created this year. Not sure if it's Google discriminating against old sites, or if the new sites just have such incredible levels of SEO-fu that a static, just-contains-what-you-want site has no hope of getting selected for results
- snowwrestler 6y agoJust for future reference: adding a URL to robots.txt will not necessarily exclude that web page from Google, especially if it has already been indexed. To reliably exclude a URL from indexing, you have to serve a “no index” instruction with that URL, either in a meta tag or an HTTP header. And for this instruction to be read, the robot has to visit that page! So disallowing the URL in robots.txt can actually be counterproductive to de-indexing it. Google also offers a tool specifically for removing URLs from their index in Search Console.
- rafaelm 6y agoThat's right, but the url exclusion tool is only temporary, so the only way to do it correctly is with the noindex tag.
- aaron695 6y agoSo does Wired and other news sites. https://www.wired.com/robots.txt https://www.wired.com/robots.txt It's concerning someone working for the news doesn't understand why. This is their power of ruining lives forever. This is a new thing, it should be taught, it's not hard to understand. I don't know why journalist think they deserve respect when these things are not fundamentally in their ethos. There should be a better process than robot.txt and some news sites are doing better. Europe has brought in laws. But if journalist want to be thought of as more than writing blog spam, they need a better answer to this.
- michaelmrose 6y agoIf he is innocent it should be sufficient to anmend the article so it shows the truth.
- DangerousPie 6y agoIf you Google a job applicant's name and find several articles reporting that they have been accused of rape (with just a little note at the end that the charges have been dismissed), are you really going to give them a fair chance? I don't think anyone would look at them in an unbiased way, even if they tried to.
- michaelmrose 6y agoWhy does it need to be a little note at the end. Change the title to so and so cleared of charges of rape if its so and yes I would give them a fair chance. What is the alternative? I do not know anything about the particular case and make no claims about it but suppose the charges were dropped for political reasons or because the victim was intimidated? Shall we in such scenarios silence them entirely in order to protect a criminal?
- curt15 6y agoShould Wired simply be required by law to delete the articles?
- 6y ago
- brianpan 6y agoHere's a relevant Radiolab episode about the right to be forgotten. https://www.wnycstudios.org/podcasts/radiolab/articles/radiolab-right-be-forgotten https://www.wnycstudios.org/podcasts/radiolab/articles/radio...
- doe88 6y agoA follow-up on one of their article. More of it. Please.
- jchook 6y agoReminds me of companies posting copyrighted material in comments only to file a DCMA takedown. If you can’t get NYT to remove it, maybe you “know someone” who can.
- cglong 6y ago(2019)
- remux 6y agoThis was changed in 2019: https://www.robots-viewer.com/robots/checksum/64dc55e4da9cc8bbf49d7d4df914a0b3 https://www.robots-viewer.com/robots/checksum/64dc55e4da9cc8... robots.txt version from 2017: https://www.robots-viewer.com/robots/checksum/8029662cfb040c58493526dae9b8bd6a https://www.robots-viewer.com/robots/checksum/8029662cfb040c...
- makomk 6y agoProbably because that's around the time he seemingly really pissed off one or more former employees who went digging for more information on him: https://www.glassdoor.com/Reviews/Employee-Review-Romio-RVW24367667.htm https://www.glassdoor.com/Reviews/Employee-Review-Romio-RVW2...
- DyslexicAtheist 6y agothe US Department of State (DoS) during 2012/2013 in its robots.txt[1] excluded around 9577 documents which leaked into archive.org (already pre Snowden). The robots.txt file now is OK but not sure if content is still on archive. [1] https://pastebin.com/raw/RE2tpyR3 https://pastebin.com/raw/RE2tpyR3 8<-----------8<-----------8<-----------8<-----------8<----------- #!/bin/bash snapshots="20120713050942 20121013154343 20121010165822 20120921054221 20130413152313 20130113162428" # orig source http://state.gov/robots.txt but also on pastebin in case they delete it: wget --output-document=robots.txt http://pastebin.com/raw.php?i=RE2tpyR3 for x in `echo $snapshots` do for i in `cat ./robots.txt|cut -d ' ' -f2 | tr -d '\15\32'` do if [ -e `basename $i` ]; then echo "$i already fetched" else wget https://web.archive.org/web/$x/http://www.state.gov/documents/$i; fi done done
- rurban 6y agoThis was one of the very rare cases where even wikileaks took down one article in the global intelligence files, after the state dpmt complained. A very high profile case. In that specific country everybody knew, what the diplomate of the other very specific country did.
- bigbillheck 6y agoIs it related to this particular incident? I'm not finding anything immediate.
- rurban 6y agoYou wont find anything, as it was taken down a few days after publication. The NY Times article was taken down later, but had not much juicy info.
- deleted 6y ago[deleted]
- michaelcampbell 6y agoI have a specific exclusion in my robots.txt file, and also a cron-scheduled grep of my logs to see if anything actually hits it. I don't care if they do or they don't, but it's a way for me to exclude specific bots that don't honor my robots.txt file.
- brailsafe 6y agoScathing Glassdoor review for his current company entitled "Just Another One of Tarik's Victims" among a sea of obviously fake ones (yes I know GD is bs)
- alerighi 6y agoUsing robots.txt to avoid search engines indexing the page is the most stupid thing you can do. Not only it's not mandated by law that search engines have to follow the rules in the file, but also you are giving to the public a known file where you put all the URL that you don't want to be public. And everyone that wants to get some information on a site the first thing that goes to see is the robots file. The correct thing would be to serve pages with the appropriate HTTP header to disable indexing. Of course search engines are still not obliged to follow the header, just as they are not obliged to follow the robots.txt file, but you are not leaking more information that you need. Really, robots.txt file is only useful to reduce the load on the server by crawlers, it shouldn't be used as a protective measure!