17 ms·
“Magic links” can end up in Bing search results, rendering them useless
- lizardactivist 4y agoNot the least surprised. By accepting the EULA and using the service free of monetary charge, you are instead paying with the contents of your e-mails and adress-book.
- feet 4y agoAnother reason to run your own email server
- iJohnDoe 4y agoIf it’s Outlook doing the dirty work then running your own mail server won’t get rid of this problem. Outlook is the client reading the links regardless of mail server. Also, somewhat a different topic, people that run their own mail servers might also still use Outlook.
- rsbadger 4y agoExactly. I also noticed Bing had accessed some non microsoft tokens too. Even gmail accounts were affected. I assume some people have connected their gmail account to the outlook client?
- matsemann 4y agoIsn't the quick fix that you arrive at some page, and there have to press a button or load some JS to do some action? AFAIK most email providers (like Gmail) will also visit links. Therefore you shouldn't do actions directly on the GET request. For instance if you have an unsubscribe link and all you have to do is visit that address, most of your subscribers will be accidentally unsubscribed. Same if you paste a link in Slack/FB/Discord/Twitter whatever, they will visit the page to create a preview. GET requests shouldn't have side effects.
- rsbadger 4y agoYup this is true - I was just being lazy. But what surprised me was that Bing actually indexed them. (even though my robots.txt said not to)
- bigtones 4y agoCan you prove that by linking to a Bing search where one of your pages show up ?
- rsbadger 4y agoyeh... https://www.bing.com/search?q=https%3A%2F%2Fshoprocket.io%2Femail-confirmation%2F&qs=n&form=QBRE&sp=-1&pq=shoprocket.io%2Femail-verification&sc=1-32&sk=&cvid=40B93AB2ADAA459CA2BB788274F18028&ghsh=0&ghacc=0 https://www.bing.com/search?q=https%3A%2F%2Fshoprocket.io%2F...
- scandinavian 4y agoThe URL in the search is this: https://shoprocket.io/email-confirmation/34b35b1 https://shoprocket.io/email-confirmation/34b35b1... I don't see that in the robots.txt https://shoprocket.io/robots.txt https://shoprocket.io/robots.txt User-agent: * Disallow: /cdn-cgi/l/email-protection Disallow: /login Disallow: /register Disallow: /404 Am I missing something?
- 4y ago
- tyingq 4y ago>As of Feb 2017 Outlook (https://outlook.live.com/ https://outlook.live.com/) scans emails Makes me curious if only the free, online, Outlook does this. There's also paid O365 online Outlook and the fat client Outlook.
- rsbadger 4y agoExactly what I was thinking - what really worries me is it also seems to have happened to a lot of @gmail users too. I still can't figure out how Bing managed to find email tokens sent to gmail. Maybe users who connected their gmail account to Outlook...?
- kotaKat 4y agoOffice 365 just seems to make links useless for security now. Our 365 instance now turns every link into this massive monolith of safelink checking URLs through Microsoft, making literally every email undeterminable if it is a phishing attempt or otherwise without turning to pasting it into one of many online 'decoders'...
- tyingq 4y agoOh, probably this thing: https://docs.microsoft.com/en-us/microsoft-365/security/office-365-security/safe-links https://docs.microsoft.com/en-us/microsoft-365/security/offi... Though that is optional and configurable.
- tikkabhuna 4y agoAt work they enabled safelinks whilst all the mandatory training stated best practice was to check the links before clicking. Its a shame those links can't have an alttext to show the real link.
- cma 4y agoYou don't want to train users to disambiguate phishing with alt texts which can be spoofed in other contexts.
- 4y ago
- foreigner 4y agoThere's a difference between adding the URLs to search engine results and accessing the URLs to scan for malware. The latter is quite common, lots of email hosts do that. It's not clear to me from the post if the former is actually happening - the author doesn't state that they found the links in Bing's results, just that they were accessed by BingBot.
- rsbadger 4y agoI found them in Bing results
- tatersolid 4y agoAre the URLs being served with a “noindex” header? Blocking crawls with robots.txt cannot de-list items from Google or other search engines. > Warning: Don't use a robots.txt file as a means to hide your web pages from Google search results. If other pages point to your page with descriptive text, Google could still index the URL without visiting the page. If you want to block your page from search results, use another method such as password protection or noindex. From https://developers.google.com/search/docs/advanced/robots/intro https://developers.google.com/search/docs/advanced/robots/in...
- rsbadger 4y agoThey are now - I didn't think I had to as all the pages are naturally behind a login, it never crossed my mind that Bing would follow email links, let alone index them in search results.
- gunapologist99 4y ago> Warning: Don't use a robots.txt file as a means to hide your web pages from Google search results. I realize that you are just the messenger and not the progenitor of that policy, so not addressing this to you, but: that is ridiculous. robots.txt is basically useless.
- 4y ago
- swiftcoder 4y agoIsn't this just the "link preview" feature, that is enabled by default in outlook? Many email clients generate link previews so that they can display a thumbnail of the webpage. Would seem to be necessary to filter out those referrers from the validate link
- rsbadger 4y agoEssentially yes - but "previewing" link and sending that link to Bingbot to crawl and index is another matter. Imagine you send someone a "private" link to a file...Bing sees that and indexes it for the world to see. Not cool.
- Gigachad 4y ago“Private links” should be covered by robots.txt. The only case I see this happening is for those “anyone with link” shares and those are easy to cover.
- rsbadger 4y agoAFAIK, most of those "private" links are actually just "unlisted" but they're still public. I'm sure Bing is indexing those too...
- Gigachad 4y agoAnything private should ideally be put behind authentication. If that isn't possible, than robots.txt. Search engines are _meant_ to index everything that is publicly accessible and not blacklisted on robots.txt.
- swiftcoder 4y agoIs this actually the Bingbot crawling for the index, or do they just use the same bingbot code to generate link previews? If I have a well-tested crawler sitting here, I don't see why I'd write a brand new one just to fetch previews in outlook...
- andix 4y agoCan’t you block it with robots.txt or some similar method?
- rsbadger 4y agoIt was blocked by robots.txt but Bing chose to ignore it. I even tried "blocking" the URLs in Bing webmaster tools today and this was the response: "Block request denied We found that the URL submitted for block is important for Bing users and hence cannot be blocked through Bing Webmaster Tools. We recommend that the best way to block URLs in this scenario is to add NOINDEX meta-tag to the HTML header of the page."
- cube00 4y agoThat is baffling logic. Sure, they think they know best and want to ignore the wishes of the owner of the web site. Why then respect a NOINDEX meta-tag instead of robots.txt?
- rsbadger 4y agoExactly - seems the safest way is to explicitly block known bots by user agent from even reaching pages you don't want indexed.
- LinuxBender 4y agoNo but one could block it with basic auth. That is how I keep Discord and Valve crawlers off my links.
- bigtones 4y agoMicrosoft does this because they're security scanning / checking all links in every Outlook email for known phishing and malware attacks. If Bing has not seen the web page before and it's not in the Bing dangerous web page index it first needs to check it to make a determination of if it's a phishing/malware page by scanning/indexing it before returning that outcome back to Outlook to flag the email as dangerous.
- rsbadger 4y agoScanning something for malware and publishing it in search results seem like 2 completely different things to me...?
- Xylakant 4y agoBut there is nothing to indicate either in the post or in the referenced SO thread that the URLs are published to the search results. They are visited by bingbot, that much seems confirmed, but there’s no example where one of these results shows up in the public search results.
- Qem 4y agoThis reminds me of that time Bing was caught stealing search results from Google queries.
- sdfhdhjdw3 4y agoGoogle reads your emails. Whenever I buy a flight, google puts the date on "my" calendar. Just lets not pretend Microfsoft is especially bad at this, ok?
- jacquesm 4y agoBut that is your calendar, not something that is normally speaking visible to the whole web.
- sdfhdhjdw3 4y agoIt's not "your" calendar, it's Google's calendar.
- NGRhodes 4y agoBy that logic they are not your emails, they are Google's.
- decebalus1 4y agoWhich would be correct, considering that ownership implies full and complete right of dominion over said entity, which you simply don't have. You could be locked out of your account with no means of getting back access, you can delete your data but have no guarantees that the data has been deleted, Google may create 'derivative works' on your data (see the terms of use) without your permission or will provide data about your account to authorities, etc.. That is not ownership, that's renting.
- jacquesm 4y agoYou know perfectly well what I meant.
- Beltiras 4y agoDoes Bing observe robots.txt? If it does, that can put your token URL out of harms way.
- api 4y agoEverything spies on you unless proven otherwise. Seems to be a rule these days.
- mordae 4y agoCan this be exploited to confirm some action automatically? Or to prevent some site from being indexed by flooding it with invalid links?
- christophilus 4y agoThis would kind of break single-use links, no? That seems like a real nuisance.
- rsbadger 4y agoExactly. From what I've heard today it sounds like most apps have an extra step between the email link and the login, usually a JS step, to check for bots. Not adding that was my downfall I think.
- Beltiras 4y agoThere's also robots.txt
- oever 4y agoThe HTTP GET method is idempotent: it should behave the same way on multiple accesses. A single use link, e.g. for resetting a password or confirming a subscription, will usually show a webpage with a form that does a POST. Once that POST has been performed, the single use link is used up. Single use links will mostly have a one-time secret that should not be leaked. Mails that contain such links or any sensitive information should be encrypted.
- weberer 4y agoHow do you send mail to an Outlook user and encrypt it so Microsoft can't snoop on it?
- deleted 4y ago[deleted]
- diegoperini 4y agoYou re-ask the password on the visited page before presenting the form responsible for the one time POST call.
- jeroenhd 4y ago
- bezoz 4y agotldr: In 2022, whether you are a paying customer or a free customer, YOU are the product and you will be squeezed for all you got (not specific to Microsoft at all)
- rsbadger 4y agoI don't think that's very surprising for most people, the real takeaway is that not only will Bing read your emails, but they may also index any links you send and serve them in search results.
- bezoz 4y agoActually, you are right. I stand corrected, this really is a new low
- phendrenad2 4y agoWow, you really think that all of the Fortune 500 companies using GSuite/Office365 are being squeezed for everything they have got?
- OnePlusfans99 4y agoIt might lead to sensitive data leak as cloud storage links can also be crawled to Bing
- FrenchDevRemote 4y ago> It might lead to sensitive data leak as cloud storage links can also be crawled to Bing no what might link to sensitive data leak is fools who store sensitive data on unprotected links
- dementiapatie 4y agoAgree, in my experience storage buckets are always private by default, and you must take several specific steps to make them public, ignoring the very big warnings sprinkled in each confirmation page along the way. Are there any cloud vendors that don't follow this approach?
- rsbadger 4y agoAlmost all of them. A good example is Dropbox link you send to someone. I could generate this link to a private file in my Dropbox, email it you, and Bing (may) index it. https://www.dropbox.com/s/vucien2ns8jktga/denim%20bodywarmer.png?dl=0 https://www.dropbox.com/s/vucien2ns8jktga/denim%20bodywarmer... I doubt many people realise this when they email "private" links...
- deleted 4y ago[deleted]
- deleted 4y ago[deleted]
- rbut 4y agoI have observed this, but also found that BingBot modifies the query string parameters of your URL. It does this by changing a character of the URL, possibly in an attempt to find new pages? I noticed this because I generate links with a signed token to ensure integrity and started receving invalid token crash reports in Sentry, always from BingBot.. To fix this I had to move the tokens from the query string into the URL itself to avoid BingBot changing it. eg. http://mysite.io/do-action?token=shvgaaehr2rnyxhh-391-1 http://mysite.io/do-action?token=shvgaaehr2rnyxhh-391-1 to http://mysite.io/do-action/shvgaaehr2rnyxhh-391-1/ http://mysite.io/do-action/shvgaaehr2rnyxhh-391-1/ Anyone else noticed this?
- batch12 4y agoI would guess that this is probably done on purpose to avoid tripping one-time-use links. Seems like a good way to hide malware from the scanner though.
- throwaway14356 4y agoi imagine one could try use the location hash. it isnt send with the request
- liam_ja 4y agoI've finally found someone else who's seen this behaviour! I've noticed this too, and I found (in my case anyway) that Bing/Outlook seems to Rot13 the keys of the query parameters - is this what you're seeing too?
- mrjin 4y agoOkay, I thought M$ was just a little bit better than $G. It turned to be as bad...
- ipaddr 4y agoWhy would you think that? If M$ had the same position as google even more things would be closed source and more connected with law enforcement and less private.
- mrjin 4y agoI was naive to think so as M$ did not have as many obvious malicious moves as $G recently. But I forgot all those companies are there for money and for sure they will do whatever they can.
- pluc 4y agoI mean you can say that Twitter and Slack for example do it too, any service that generates a preview of your links, they'll crawl the URL you provide whether it's secret (eg sent in a private message) or not. Very very very few will stop at the "og:image" tags and such because why would they discard data about you?
- batch12 4y agoI have observed that twitter's bot hits links within seconds of being tweeted. The traffic comes from several locations, not all twitter ASNs. One interesting source is Apple. Their bot/scanner hits soon after.
- Aissen 4y agoAnyone paying for the firehose access can do this.
- batch12 4y agoThat's true of course. What's interesting to me is that they've decided to pay for this access and visit the links so quickly. It must be pretty expensive or hard to get if only around two-dozen companies pay for access to the data[0]. [0] https://www.washingtonpost.com/technology/2022/06/08/elon-musk-twitter-bot-data/ https://www.washingtonpost.com/technology/2022/06/08/elon-mu...
- Aissen 4y agoIt might be that Apple is paying for the firehose as a data-source to bootstrap its search engine. Don't they have one accessible via Siri already ? (I don't follow Apple tech very closely).
- baisq 4y agoI remember sending one-time use URLs in emails to customers and they would've expired by the time they clicked them because Outlook was opening them before they did. Yeah yeah GET is idempotent and I shouldn't do that blah blah. That's not the point.
- logifail 4y agoEven if Microsoft claim this is about security scanning, isn't it fairly trivial to configure your webserver to serve up different content depending on the User-Agent request header? BingBot scans the link, gets a dummy page with 'clean' content, Microsoft delivers the email message to the user, user clicks through the link with actual browser, gets phishing / malware content...
- hazmazlaz 4y agoYes, that is exactly what a motivated attacker would do to avoid their phishing site getting flagged as malicious. Here is a good article about how that is accomplished: https://rhinosecuritylabs.com/social-engineering/bypassing-email-security-url-scanning/ https://rhinosecuritylabs.com/social-engineering/bypassing-e...
- eli 4y agoSure, or even just ignore user agents if you know your target has this scanning in place, just send the malware to the 2nd click. It's not just MS. Lots of enterprise email security stuff works like this.
- plasma 4y agoIn my experience Gmail does this from time to time too, so any “load once” links won’t reliably work.
- sandermvanvliet 4y agoI've had to deal with this with e-mail verification links and Auth0. The user clicked the link after getting it in their mailbox but then Auth0 throws up an error page because the e-mail address has already been verified (by Outlook scanning). The problem becomes worse if for some reason the mail ends up in the junk mail folder so the user thinks they've never received the mail but when you check it looks like the e-mail address was verified successfully. That has caused a lot of annoying back and forth trying to figure out what the hell is going on. We ended up adding a custom page to handle e-mail validation so we could handle the situation where the user lands on the page and the address has already been verified. Super annoying.
- heipei 4y agoHonestly that's why I built the email verification page so the user still has to click a button on that page.
- dmw_ng 4y agoLinks like this are stupid regardless of Outlook's behaviour because they require a perfectly reliable client and network and user in a perfectly undisturbed flow. If I can't F5, if I double-click, if my mouse is wonky, my wifi is bad, my power goes out, my computer hangs, my DSL dies just after a click, if I accidentally close the tab.. there are any of a thousand reasons why abusing GET for a one-time-use page or redirect is horribly wrong. It takes incredible arrogance to continue using them in order to "improve usability" given all the obvious and common cases where they completely destroy usability. The difficulty for a provider to verify they aren't sending you to a phishing or browser 0day page barely scratches the surface.
- rsbadger 4y agoThe only purpose of this link was to verify that the email address is valid. Once it’s verified, you can login.
- TremendousJudge 4y agoI have seen services where you have to click a link every time you want to log in
- heipei 4y agoTo all the folks that suggest preventing opening of single-use links by robots.txt or user-agent detection etc: Just don't. There are dozens of tools at use throughout the various stages of an email with URLs being delivered that will go out and fetch websites. You have to design any confirmation dialog so the user still has to click a button to confirm, otherwise any one of these tools might inadvertently trigger your confirmation.
- mikro2nd 4y agoJust wondering how this might be an attack vector for fucking with Bing... I can think of a couple of avenues; the most elegant would be if the URL itself triggered something within the scanner/URL processor; next up would be the content of the target page attacking the Bing infrastructure. I'd guess the backend processing is sandboxed, but it seems like an interesting avenue that a malicious actor might explore. Don't try this at home, kids. :)
- aasasd 4y agoSure, we'll do this right after cracking Google through Googlebot.
- AtNightWeCode 4y agoMore likely scanning for vulnerabilities or generating of previews. Magic links don’t work anymore. Therefor most services sends a code or something that you have to enter on a generic page.
- mcv 4y agoSounds like anyone dealing with any sort of vaguely sensitive information through email, and certainly any corporation, should avoid using Outlook for anything. The article is about email verification links, which is a pretty clear case where this can be dangerous, but tons of other links can get emailed without being intended for a wider audience. Besides, the fact that Outlook shares anything related to the content of your email with the outside world is just completely unacceptable. (Should private links be sent over unencrypted email? Probably not. But lots of stuff gets emailed that's not super secret and yet also not meant to be shared outside the company.)
- rsbadger 4y agoExactly that.
- phendrenad2 4y agoOr maybe you shouldn't rely on security through obscurity and instead should add a robots.txt as has been in the web standard since 1997.
- rsbadger 4y agoI do have a robots.txt to block this directory. But Bing only listens to that for what to crawl, not what to index.
- 4rb1t 4y agoproofpoint does this too. Wonder if they add a separate header or query string param to distinguish from an actual user.
- ck2 4y agoAnd google harvests all your online purchase emails to log everything you've bought. https://www.techspot.com/news/80134-google-uses-receipts-sent-gmail-log-online-purchases.html https://www.techspot.com/news/80134-google-uses-receipts-sen... Just a reminder any email left on any online service over six months in the USA is allowed to be read by any law enforcement agency without a warrant. You'd think these services would have a six-month auto-delete feature but nope. There's good reason why there was a email server in the basement, everyone should have their email server where at least a physical warrant is needed.
- kornhole 4y agoI struggle to understand how private companies like mine are OK with MS reading all employee email and processing it through their AI. I get these daily creepy emails from MS saying that you said you would do this yesterday.. I have resorted to using burnernote.com, not to hide anything from my company but to hide it from MS who competes with us on some products. I guess burnernote.com will also not work anymore since it creates one-time links. We are monitoring you for your protection.
- deleted 4y ago[deleted]
- jeroenhd 4y agoOutlook will only send GET requests, which are idempotent unless you're ignoring the spec. A message saying "this code has already been used" after sending a GET request is a bug. I don't see the problem here, all services need to do is add a page that's says "welcome back, $Username, click here to log in!" that sends a POST request to do any serious confirmation without breaking any specifications. Microsoft claims the visiting not is BingBot but it's probably just SmartScreen system checking for malicious links/downloads/etc. like many cloud integrated security products do these days. I can set my browser to pretend I'm BingBot, you can't derive anything meaningful from the user agent. Unless you find your secret URLs in Bing's search results, your secret links aren't actually being monitored by a search engine.
- rsbadger 4y agoI’m fairly certain they are. My links ended up indexed in Bing search results. The only place they were ever rendered was in private emails to users. Bing should not be indexing that.
- jeroenhd 4y agoYou're right, it shouldn't. It's possible that they're fetching these URLs from their customers' browsing history and submitting those (external submissions follow different crawling rules, sometimes bypassing robots.txt). Bing's webmaster information says so, at least: https://www.bing.com/webmasters/help/webmasters-guidelines-30fba23a https://www.bing.com/webmasters/help/webmasters-guidelines-3... For a bit of added "fun", Google will do the same, but if you add a page to robots.txt and set noindex then they won't process the noindex parameter and external indexing sources might still generates search results: https://developers.google.com/search/docs/advanced/crawling/block-indexing https://developers.google.com/search/docs/advanced/crawling/...
- elric 4y agoThat's a bit of a narrow view on this problem. When sending a link to someone, you expect that someone to view the link. Not some random mail service. Who gave the mail server permission to access the page? What if it contains copyrighted material? What if it's one of the millions of pages which don't follow the HTTP design philosophy to the letter? This is a can of worms.
- progman32 4y agoOur company occasionally does "test phishes" to see how well people resist them. Every time, some of our most security minded engineers end up on the "clicked on the malicious link" lists, when all they did was forward the message to IT to report the phish. I'm wondering if the bingbot leak is the reason.
- martin_drapeau 4y agoIn the B2B SaaS where I work we started using single use codes to log in for certain account types (non-admins). No password. We send you an email or an SMS with a 6-digit number. Copy/paste it to log in. Very much like 2FA except there is no password. The session lasts 30 days. The user can disconnect of course. Curious what HN readers think. Is this secure? Sufficient?
- midislack 4y agoDon’t Windows users already just accept whatever from MS? This is what you get with a proprietary operating system. Whatever you’re served. Now stop complaining and look at the new ads in the Start Menu.
- anonym29 4y agoIf you're concerned about privacy, you shouldn't be using any Microsoft products, period.
- hansel_der 4y agotrue, but for the last 30ish years nobody cares about that opinion because it is profitable to use the smallest common denominator, get shit done and call it a day.
- zo1 4y agoJust wait till someone figures out how this "leaks" personal info. They'll very quickly remove it.
- egberts1 4y agoThat’s why I shield my URL with a casual password prompt because external email scanners has no business looking into the Enclosed email URL. Email address domain and enclosed URL domain are the same parent domain.
- creeble 4y agoHow long does it take for them to check the link? I sent an email with a unique link in it to my @Outlook.com account 6hrs ago, and there have been no visits to the link. The email is in my inbox (though I have not opened it). Does this only happen on opening the email (in the Outlook web ui)?
- windows2020 4y agoI've noticed Office 365 Safe Links makes an OPTIONS request, not GET. So, restricting the endpoint to GET, via [HttpGet] decorator for example, may be a quick resolution.
- HeavyStorm 4y agoWhile I think a Turing check can easily solve the problem without much friction, this only increases my hatred for Outlook scanning. The worse part - to turn it off, you also have to turn off junk mail protection (well, used to, it's been a while since I tried). Now, having my private links indexed by Bing is a bit too much!? I sincerely hope OP is mistaken and Bingbot is actually the outlook scanner.
- rsbadger 4y agoUnfortunately not - the links were indexed and shown in Bing search results
- yencabulator 4y agoFor what it's worth, here's Google admitting that GoogleBot causes POSTs: https://developers.google.com/search/blog/2011/11/get-post-and-safely-surfacing-more-of https://developers.google.com/search/blog/2011/11/get-post-a... Automatically triggered POST is not sufficient to keep the bots at bay. They seem to be implying that only automatically triggered POST are acceptable, but that was also >10 years ago. With the way things are going, it might be that any on-page confirmation buttons won't be sufficient to keep the bots at bay. Maybe it's time to fight back, check the user-agent, and serve the bots a CAPTCHA?