10 ms·
> E-mail marketers will no longer be able to get any information from images—they will see a single request from Google, which will then be used to send the ima
by codeflo 13y ago
> E-mail marketers will no longer be able to get any information from images—they will see a single request from Google, which will then be used to send the image out to all Gmail users. Unless you click on a link, marketers will have no idea the e-mail has been seen.
Absurdly wrong, marketers already use a unique image URL for each email recipient, and Google has no way to know that all of those point to the same image. So they won't see "a single request from Google", they'll see one request from Google per successful delivery to an inbox.
Now, an open question is if Google will make that request when the email is actually opened, which would allow marketers to determine if and when the email was read by the user, or if Google will make the request as soon as the email is received. The latter would enhance users' privacy at the cost of bandwidth for Google, but early tests indicate that they don't actually do that, waiting for the user to click the email to make the request.
I'd like to add that there's no possibility the Gmail team is stupid enough to not have considered this. They must know full well what they're doing, and marketing this as a privacy enhancement when it's actually detrimental to privacy is willfully dishonest.
- ceejayoz 13y agoIn fairness, it's a partial privacy improvement - it masks IP and user agent, as well as repeat opens.
- ben0x539 13y agoIf no actual http request between the email recipient and the sender happens, doesn't that also imply that the sender has less opportunity to do all the regular http user tracking stuff to associate the browsing session with that email address? That seems vaguely beneficial.
- vilhelm_s 13y agoIt doesn't even mask repeat opens: as buro9 found in the sister comment, the image does not get cached, Google's proxy will request it again each time it is viewed.
- gabriel34 13y agoI believe if you want my location and a confirmation I've read your email you should ask me, not shadily collect information
- mvanveen 13y agoGoogle was already image caching external email images, IIRC. So as far as amount of raw information being leaked, I think today's feature launch represents a step backwards.
- deleted 13y ago[deleted]
- deleted 13y ago[deleted]
- ryankshaw 13y agoit does, indeed, send the request when it is actually opened. see: https://news.ycombinator.com/item?id=6895876 https://news.ycombinator.com/item?id=6895876 /me goes to go find and turn on the "block third party images" feature.
- xux 13y agoIf they wait until the user opens the email to cache the image, then it's actually BETTER for marketers. This essentially allows the marketer to track whether the email was opened by default.
- slig 13y ago> This essentially allows the marketer to track whether the email was opened by default. And the spammer. Unless they decide to not load images by default from untrusted senders.
- trevyn 13y ago> Google has no way to know that all of those point to the same image. Try again. :)
- slg 13y agoYeah, I am not sure where the original claim is coming from. It isn't that hard for Google to simply follow the links and compare the files as part of the caching process. So unless marketers start customizing the images in addition to the links, there isn't any reason why Google can't cache the images together. And even if marketers do start customizing images, hasn't Google gotten pretty good at comparing very similar files? Isn't that how Google Music works without having a copy of every single individual upload?
- codeflo 13y agoOnce they followed the links, the tracking has already happened. Any deduplication after that only helps to reduce Google's storage costs.
- taeric 13y agoI believe the claim was more of "there is no way for google to know those all point to the same image without following the link."
- untog 13y agohasn't Google gotten pretty good at comparing very similar files? Sure, but they'd have to download the file first. At which point the tracking has succeeded.
- jessaustin 13y ago...the tracking has succeeded. Succeeded in what, confirming that Google is still in operation? Please note that Google doesn't even have to confirm the receiving email address is valid in order to get the image.
- 13y ago
- surrealize 13y agoGoogle could retrieve and cache a random sample of the images and hash them. Then they could note which image links go to the same images. And they could identify image links by noting that a bunch of emails have the same structure except for one image link that's different for every recipient. So I think they actually could take a pretty good stab a deduping images with unique tracking URLs. It might not be perfect, but even if it works 95% of the time they could still kill the profitability of the unique tracking image technique.
- codeflo 13y agoEasy: Make the "Dear <username>" part of the email an image. Boom, deduplication fooled. (Also, they don't currently seem to do anything like you're suggesting.)
- surrealize 13y agoWith google's machine learning brain trust, I think they could still do a pretty good job deduping. Maybe not perfect, but I'd bet on them to win an arms race. Edit: ah, codeflo and EGreg are right. I was just thinking about the task of determining that the images serve the same role in each message (which I'm sure google could do a good job of). But (as they point out) in the "Dear <user>" case they'll still have to show the right image to the right user. Although, as Nacraile and jaxn say, if they load all those images eagerly they'll remove the value of those unique tracking images, and impose a cost on the sender.
- codeflo 13y agoThere's a huge difference between something that they might do, in theory, at some point, using vaguely magic machine learning technology, and what that they are currently doing, right now, to address the privacy concerns over a change that they already rolled out to millions of users.
- EGreg 13y agoAre you saying Google will try to guess and reconstruct "Dear Marge" in the same font as "Dear John" instead of requesting both from the origin server?
- MrClean 13y agoConsider this @codeflo: 1. Google may cache all images in all emails sent to gmail.com instantly and regardless of the existence of the address. This would remove the possibility for marketers to check user timestamp, remove user data from request and hide user email existence. 2. Google does _not_ need to save each image from each unique URL separately, all they need to do is fetch each image and check against an already existing (mega)array of images they've fetched. This greatly reduces storage needed, but doesn't do much for the bandwidth requirement, but they won't care about bandwidth in all their Googleness. 3. The single most important aspect of this change has been omitted in the article, and in your comment: This change completely eliminates the risk of CSRF attacks by spammers and the likes. CSRF attacks are still number 8 on OWASPs list of top 10 attacks. My three cents ;)
- j_s 13y agoAlso, no more cookies that were set when loading those images
- cstrat 13y agoI thought that 99% of the images included in emails for tracking purposes are single pixel transparent GIFs - so no biggie in working out which ones those are...
- martin-adams 13y agoThen they switch to two-pixel transparent GIFs as a workaround.
- cstrat 13y agogod damn marketing geniuses.
- vidarh 13y agoIn the e-mail campaigns we send every image is a tracking image. It all just goes into our log files an is then post-processed, so the additional cost of processing every image is minimal compared to the rest of the cost of the send. Using a separate tracking pixel is pointless unless you for some reason want to let some third party track the opens (which some people might, e.g. to prove certain open rates)
- logicallee 13y agoThere is a simple way around that. You are saying that the fact that Google retrieves images means that the mail was delivered. But this is only true if Google retrieves images only from delivered mail. If Google retrieves every image in every mail sent to @gmail.com, @googlemail.com etc, then the only thing the retrieve tells you is that you spelled "gmail.com" correctly - nothing about whether there is a mailbox there or whether it was delivered.
- jlgaddis 13y ago> If Google retrieves every image in every mail sent to @gmail.com, @googlemail.com etc, ... But they don't. If they retrieve the image, the account exists (they reject mail after the RCPT TO: stage for non-existant accounts).
- logicallee 13y agoand presumably bounce it in other cases? I mean this could be a way of ensuring that certain addresses exist without needing to be able to receive the bounced mail if they don't. But how is that an interesting or useful vector? If Google does this with every mail regardless of the inbox it goes to (spam etc) then it doesn't tell you any more information than you learn from not receiving a bounce. However, I could imagine a scenario where the bounce address is wrong anyway (spoof) - is this really that useful for anything? I mean, presumably most combinations of common first and last name plus two digits go to a registered mailbox. How does being sure that it is registered (but knowing nothing else) without having to be in a position to receive the bounce, mean a compromise? I'm open to the possibility that it does - but I'm not seeing it. EDIT: another possible area of concern is that you can get Google to visit an address just by sending mail to johnsmith@gmail.com and calling the link an image. But can't you already do the same thing with the Google bot by including a link causing it to probably visit? This could be more instantaneous and hide the actual referring source of the visit behind an email, but I don't really see how this can be used for anything. For example if an extremely malformed server performs actions on the basis of a simple http GET then I guess you could craft that command into an image url, send it to any gmail address, and then Google will do your dirty work of actually visiting that link. But, really, is this a vector that is dangerous for anything? Don't URL's already get random Google traffic?
- 300bps 13y agoAs someone who co-invented email image bugs with a million other programmers over 15 years ago, you are absolutely correct. (for the curious, my reason for coming up with it was to tell when a customer who requested a car insurance quote from the company I worked at read their email which instantly initiated an outbound call to them). The entire article is just plain wrong. For instance: This move will allow Google to automatically display images, killing the "display all images" button in Gmail. Go ahead and do that, Google and you'll bring Web bugs back completely. How about this: 1. Marketer embeds a "jpg" file whose filename consists of a GUID that matches back to a user. 2. When you load that "jpg" file, it gives you an image unique to that user - maybe an MD5 hash of their GUID filename or some other thing unique to them that holds no value. This would uniquely identify that a user reads their email. How would Google stop this? They can't cache it for multiple users. Both the filename and contents are unique to the user. If you think they could otherwise detect it for this particular situation, I could think of 100 other ways to do this that would not be able to be tracked in the same way. There's no way Google or any other email provider can legitimately automatically load email images and not open the door to web bugs. No way.
- EGreg 13y agoThere is ONE way: fetch every single image right away regardless of whether the email is even valid. And then don't store the cached version if there is no valid email. Otherwise store it and display to the user. Since every result will be a positive - false or not - no information is revealed to the marketer AND the images are displayed.
- 300bps 13y agoThere is ONE way: fetch every single image right away regardless of whether the email is even valid. And then don't store the cached version if there is no valid email. Otherwise store it and display to the user. Sure, in theory that works. In practice, you make an easy may to do a Denial of Service attack against Google or innocent third parties so it in actuality would never work. Send a million emails to Gmail accounts that each have fifty links to 1 MB JPG files hosted all over the Internet. The size of the file you send to Google in the email is what - 1k? The size you are making Google download is 50 MB. This is a 50,000:1 attack ratio. You could take Google down with a 56k modem. You could also launch a denial of service attack against any other target, courtesy of Google. Send a million emails to Gmail users that downloads JPGs on a target web server. You can even make up the JPG names to be non-existent. Again, figure your email is costing you 1k in transmission and Google is putting tremendous strain on the target server downloading 404 error messages.
- buro9 13y ago> Now, an open question is if Google will make that request when the email is actually opened, which would allow marketers to determine if and when the email was read by the user, or if Google will make the request as soon as the email is received. The latter would enhance users' privacy at the cost of bandwidth for Google, but early tests indicate that they don't actually do that, waiting for the user to click the email to make the request. I've just tested this. The image was retrieved when I viewed the email in Gmail. The tracking info basically comes back as "anonymous" and viewed from an unknown location. The image was retrieved twice even though I only viewed the email once. Currently I'd say that seeing the image being viewed is still valuable. I'm sure Google could move to proactively fetching the images in future destroying even that value. The image is cached (via browser headers), but isn't aggressively cached (via reverse proxy). Fresh views on different browsers or a later session would still result in a request for the image back to the source server... registering the view again. My personal view on all of this is that this is a bit Microsoft... you know, convenience and features over security and privacy. For me, this "feature" leaks data about what I view to 3rd parties where today I block all images and do not leak that info.
- curveship 13y agoWait, do I understand you correctly that Google fetched the image even before you requested that images be displayed in the email? That seems like a boon to marketers, not a bane.
- buro9 13y agoA dialog popped up when I logged into email, and before I'd seen this thread. That dialog was much along the lines of "We'll now show you images in your email automatically." with a big "OK" button. I don't recall whether there was a less prominent "No, thanks" as I was only logging on to reply to one question really quickly. I suspect this is a UX anti-pattern. I've gone back into my settings and changed it back to "Don't display images".
- 3JPLW 13y agoThe alternative to "OK" is "Settings," with a link that just dumps you into the main settings view… leaving you to find the Images section to change it back. Quite an opt-in.
- jfoster 13y agoI think any email marketers who want to get around this easily can. Just change robots.txt to not give permission to Google fetching the images. Copyright will laws will prevent them from wilfully ignoring that, presumably.
- cdash 13y agoSo they strip the images and show you a text only email and tell you to deal with it.
- jfoster 13y agoTrue, that would close off this avenue. Would be interesting to know if that's how they handle this currently.
- Dylan16807 13y agorobots.txt is for crawlers, it would not stop an email client from rendering the email on behalf of its user upon receipt.
- jfoster 13y agoSo then what gives Google the legal right to fetch, store, and re-serve the images?
- Dylan16807 13y agoThey're just acting as an email client. Alice sends email to Bob, Bob is granted rights to fetch and view images. Bob uses gmail, so he passes the rights onto google to do part of that serverside.
- jfoster 13y agoI'm not a lawyer, but I don't think copyright law allows for Bob to pass on that right.
- deleted 13y ago[deleted]
- deleted 13y ago[deleted]
- FiloSottile 13y ago> Now, an open question is if Google will make that request when the email is actually opened, which would allow marketers to determine if and when the email was read by the user, or if Google will make the request as soon as the email is received. I reply to this question here. Sadly, a request is done each time and only when a user loads the image. https://news.ycombinator.com/item?id=6898087 https://news.ycombinator.com/item?id=6898087 Which makes the Ars article even more wrong.
- bloaf 13y agoIf google cached them as soon as they were received wouldn't they, in some cases, effectively perform a DoS attack on whatever was hosting the image? I.e. if a spammer sent out 1,000,000 emails w/ images they were hosting, they would immediately receive 1,000,000 requests for the image.
- FiloSottile 13y agoI think such a behavior is really easily spotted and filtered, if not already blocked by spam filters. Google has the technology to do that, the question is whether they want to put in the investment of storage required to actually guarantee users privacy, of if they just want to spend the least amount they can get away with...
- anonymic 13y agoGoogle could technically do a similarity measure between images coming from similar links (or similarly worded emails) and then provide the cached version. But then advertisers could hide random information in the image (steganography), making two visually similar looking images but with dissimilar enough content, that Google will be have to up the ante. And so on and so forth. But you are spot on them being "willfully dishonest". The way I look at it, they are at this point trying to push the barrier and see where users will protest enough that they need to roll-back.
- twakefield 13y agoI just did a quick test using Mailgun and recorded the results [1]. The TLDWatch is that Google does give you back the open event and they also give you the email address of the open since this data is encoded in the URL. The data that is not accurate is what you would expect: IP address, geo-location and user-agent string. We'll hopefully write a more extensive blog post shortly. [1] http://blog.mailgun.com/post/gmail-open-tracking-test-at-mailgun/ http://blog.mailgun.com/post/gmail-open-tracking-test-at-mai...
- dspillett 13y ago> The data that is not accurate is what you would expect: IP address, geo-location and user-agent string. Also any cookies present in the user's local environment from other actions (images from the same ad network in other emails or from visiting web pages that use the same ad network) are not going to be sent, so tracking you between locations is going to be neutered somewhat.
- drawkbox 13y agoGoogle gets huge benefits from it but I think in the end this helps advertisers as well as still being able to track via the image url and a unique url/id. Good deals usually help both parties with some give and take. Pros: - Since all images flow through google, phishing and other malware attacks could be subverted. - Images will be hosted faster in many cases (possibly less cost to run newsletters). - Less connections on your server/cdn from multiples sources but google singularly. - Still able to identify users legitimately but new users and newsletters will have more trouble getting your information initially. Cons: - Re-views later will not be tracked if out of cache, it will come from google the second time if is hasn't been purged (re-views are not big on newsletters anyways) - Google getting all this data as well as your company - The obvious 'national security' reasons - Limited location and meta information
- gohrt 13y ago> 'national security' reasons Which? The email already has the image URL. What extra leak is there if Google fetches the image?
- tjoff 13y agoFewer destination IPs is not a pro. Also, google already has this data so why is that a con? There is no national security reason. Bottom line: Christmas came early for spammers this year.
- malandrew 13y agoI would hope that they cache the images before delivery to any inbox regardless of whether the inbox exists or not.
- nikcub 13y ago> and Google has no way to know that all of those point to the same image. They could do content hash-based caching rather than URL-based caching. It would be more private, as the email senders would have to generate a unique image for each recipient.
- ukd1 13y agoIt's impossible to know if the content for a different URL is the same without fetching it first, your point won't work.
- jfoster 13y agoThe behaviour might change in the future, but I think an article in the Gmail help centre has some answers on the initial implementation of this: https://support.google.com/mail/answer/145919?hl=en https://support.google.com/mail/answer/145919?hl=en "In some cases, senders may be able to know whether an individual has opened a message with unique image links" suggests Google (at least for now) fetches the images upon opening of the email.
- arkj 13y ago"Absurdly wrong, ....,and Google has no way to know that all of those point to the same image." I think the guys who implemented image search has better ways to figure this out.
- vladtaltos 13y ago> Absurdly wrong, marketers already use a unique image URL for each email recipient, and Google has no way to know that all of those point to the same image. wrong. you can of course do simple image processing and identify similar images. if the marketers only change the name(url) of the file and not the content one-bit, you can trivially compare the hash of the file... even if the marketers change content, assuming the marketer sent the image %90 same and 10% customized per person, you can borrow techniques from image compression domain to compress this humongous data very efficiently.
- ukd1 13y agoAnd how do you get the image to compare it? By requesting it, which means you have to ask for it from a server, thus identifying that the image has been loaded.
- tomkin 13y agoA few years back I shared office space with an "email marketer", and even then he would tell me that each image URL contained a hash as part of the file name. This hash is linked to the email address. How would they be able to prevent that? Even if Google pre-fetches it, it still would (possibly harmfully) confirm they had saw it. Even if they didn't.
- getdavidhiggins 13y ago> They'll see one request from Google per successful delivery to an inbox But if the images are pre-cached before the user opens it, there is still no way of knowing the E-mail was read. Unless the image is cached upon opening, which is a bit counter-intuitive.