9 ms·
How I attacked myself using Google and I ramped up a $1000 bandwidth bill
- ltcoleman 14y agoThis was an extremely interesting article. I hope Google rectifies this type of behavior. This is actually pretty scary.
- typpo 14y agoThis isn't Google's fault. I could get a couple machines and do the same thing if I had a list of 250 gigs of files. And I wouldn't be limited to once every hour.
- bri3d 14y agoBut those machines wouldn't be Google's, and you would probably either be committing a crime (building an exploited botnet) or using someone's money (be it your employer, university, or yourself) to run them. This is pretty clearly a different league of attack: rather than attacking systems or spending money, you'd be exploiting a Google feature to use Google's resource for free to incur a giant amount of data transfer.
- tantalor 14y agoIt would almost certainly be criminal regardless of the method (Google, EC2, botnet).
- igorsyl 14y agoTry http://cloudability.com http://cloudability.com to keep track of your cloud costs. Note: I am not affiliated with them.
- ltcoleman 14y agoThanks for the link!!
- TwoBit 14y agoThat such a simple thing on your part could result in such an extreme and expensive result implies something is wrong with the design of the system you're using.
- ma2rten 14y agoI was going to post exactly the same thing. It's very lenient of him to blame himself instead of Google for this.
- Yarnage 14y agoI'm pretty surprised Google didn't have the client download the images instead. Wouldn't that be a better solution or am I missing something here? Pretty interesting though and if this becomes a big enough story you can bet Google will be changing something; the last thing they need is someone using Google Docs to DOS websites.
- recursive 14y agoPerhaps the server needs to know the image dimensions for page layout.
- Yarnage 14y agoThey use quite a bit of HTML 5 so they should be able to do this client-side I would imagine. But I dunno.
- bibinou 14y agoGoogle Docs is in HTTPS so they need to proxy the assets like Github does : https://github.com/blog/743-sidejack-prevention-phase-3-ssl-proxied-assets https://github.com/blog/743-sidejack-prevention-phase-3-ssl-...
- Yarnage 14y agoInteresting. I never thought of that. It seems almost silly that you're bringing insecure content over a secure channel like that but then again it is only images.
- EvilTerran 14y ago"only images" As long as they don't do WMFs, with their code-injection-by-design functionality... http://en.wikipedia.org/wiki/Windows_Metafile_vulnerability http://en.wikipedia.org/wiki/Windows_Metafile_vulnerability
- jonknee 14y agoThe function lets you specify height/width, passing it through makes some sense. It's also usually (though not in this case) to not hot link--a document can be viewed by thousands of people. There are some privacy issues too, viewing a Google Doc could have your browser making requests to unknown servers that could return whatever they want.
- timwang 14y agovery interesting findings, good read.
- gojomo 14y agoBut, why re-downloading every hour? Does merely having the spreadsheet passively open in a browser trigger that, or was some other process re-loading the spreadsheet every hour? (If the former, I wouldn't be as forgiving of Google. I understand the desire not to cache possibly-private data, but proper URL design and conditional GETs should be able to prevent the entire download on an automatic hourly schedule. And even if the latter – the author had chosen to reload the spreadsheet each hour – I'd want Google's design to allow browser-side caching to work for such embedded and/or generated images.)
- ardillamorris 14y agoThere must me something else to it because if the spreadsheet is sitting passively open there would be no need for Google to go an pre-fetch again.
- deleted 14y ago[deleted]
- jonknee 14y agoAbout an hour is the standard time for external data in Google Spreadsheet to be refreshed. I've come across this with JSON data.
- gojomo 14y agoEven if noone has it open in a browser? Or they do have it open, but haven't interacted with it? In either case, this seems odd, unless the URL is especially noted as 'volatile', and/or there are other parts of the spreadsheet that might trigger conditional calculations/notifications based on that URL's contents. (And don't S3 resources have last-modified-dates or etags for conditional GETs?)
- jonknee 14y agoI don't remember, but it wouldn't surprise me, that way it updates correctly for offline mode and if you have conditional logic it can work on changed data. It also improves page load time, which is very important for Google in general but doubly so for an online office suite (Excel opens quickly, so should a Google spreadsheet).
- RobertKohr 14y agoCan you limit bandwidth with AWS? Also, why would the spreadsheet be calling these images every hour. Did you have the spreadsheet open? Does google do this call even when no one is viewing the spreadsheet?
- justincormack 14y agoYou can put a robots.txt in the bucket.
- Panos 14y agoWhich will be ignored by Feedfetcher :-) Plus you cannot put a robots.txt at s3.amazonaws.com so if the url is accessed through the https://s3.amazonaws.com/... https://s3.amazonaws.com/.... url, the robots.txt will not work.
- simonw 14y agoYou could put robots.txt in the bucket if you address it using the http://mybucket.s3.amazonaws.com/ http://mybucket.s3.amazonaws.com/ alternative URL scheme - a robots.txt in the root of the bucket would then be available at http://mybucket.s3.amazonaws.com/robots.txt http://mybucket.s3.amazonaws.com/robots.txt
- Panos 14y agoYes, that would solve the issue of not being able to have your own robots.txt file and I did not know about that. On the other hand, Feedfetcher would still ignore the robots.txt
- justincormack 14y agoGoogle's justification for ignoring this is very weak.
- 14y ago
- deleted 14y ago[deleted]
- SagelyGuru 14y agoScary indeed. Many thanks for the warning. I was contemplating starting to use Amazon Cloud and some Google tools but definitely will not touch any of it now. Who knows how many other traps like this are laying in wait for the unweary? All this automation is very well but this illustrates the dangers of running stuff on other's machines of unknown complexity and out of your own control, while having to pay for whatever may happen. Not for me, thanks.
- dasil003 14y agoI don't think the OA would want you to take this as FUD. These are incredibly powerful tools that give you more control than the alternatives (AWS anyway), and yes unintended consequences are possible, but this really is a freak occurrence.
- K2h 14y agoI loved your reference to the huge Russion bomb Tsar Bomba. I for one had never heard of it, and it made a great metaphor. [1] http://en.wikipedia.org/wiki/Tsar_Bomba http://en.wikipedia.org/wiki/Tsar_Bomba
- DanBC 14y agoCheck out the Nuclear Effects Calculator; You can see what effect various historical bombs would have. (They include Tsar Bomba.) (http://nuclearsecrecy.com/blog/2012/02/03/presenting-nukemap/ http://nuclearsecrecy.com/blog/2012/02/03/presenting-nukemap...) (http://news.ycombinator.com/item?id=3624714 http://news.ycombinator.com/item?id=3624714)
- tocomment 14y agoThis really underscores Amazon's glaring omission of a billing cutoff on Amazon web services. How hard would it be for them to let me say, cut off my services at $100/month? This is the main reason I'd never use AWS to host anything public.
- deleted 14y ago[deleted]
- deleted 14y ago[deleted]
- eli 14y agoI guess that would be a nice feature to offer as an option, but most people would not want their web service cut off if they get a spike in traffic.
- adgar 14y ago> most people would not want their web service cut off if they get a spike in traffic. Only if they are sure enough about their finances and their web service to be confident that the spike will be profitable.
- scott_w 14y agoThat's where the "option" part comes in. People who aren't worried about the cost simply don't turn it on. Those of us who can't afford an extra £1000/month (or more if it's a dedicated attack) would love to be able to just turn it off when a threshold is hit. Users may lose access, but at least you don't go bankrupt in the process.
- tybris 14y agoEnd of the article: "PS: Amazon was nice enough to refund the bandwidth charges, as they considered this activity accidental and not intentional. Thanks TK!"
- pavel_lishin 14y ago
- richardlblair 14y agoIt's really awesome that Amazon was reasonable and refunded the charges because they were accidental. I mean, technically it was still your fault, so it would have been easy for them to be jerks about it.
- sohn 14y agoThey wouldn't be jerks if they asked for the charges; you still generated the traffic and they had to pay for it.
- EvilTerran 14y agoJust as "legal" does not mean "ethical", "contractually permitted" does not mean "not a jerk move".
- jmj42 14y agojust like asking you to pay for the services you used is not a jerk move. Amazon did a nice thing, but the services were used. Asking to user to pay for those services (even if it was a mistake), would not have been a "jerk move"
- Karunamon 14y agoBandwidth pricing is a funny thing, in that it's only metered because it's convenient to do so. You haven't consumed any kind of finite resource by moving 1GB or 1000GB. There isn't any "use". And it makes good business sense besides. They can let the guy off and eat the probably less than a hundred or so bandwidth this guy actually cost them due to their upstream providers, get a good writeup and look better as a result, ..or kill his account, take him to collections, etc, etc, etc, (which would probably cost them more than $1K anyways), and lose a customer and get a PR black eye while they're at it. So yeah, it would be a jerk move, and a pretty dumb one at that.
- icebraining 14y ago
- rachelbythebay 14y agoFeedfetcher strikes again, huh? I never did find out why they were so interested in one of my images. http://rachelbythebay.com/w/2011/10/27/wtfgoog/ http://rachelbythebay.com/w/2011/10/27/wtfgoog/
- RandallBrown 14y agoIf I wanted to launch an attack on something like Instagram all I would need to do is put a bunch of images (hosted on instagram) into a Google Spreadsheet? Then the google crawler will come through and download them all once an hour?
- BryanB55 14y agoThat's what I'm wondering too, I don't see how this would work on a normal website since a normal website probably wouldn't be whitelisted by Google like AWS is... What am I missing here?
- RandallBrown 14y agoI don't think you'd be able to take down a website with this strategy. The OPs site never went down, it just cost him a lot of money. Even pretty small webservers should be able to serve up static image files pretty easily. But you could do a denial of service by making the service too expensive and all of your work would be hidden behind the anonymity of the google bot.
- OzzyB 14y ago+1 For Amazon for kindly reimbursing the overage charge. -1 For Google for creating what is the biggest threat to content providers by enabling easy-to-use DDOS attacks across the entire interwebs. </hyperbole> Seriously, is this what we have to look forward to when Google Spreadsheets, and God knows what else, become ever-more popular? Think about all the additional onerus costs that would be incurred by content providers as more and more Google Spreadsheet users hotlink images, mp3s, videos... This has to be a bad design decision by Google, there's no need to redownload assets by-the-hour, on-the-hour, regardless of whether the user's spreadsheet is open or not. Is it time to go back to the days of putting your web assets behind $HTTP_REFERER?
- hm8 14y agoLike the author notes, I don't think the problem was with Google. It was no fault of theirs. The reason: It's a fine line between maintaining privacy policies and managing such events. If google were storing/caching these links, there would have been an outcry from those worried about user privacy and stuff. About the by-the-hour downloads, well there again is a trade-off between providing data quickly and doing a lazy evaluation. I think, as the author notes, it was just an unfortunate event that was a consequence of good design decisions gone bad circumstantially and wreaking havoc for the author. It was certainly nice of Amazon to have made that refund. The resources (read bandwidth) were used after all.
- mooism2 14y agoGoogle are presumably storing/caching these photos for an hour before expiring and redownloading them. I don't see why Google don't redownload them using If-Modified-Since, and reusing the original downloaded photos in the case that the origin server says they've not changed.
- ricardobeat 14y agoI don't understand the privacy concern - a publicly accessible URL doesn't offer any privacy. It's the same as a transparent proxy.
- damoncali 14y agoReminds me of the email cannon we built on Gmail, not exactly on purpose. We needed a full Gmail account to do some testing on our company's email backup system. At the time that meant 8GB of email. And you couldn't just send a bunch of huge attachments, as it had to be like regular email. So we took a gmail account and signed it up for dozens of active linux related email lists (because they were easy to find). The result was an email account that got an email message every second or so in a variety of languages. We quickly realized that you could have some fun by forwarding that address to someone else's email account. Fast forward a few months, and Gmail smartly requires a confirmation before allowing you to forward.
- creamyhorror 14y agoOnce in middle school in the late '90s I created an email loop between two free email accounts that forwarded an extra copy of the email every iteration. It was an experiment out of curiosity and mischief. I don't think the server actually went down when I started the loop (at least, I can't remember if it did) - since the mailbox would have failed up and messages fail to be delivered. I think I was hoping to produce an "email cannon" similar to yours - an email spammer to be directed at whoever I liked.
- joelhaasnoot 14y agoAh yes, the days of computer labs in high school and figuring out that outlook lets you copy and paste emails in the Outbox. Just get a big enough collection and keep copying and pasting. Easy recipe for getting hundreds - thousands of emails.
- rorrr 14y agoKudos to Amazon for refunding the guy. I don't think any other hosting company would just take the bullet for the customer.
- anthemcg 14y agoQuite an interesting article. I didn't even think of it as a potential DDoS tool until he mentioned it. It makes sense but it seems like Google should have a built a fail-safe for something like that situation.
- deleted 14y ago[deleted]
- lubujackson 14y agoHere's the big money question, since I haven't used S3 - do you have any throttling control? Can you simply block the feedfetcher bot, or block repeated hits like this at all? It seems like S3 is a nightmare $$$ hole if there aren't some really robust tools to manage this sort of problem.
- drivebyacct2 14y ago>"But how come did Google download the images again and again?" "But how come did" indeed.
- drivebyacct2 14y agoCommunication skills matter, sorry.
- arunoda 14y agoBut you get the idea what he trying to say no? We all here are not speaking English as the mother language. Please excuse :)
- vladd 14y agoI cannot help notice that Hetzner offers 5000 GB/month AND a full dedicated server, for 39 EUR (51 USD) [1], so his traffic would have cost him at that rate a total of 100 USD if he were to use a dedicated server instead of Amazon. (before mentioning Amazon's scalability, consider that Hacker News is ran on a single dedicated server, and the moral of the story seems to be how not to scale especially when you don't want to) [1] - http://www.hetzner.de/en/hosting/produkte_rootserver/x3 http://www.hetzner.de/en/hosting/produkte_rootserver/x3
- robryan 14y agoOf course if you want to do away with any kind of built in redundancy and take over the sysadmin duties involved you can get the raw space and transfer cheaper.
- krelian 14y agoI have a VPS on buyvm which includes 2TB/month for $6. Each extra TB is another $2.50. This is not a special case, you can other providers with very good pricing for data transfer. I don't understand how the difference in bandwidth pricing can be so large. I'm also certain that Amazon doesn't pay as much for bandwidth as the guy from buyvm.
- true_religion 14y agobuyvm is overselling their bandwidth to you. Real uplink pricing is not 2 dollars per terabyte. If you actually tried to use say 13 terabytes from a single vm, it would work eithe because of rate limmiting on your uplink or they would cut you off.
- krelian 14y agoI realize they bank on their users not fully utilizing the allocated bandwidth but you can take as a somewhat similar example whatbox (or any other seedbox provider), they sell bandwidth pretty cheap and their users are in general huge bandwidth hogs.
- devs1010 14y ago"What I find fascinating in this setting is that Google becomes such a powerful weapon due to a series of perfectly legitimate design decisions." I call in to question that these are "perfectly legitimate design decisions", basically, if google thinks that the data is private or too sensitive to cache, then it shouldn't be this easy to have it automatically keep hitting a site like this. Google should have realized the potential for abuse here. I'm guessing it truly is an oversight on their part as I can't imagine they would want to waste all this bandwidth either, however its something they should figure out a solution for.
- cabalamat 14y agoI have a VPS. When i read stories like this, I think there isn't any point in going over to GAE or AWS. But lots of people do use these services, so what am I missing?
- stretchwithme 14y agokudos to Amazon for running time backward and letting you pull your foot out of the way. I can't help but think that if benign decisions lead to disasters like this in the cloud, how much destruction could robots wreak in the future due to similar benign choices?
- Freaky 14y agoAmusing thing about that page - it's full of '\$100', and uses Javascript to strip the \'s out, replacing them with empty <span> elements. Not sure I really want to know why...
- jyothi 14y agoSometimes a 509 Bandwidth limit exceeded helps! AWS can actually do a setting for max bandwidth per hour or so & alert early if there is suspicious activity.
- otterley 14y agoI'm curious to know whether the objects were stored with Cache-Control: or Expires: headers. Does having such headers make a difference? Clearly the client's not presenting If-Modified-Since: pragmas as I believe S3 honors those.
- arunoda 14y agoThis is the same kind of attack. I've demonstrated here. But for users of a popular analytic service. http://news.ycombinator.com/item?id=3873774 http://news.ycombinator.com/item?id=3873774
- arunoda 14y agoThis is the same kind of attack. I've demonstrated here. But for users of a popular analytic service. http://news.ycombinator.com/item?id=3873774 http://news.ycombinator.com/item?id=3873774
- _k 14y agoHe's getting 100+ requests per second. You could rate limit the IPs. The question is how many IPs is Feedfetcher using ?
- matthieupiguet 14y agoSeeing how popular the story is, amazon could not have made a better $1000 PR campaign than being classy and reimbursing him!
- nfonrose 14y agoYou can use http://cloudcost.teevity.com http://cloudcost.teevity.com to protect yourself from such situations (disclaimer : I'm the founder and CEO)