17 ms·
Apple downloads ~45 TB of models per day from our S3 bucket
- ebg13 7y agoIf you don't want someone else to do something that costs you money, you're going to have a bad time if you don't prevent them from doing it.
- Topgamer7 7y agoAmazon has the ability to charge the requester. Pass the buck on.
- mtnGoat 7y agoi dont think this works on public files though, does it?
- paulddraper 7y agoObviously not.
- idunno246 7y agoThis is the main purpose of the oft maligned allauthenticatedusers permission in the s3 console - anyone can read it they just need an identity to pay
- privateSFacct 7y agoYou can allow anyone to read it and pay as long as they authenticate with AWS so aws can bill them.
- bryan_w 7y agoHave you considered cloudflair?
- toomuchtodo 7y agoBackblaze B2 + Cloudflare = free outbound transfer (bandwidth alliance). Making objects in S3 publicly available and then complaining about their extortionate bandwidth charges is...silly.
- julien_c 7y agoI had intended to look into Cloudflare before but didn't get a chance. Looks like it's a good time now!
- joecot 7y agoBackblaze B2 + Cloudflare would be a perfect combination for hosting a static site. Unfortunately there's no way to map a Backblaze bucket to a domain, so even if you use cloudflare to point www.mydomain.com to it your files still show up at www.mydomain.com/path/to/bucket. But it certainly works if you need to CDN a bunch of large files.
- jrullman 7y agoYou can use Cloudflare Workers to rewrite the path.
- zenexer 7y agoThose get fairly expensive if you have a lot of requests. Page rules are free if you only need a few, though.
- joecot 7y ago3 pages rules are free. You need at least 2 page rules per website you want to use with Backblaze and Cloudflare, and that's for bare minimum workability. For a commercial company paying $5-10 a month to save TBs of bandwidth via Backblaze & Cloudflare is a no brainer. It's just unfortunately not very workable for small projects. Github Pages works with Cloudflare for static sites just fine though.
- soared 7y agoCharge, them, money?
- hi41 7y agoI read the Twitter post but did not understand what is happening. Can someone please explain.
- ProAm 7y agoApple employees are using their product, downloading lots of data, not paying for any of it, and the OP doesn't like it or can afford it.
- microtherion 7y agoI don't think it's employees as such — even Apple does not have THAT many machine learning people, and they wouldn't download models daily. Maybe a server farm, where each instance downloads a model when spinning up?
- phire 7y agoIt sounds like a misconfigured build pipeline which re-downloads the model for every single build.
- hinkley 7y agoIn the last days of my time spent in the XML salt mines, I got in on a conversation with the web masters at w3.org. You would not believe how many people and how many libraries pull from primary sources directly instead of using local copies of common resources. I found this conversation because I'd just finished fixing that in our code and taking about 5 minutes off the build process. Let me restate that: We were spending 5 minutes just downloading schema files. In an automated build. Every time, sometimes on several machines at once. At one point we were trying to convince him that intentionally slowing all requests down by say 500 ms would get the attention of people who were misbehaving. Anyone who was downloading it once and caching would hardly notice the 500 ms. Those running it once per task would be forced to figure their shit out.
- angrygoat 7y ago
- lacker 7y agoWell, you could contact them and make a very-likely-to-succeed case that they should pay you some money, or you could complain about it on Twitter.
- dharmon 7y agoI don't have high hopes for his business prospects if this is how he handles one of the richest companies in the world clearly having a high need for something his company offers. Maybe spend less time on Twitter and more on your business model?
- deleted 7y ago[deleted]
- httpz 7y agoThey're basically bragging they have something Apple really wants. Now they have a bunch of people at least interested in what they got. I'll say that's not a bad PR.
- joatmon-snoo 7y agoAlso a great way to throw out massive red flags to any enterprise user that cares about privacy and non-disclosure. IP address data is pretty sensitive information, and throwing it out there like this, even in aggregate, is not OK because of what it shows. No matter how much PR this gets, this goes both ways.
- capkutay 7y agobig companies are also notorious for reaping whatever they can take from smaller companies...and when its time for the smaller company to monetize..."whoops we don't have budget for that."
- mariomariomario 7y agoThis isn't sensitive information. Anyone with a BGP session can have this information. [1] [1] = https://bgp.he.net/AS714#_prefixes https://bgp.he.net/AS714#_prefixes
- vonseel 7y agoPretty amazing that a single company can own an entire block of IP space, if I understand this correctly. Approx how many addresses is this?
- alphagrep12345 7y agoWhat does hugging face do? Do they implement models from papers and make them available for free?
- physicsyogi 7y agoYes. I don’t know if that’s all that they do. They often port new Tensorflow models to PyTorch as well. They provide straightforward APIs, nice documentation, and clear tutorials. I use their stuff pretty regularly.
- codesternews 7y agoDo you know their business model? Looks like they are open source company. How they earn money?
- JustFinishedBSG 7y agoThey are a startup, 1 year ago their future product was chatbots / text platforms for entreprise. Don't know if they pivoted but I don't think so. My guess is that they don't earn money yet. They more or less are just out of school so it's a young startup. Actually studied with some of them and I'm still a student.
- rasz 7y agoTheir business model is someone like Apple buying them out.
- cuillevel3 7y agoAre those full downloads or just HEAD or range requests from some CI?
- ecnahc515 7y agoWhy can't they use cloudfront?
- ajay-d 7y agoAren’t all the authors of that paper from Apple?
- mrfusion 7y agoWhat’s the backstory on this? (Is it something I should already know)
- fitzroy 7y agoIn a few weeks he can just point Apple's IP range to a shared iCloud folder.
- emeraldd 7y agoThis looks kind of interesting: https://github.com/huggingface/pytorch-pretrained-BigGAN/blob/6ae20a35a051816d66811d85597033623a8ac888/pytorch_pretrained_biggan/model.py#L29 https://github.com/huggingface/pytorch-pretrained-BigGAN/blo... When you look further down you find: https://github.com/huggingface/pytorch-pretrained-BigGAN/blob/6ae20a35a051816d66811d85597033623a8ac888/pytorch_pretrained_biggan/model.py#L253-L281 https://github.com/huggingface/pytorch-pretrained-BigGAN/blo... And that's just a quick search for s3 in the repo. It would not surprise me in the least to discover a `from_pretrained` that points at one of the s3 resources being pulled. There's probably other stuff like that as well in the code that could be causing equally nasty heartache .. especially if non-persistent containers are involved.... (This is a WAG aka Wild A Guess) EDIT: Dug a little more and found: https://github.com/search?q=org%3Ahuggingface+s3&type=Code https://github.com/search?q=org%3Ahuggingface+s3&type=Code Unless I'm mistaken here, there's a crap ton of code that could be downloading models at runtime ... Which seems significantly less than ideal ...
- emeraldd 7y agoBased on this, they might not even realize they are downloading this stuff ...
- madisonmay 7y agoIt's also possible it's part of a docker build step or similar. Even if they're aren't downloading models at run time they may be loading s3 if their pytorch-transformers lib docker cache gets invalidated frequently.
- kortex 7y agoDocker builds are crazy wasteful in terms of bandwidth and compute. Right now I'm struggling with a project that builds ITK on demand every heckin time. I'm working through how to best integrate apt-cacher-ng, sccache, and a pip cacher. It costs me nothing to hit apt or pypi, but like, somebody is paying that bill. A little perspective goes a long way. I wonder if I could do something to just proxy all requests and cache those on a whitelist and stick it on my CI network.
- cs702 7y ago"Almost everyone" working on NLP uses one of hugginface's pretrained models at one point or another, sooner or later: https://github.com/huggingface/pytorch-transformers https://github.com/huggingface/pytorch-transformers It's so damn convenient, and so nicely done. And they keep doing neat things like this one: https://github.com/huggingface/swift-coreml-transformers https://github.com/huggingface/swift-coreml-transformers Kudos to Julien Chaumond et al for their work!
- megaremote 7y ago> Swift Core ML Why does he call them Swift Core ML? They are core ml, usable in swift and objective-c.
- jph00 7y agoThat repo is written in Swift, hence it is called Swift Core ML Transformers.
- BlueGh0st 7y agoFor anyone else initially confused, NLP in this context is "Natural Language Processing."
- ShteiLoups 7y agoAs opposed to?
- btown 7y agoA brief reminder: Whenever you publish code or documentation that might be used/scraped by the outside world, ALWAYS use a domain you own. If you're on Cloudflare you can instantly (and for free) create Page Rules to use Cloudflare as a CDN, redirect to another CDN, or black-hole or reroute traffic anywhere you want.
- dehrmann 7y agoWhen working on APIs meant to be used client-side (especially mobile clients) by different customers are partners, use one subdomain per integrator. If there's a bug in their integration, it could easily DDoS your servers, but DNS is an easy way to have a manual kill switch.
- btown 7y agoLiterally had this happen to us (website, not API) from a misconfigured partner last week - they accidentally misrouted unrelated click traffic through our servers. A 2 minute Page Rule and we not only saved our servers, we protected our partner's brand until they could hotfix. We could have done this with a rule looking at the path, but not easily something looking at obscure auth keys. Segmented traffic is happy traffic.
- techslave 7y agohuh? why publish if you don’t want it used?
- nathantotten 7y agoNot to mention that if Cloudflare CDN was in front of it this traffic would be free.
- lgats 7y agoI'm skeptical of the number of 500+ MB files the CloudFlare CDN would actually cache... Does anyone have any numbers on this?
- kelnos 7y agoIf a company the size of Apple finds this that useful, perhaps you should consider charging for your service, rather than just complaining on Twitter about the free usage you appear to have willingly given away? Or perhaps you have reached out to them, but are for some reason still complaining on Twitter to drum up PR or something? Regardless, this posting is ridiculously context-free to the point of being click-baity. (But hey, good job, I clicked on it anyway.)
- rtkwe 7y agoOn the flip side huge corporations like Apple should be more careful about not abusing free services set up by people not just because it's a dick thing to drown free services in traffic just because they're free but also it's a pretty big security issue downloading (seemingly at runtime given how much they're downloading each day) from a random bucket you don't control.
- JustFinishedBSG 7y ago"It's your own fault if you didn't foresee a trillion dollar company exploiting the product you made free and open source to help researcher and now are incurring 15K$ bills per month" And then people wonder why people don't want to make their stuff free/open source. Even when it's free people still think you're somehow entitled and overcharging.
- kelnos 7y agoI mean... yes? It doesn't have to be a trillion dollar company. You put tons of useful data on a public S3 bucket and publicize it, and people are going to download it. S3 data transfer isn't free, so I think it's reasonable to expect that, over time, the cost of serving the data is going to be prohibitive without funding it in some way. > Even when it's free people still think you're somehow entitled and overcharging. I just explicitly advocated for the opposite of that, so I'm not sure where you're getting that.
- paxys 7y agoThat's about $4000/month in bandwidth costs, assuming retail pricing. FYI he is bragging, not complaining. There are a dozen ways to reduce or eliminate this problem.
- m-p-3 7y agoCould be a lucrative way of eliminating that problem.
- qes 7y ago> That's about $4000/month in bandwidth costs You're an order of magnitude off. 45 TB per day is 1,350 TB in a month, or 1,350,000 GB. Show me somewhere you can get a petabyte of egress inside a calendar month for 4 figures USD... Let's suppose you even used the cheaper egress from Cloudfront rather than serving from S3 (lol @ your wallet if you serve 1 PB doing that). https://aws.amazon.com/blogs/aws/aws-data-transfer-prices-reduced/ https://aws.amazon.com/blogs/aws/aws-data-transfer-prices-re... The first petabyte costs an average of $0.045 per GB - $45,000. The remaining 350 TB costs another $10,500 for a total of $55,500. Serving off S3 directly? Yeah, that'll be more like $150,000.
- heyoni 7y agoHow the hell are they paying that bill then??
- pzmarzly 7y ago> Show me somewhere you can get a petabyte of egress inside a calendar month for 4 figures USD Correct me if I'm wrong, but most colocation/dedicated server providers offer such prices. E.g. hetzner.com @ €1/TB, sprintdatacenter.pl @ €0.91/TB, dedicated.com @ $2/TB (or $600/month for an unmetered 1 Gbps connection). But if you want S3/CDNs/<insert any cloud offering here>, then yeah, they're expensive. BTW per Cloudflare ToS[0]: > Use of the Service for the storage or caching of video (unless purchased separately as a Paid Service) or a disproportionate percentage of pictures, audio files, or other non-HTML content, is prohibited. [0] https://www.cloudflare.com/terms/ https://www.cloudflare.com/terms/
- CobrastanJorji 7y agoIf you host large, publicly available data in a cloud blob service, but you don't have a budget for it, one option is to use the "Requester Pays" feature that Amazon and Google provide. This makes the data available to anyone to download, but they need to pay the download cost themselves. This is at the tradeoff of making your data significantly more irritating to access, as it's no longer just plugging in a URL into a program, plus everyone who wants your dataset needs to set up a billing account with Amazon or Google.
- baroffoos 7y agoOr just post a magnet link.
- CobrastanJorji 7y agoSure, that's a great option for helping to reduce the cost for well-meaning general use, but the other way makes your costs 100% predictable, which is great if you're on an academic budget (but, again, way more annoying for the downloaders unless they're also using AWS).
- delfinom 7y agoSo if there's no other seeders, you end up eating the full cost as the only seeder....
- fs111 7y agotorrents were supported by S3 in the past https://docs.aws.amazon.com/AmazonS3/latest/dev/S3Torrent.html https://docs.aws.amazon.com/AmazonS3/latest/dev/S3Torrent.ht...
- rtkwe 7y agoThat probably wouldn't work here. That s3 bucket is hosting models downloaded at runtime/startup [1] and looks like under normal runs it would be cached. If this is being used in Apple's CI pipeline though the whole thing is being torn down between builds so every build and test has to fetch it again. [1] https://github.com/huggingface/pytorch-transformers/search?q=s3&unscoped_q=s3 https://github.com/huggingface/pytorch-transformers/search?q...
- rhacker 7y agoI'm guessing someone at apple internally distributed a dockerfile that pulls that down.
- jijji 7y agoHosting terabytes of data on an S3 bucket where people would download 45TB per month ($0.023/GB == $1000+/month) sounds like a really expensive way to distribute your data to people...
- xfitm3 7y agoYou could configure buyer pays, if you wanted to.
- rtkwe 7y agoThat makes it harder for everyone though where companies like Apple should proxy and/or cache those requests to their own internal version rather than hitting that S3 bucket every time. Requiring requester payment would mean it would only really be used by corporations where the author clearly wants a service open to anyone without having to open an AWS account to pay.
- thenickdude 7y agoMake the bucket buyer-pays, but offer Torrent links as well. Businesses doing CI will pay for S3 usage so they don't have to deal with torrents, end-users will get free torrent access, everybody wins.
- Hitton 7y agoI just skimmed AWS requester pays and it seems that you can't set the price, the requester always pays only the AWS's download cost, so why bother setting it as requester pays at all if it doesn't bring you anything?
- corint 7y agoIt brings you the absence of GET requests and the bandwidth charges on your AWS bill. In this case, 45TB a day is going to add a lot to an AWS bill, you can shift that cost to the user that's downloading the files from you.
- deleted 7y ago[deleted]
- dlasek 7y agoThey're the ones that made Amazon get those Data Trucks lol
- idlewords 7y agoThis is what success looks like if you charge money for a good or service.
- master_yoda_1 7y agoSo these jokers at apple publish a paper by using code from huggingface.
- yalogin 7y agoIsn't it likely that someone wrote a script for testing some regression and it keeps running in a loop? I can almost bet that will be the case.
- nurettin 7y agoThis is probably apple's continuous integration tests, lazily written to download the whole thing every time someone merges a commit.
- mister_hn 7y agothat's really stupid. I mean, I would have set a cache repository (SonaType Nexus maybe?), download everything there and use that repository. In the tweets the author says they've blocked the download from Apple IPs, so now their pipeline is broken.
- Manfred 7y agoUnfortunately it's very common for companies to set up their CI without any form of caching. I think it's mostly because developers are under time pressure from their managers. In some cases it's because CI is set up by juniors who don't fully understand the tools and the consequences of setting them up at this scale.
- coldtea 7y agoIt's also not a problem until it is. Then, you can devote resources at it, but meanwhile you got things up and running for months/years faster than if you tried to get everything setup just so from the get go...
- coldtea 7y agoSo they can fix it, now, that is is a problem, and they didn't have to spend time worrying about that before. Sounds very smart on their party: move forward with what matters (building your codebase, tests, etc) and don't do something (like an internal cache), unless you have to -- the ops equivalent of lazy-loading...
- falsedan 7y agoEvery medium-sized org I’ve worked at has put a caching layer in front of their build dependencies, so builds aren’t blocked when GitHub/PyPI are unavailable. No build/release engineer would leave that trivial door open if they were responsible for the build Sounds like the build pipeline was set up by a regular dev.
- dymk 7y agoNeed to distribute large static content? Looks like a good job for a torrent.
- yellowsir 7y agoor ipfs using ipns ... so many good solutions, but people suggest bucket where the downloader is paying... or cloudfear /sic
- codesternews 7y agoLooks like open source company. What's their business model? Does any one know, How they earn money?
- cpach 7y agoIsn’t this a use case where BitTorrent would shine?
- michaelt 7y agoThe problem here is "Someone's CI pipelines redownload the same models on every build" I'd say there's only a 10% chance Apple's firewall would let BitTorrent through, and only a 3% chance the CI servers would maintain a positive seed ratio. Possibly it might solve the problem because users would cache the resources themselves to avoid the hassle of getting BitTorrent into their CI pipeline...
- delfinom 7y agoNot magically? Torrents need seeders. If you are the only seeder, then you will still get the full bill from AWS all the same.
- StreamBright 7y agoPaid by requester is the feature they are looking for. https://docs.aws.amazon.com/AmazonS3/latest/dev/configure-requester-pays-console.html https://docs.aws.amazon.com/AmazonS3/latest/dev/configure-re...
- martin-adams 7y agoThat's a very neat feature I never knew existed. I only suspect this will hurt their larger mission at helping many smaller teams and individuals to use the models.
- StreamBright 7y agoNot really, if you are a small user your cost is negligible. You can calculate how much it would be for a small team but my guess is couple of dollars per month. Apple's use case is still very reasonable and the cost for them also not as bad. They could also split out different customers to different buckets and have big guys pay for it while smaller companies have it for free. There are many options.
- martin-adams 7y agoYeah that makes sense. I was thinking in terms of ease of access. If a large organisation makes everyone pay, that means everyone has to arrange accounts and payment methods. By splitting it out you have no guarantee that the big guys will just use the free one.
- z3t4 7y agoApple are probably doing "continuous integration" where all assets are re-downloaded from the Internet in each iteration. Tip: put your stuff on Github :P
- coleca 7y agoWith models that large you would be paying for GitHub's LFS credits. Those aren't cheap as I recall. Napkin math at $5/50GB of bandwidth per month for 45TB per day it would cost them $135k/mo to use Github. That's over 2x more than the S3 egress charges would be.
- peterwwillis 7y agoIf your CI/CD is re-downloading and re-building everything on every single run, you are not only being wasteful, you're actually more likely to have an outage due to not storing dependency artifacts needed for deploy. Use a local artifact store to be more resilient to failures of servers you don't control (and also save everyone money and time).
- tnolet 7y agoIs this what they call product market fit?
- half-kh-hacker 7y agoIt's surprising that nobody here's mentioned Wasabi, since they have free egress.
- koolba 7y agoThere's no such thing as a free lunch. Wasabi's own docs[1] mention a ballpark definition of "reasonable" as monthly transfer being less than total storage. Above that you'll get a call. I'm sure they'd still be cheaper than AWS, but it's not going to be $5/TB/month to service the entire world. [1]: https://wasabi.com/pricing/pricing-faqs/ https://wasabi.com/pricing/pricing-faqs/
- ChuckMcM 7y agoAnd now the twitter post is gone? I'm guessing the west coast woke up and someone at Apple said "Wait, you could infer some proprietary information with that information ..."
- deleted 7y ago[deleted]