7 ms·
Downloading files from S3 with multithreading and Boto3
- dddw 5y agoNice read
- incomplete 5y agoagree, i found it quite helpful!
- tiltrus 5y agoAgreed!
- catillac 5y agoExcellent walkthrough, love boto. We’ve recently been using s5cmd which we’ve found is ridiculously faster than boto without any extra boto tricks. https://github.com/peak/s5cmd https://github.com/peak/s5cmd
- thefourthchime 5y agoSame!
- __turbobrew__ 5y agoAre parallel downloads faster in S3? My intuition says you will hit network or disk bottlenecks before being able to saturate a single core with work?
- aloer 5y agoExtreme case can be when you make heavy use of s3 for ec2 workloads and want to saturate your instance connection. This can go up to 40g afaik Edit: I read this recently and if I remember correctly there’s a limit of like a thousand parallel connections to s3 AWS internally probably has higher limits for some of their services, e.g. when you query data in s3 with Athena
- __turbobrew__ 5y agoMake sense, thank you. Wouldn't you want to multiprocess the downloads instead of multithread? I imagine you would run into the Python global interpreter lock before being able to push through 40gb?
- ak217 5y agoYes, you would. The CPU and memory overhead of multiprocessing for this application is why we ended up migrating away from boto3 and to the AWS Go SDK for this specific purpose (https://github.com/chanzuckerberg/s3parcp https://github.com/chanzuckerberg/s3parcp as I mentioned in another comment). We still use boto3 in other areas, but for maxing out the network connection, golang is far more scalable.
- Galanwe 5y agoUsing multiprocessing I've been able to quite easily saturate a 20GBps Ec2 NICs in python. https://github.com/NewbiZ/s3pd https://github.com/NewbiZ/s3pd There is no reason why multiprocessing for IO in python would use _crazily_ more memory than in an other language, when done properly.
- banana_giraffe 5y agoYes, parallel downloads are faster in some cases. There are even extreme cases, mostly around giant EC2 nodes where you can take this idea a step further and spin up multiple processes to each download parts of a file, and really saturate your network or disk. My favorite version of this is when you start to use shared memory of some fashion to move terabytes of data from S3 to EC2 to work on it without ever hitting a disk. Not for everyone, and for sure many times the extra milliseconds saved won't matter, but sometimes you really do need to get hundreds of gigabytes or terabytes of data moved as quickly as possible.
- Galanwe 5y agoThat's what my small python package does (https://github.com/NewbiZ/s3pd https://github.com/NewbiZ/s3pd). Definitely not perfect (not saturating cores because no use of event loop per process) but download is split in multiple processes and stored in shared memory. I've been able to saturate 20GB NICs on Ec2 with it (32 cores)
- killingtime74 5y agoYou just described Apache Spark
- twotwotwo 5y agoWhen you have a lot of small files, the latency on each operation, though not huge in absolute terms, can end up taking up more time than the data transfer and whatever other work is being done on the content. (At work, we had an upload job with ~800k files, ranging from <1kb to >100kb. I looked at rearranging how we stored things to avoid small files, but it ended up a straighter shot to continue to use little files but use a worker pool to make the transfer parallel.)
- twotwotwo 5y ago(Er, some files >100MB not >100KB. If they maxed out around 100KB we probably wouldn't have picked S3 to store them!)
- throwaway340953 5y agoDepends where you're downloading from. S3 single stream GET throughput is throttled to ~40MB/sec. An HDD can write at ~200MB/sec and SSDs can write ~500MB/sec. NICs on old EC2 instances can do 10Gbps and the new ones can do 25Gbps. If you're on a home internet/mobile connection downloading large files, a single download will likely saturate your connection. If you're on an EC2 instance, you should be able to do 10-100x better parallelizing. Source: I used to work on S3.
- ak217 5y agoAside from what the other replies mentioned, one of the network bottlenecks that you're referring to is the per-TCP connection packet rate limit and congestion avoidance artifacts that many networks impose. The S3 frontend node that serves your request may also experience congestion from "hot" objects it has been assigned, so amortizing your connections over many S3 nodes can help. Finally, for many small objects, having multiple connections helps amortize the per-object latency overhead. As far as the disk overhead, modern NVMe SSD drives can easily sustain millions of IOPS and multiple gigabytes per second of bandwidth, more than keeping up with a 40 gbps link (such as on a large EC2 instance that does have the connectivity to talk to S3 at that rate).
- emasquil 5y agoIn this case I needed to download about 1k files of size ~ 10mb. Doing this with multithreading was so much faster than downloading all the files sequentially.
- mathgladiator 5y agoyes, and if you leverage the multipart upload then downloading along those boundaries will have zero contention and you can go extra fast.
- bassdropvroom 5y agoThere is aioboto3 which wraps boto3 in asyncio calls. It's a lot simpler than all of this.
- ak217 5y agoHave you benchmarked this solution on a high bandwidth connection? If so, what's the max throughput that it achieved?
- emasquil 5y agoUnfortunately I haven't. If someone wants to benchmark this solution I'd be glad to hear about that!
- emasquil 5y agoIn this case I just wanted to keep it simple and do it the "native" way. I've been hearing a lot about aioboto3 and I'll surely check it out. However, it seems to have some limitations like https://github.com/boto/botocore/issues/458 https://github.com/boto/botocore/issues/458
- ben509 5y agoThat issue is just botocore not supporting async, which aiobotocore addresses. It's also very long; is there a particular comment on that issue that you wanted to point out?
- emasquil 5y agoI just mentioned the issue because a colleague of mine did some quick tests with aioboto3 and he told me it was downloading the files sequentially. We thought it might be related to that. You made a great point about going async if you need a working command channel.
- ben509 5y agoHaving used aiobotocore, it's pretty good. It can be a bit sketchy about releasing connection pools, though. Normally you shouldn't run into any issues, but we've had trouble with running out of filehandles on AWS Lambda. I would also point out that for many cases, using a ThreadPoolExecutor works fine. I maintain a simulation engine and it's distributing work out to AWS Lambda. So I had two distributors, one thread based and one async based. (Also one based on multiprocessing that runs work on a local machine.) From my tests, the performance is basically equivalent, which is not surprising: most of it was just waiting and then processing incoming responses. Threads work great at this, and botocore is designed to work with threads. I eventually went with async because Python doesn't let you prioritize threads. That means that if you get a "stop" signal, in the threaded model, the command thread is competing with many worker threads. In the async model, the workers are all in one thread, so the command thread will be woken up per the switch interval[1]. So, broadly, if you're going to need a command channel that must respond in a timely fashion, I'd recommend piling workers into an event loop through async. The other possibility is to have a command channel run in another process, but then you need to get the fork right, do IPC, etc. [1]: https://docs.python.org/3/library/sys.html#sys.setswitchinterval https://docs.python.org/3/library/sys.html#sys.setswitchinte...
- ak217 5y agoThe AWS Go SDK now has a connection pool based S3 download/upload manager API that allows saturating your (e.g. 40Gbit/s EC2-S3) network connection using far less memory and CPU than is possible with Python. A colleague of mine developed this tool to make this functionality available in a CLI: https://github.com/chanzuckerberg/s3parcp https://github.com/chanzuckerberg/s3parcp
- Xunjin 5y agoDo you know if there is a "sync" function just like the aws-cli?! I've been thinking of starting using Go to deploy some stuff that doesn't need python as dependency, and statically compiled :P Edit: In this case for DO Spaces. Way more cheap.
- ak217 5y agoI'm not aware of a supported API in the AWS Go SDK for this, but there is a sketch here: https://github.com/aws/aws-sdk-go/tree/main/example/service/s3/sync https://github.com/aws/aws-sdk-go/tree/main/example/service/...
- darthShadow 5y agoRClone does all that you require and much more. RClone: https://rclone.org/ https://rclone.org/ S3 Backend: https://rclone.org/s3/ https://rclone.org/s3/
- Galanwe 5y agoBe cautious though, as rclone "sync" is based on file metadata (e.g. last modified), it does not recompute local etags to know which files need to be sync'ed. For instance, if you "cp -a" a directory and then apply sync, it could do nothing and return success if the copied files were last modified before the ones in S3. For our use case at work, we wanted to be _sure_ that sync always work as intended, and thus ended up recomputing etags locally and compare to the ones in S3 to know what to sync (got bitten by the issue of last modified before)
- 2ion 5y ago
- vtuulos 5y agoIt is a bit of a hidden gem but Metaflow includes a Boto-based highly parallelized, error-tolerant S3 client that Netflix uses routinely to get 10-20Gbps throughput between EC2 and S3. Technically it is independent from Metaflow, so you could use it as a stand-alone, high-performance S3 client. See docs here https://docs.metaflow.org/metaflow/data#store-and-load-objects-in-a-metaflow-flow https://docs.metaflow.org/metaflow/data#store-and-load-objec... And code here https://github.com/Netflix/metaflow/tree/master/metaflow/datatools https://github.com/Netflix/metaflow/tree/master/metaflow/dat... (I wrote it originally - AMA if curious)
- nicornk 5y agoReally interesting going through your code, thanks for sharing. Why did you opt for running the parallelization code in a separate python process (s3op)? You might want to update section ‚Caution: Overwriting data in S3‘ in the docs since S3 offers strong read after write consistency since dec 2020. https://aws.amazon.com/blogs/aws/amazon-s3-update-strong-read-after-write-consistency/ https://aws.amazon.com/blogs/aws/amazon-s3-update-strong-rea...
- vtuulos 5y agore: separate process - fault-tolerance is a key requirement. There are a myriad of ways how a highly parallelized, network- and data-intensive code can fail, so isolating it in a separate process is a safer approach than trying to try-except everything and hope it works. Good catch re: the warning about consistency! The docs were written before the change :)
- brian_herman 5y agoCan you do a blogpost please!
- vtuulos 5y agoThe topic is close to my heart so maybe one day :)
- jcims 5y agoSorry to do this here, I know it’s not Amazon support. Is there a way to copy S3 objects from bucket to bucket without sending them through the compute that’s doing the copying? We have a use case for copying terabytes of content to buckets with different owners and it just seems wasteful to run everything through a client.
- YawningAngel 5y agoYou can probably use https://docs.aws.amazon.com/AmazonS3/latest/userguide/replication.html https://docs.aws.amazon.com/AmazonS3/latest/userguide/replic... to accomplish this
- derekdb 5y agoaws cli cp can copy between buckets.
- wikibob 5y agoThat’s true but data flows through the machine executing aws cli. The parent was asking how to copy data without it flowing through a compute instance.
- toomuchtodo 5y agoI’d double check the cli code path for that parameter. The s3 api does have a copy method, which performs the operation within s3 without compute acting as an intermediary. If that’s not the case for the copy parameter, sounds like a bug that needs to be fixed in the cli tooling.
- ben509 5y agoIt is directly checking for s3 to s3[1] and indicates that it wants to copy... I've read over it and I'm reasonably sure that it's going to issue CopyObject, but it would take me actually getting out paper and pen to really track it down. The AWS CLI and Boto are a case study in overdoing class hierarchies. Not because there's any obvious AbstractSingletonProxyFactoryBean, but rather that there's no instance that stands out as "this is where they went wrong" and nevertheless the end result is a confusing mess of inheritance and objects. [1]: https://github.com/aws/aws-cli/blob/45b0063b2d0b245b17a57fd9eebd9fcc87c4426a/awscli/customizations/s3/subcommands.py#L970 https://github.com/aws/aws-cli/blob/45b0063b2d0b245b17a57fd9...
- hiyer 5y agos5cmd[0], written in Go, performs cp in parallel to/from s3 too. I've been using it for some time and it's plenty fast. https://github.com/peak/s5cmd https://github.com/peak/s5cmd
- e12e 5y agoNo mention of how it compares to s4cmd (which like s3cmd is python). https://github.com/bloomreach/s4cmd https://github.com/bloomreach/s4cmd
- hiyer 5y agoI had tried both a couple of years back and at that time s5cmd was significantly faster when downloading multiple (100s) of files, hence my loyalty to it. Of course both are actively developed projects, so things might have changed since.
- dukeofdoom 5y agoI used Boto to upload static images from Django to s3. Ran into a few problems where the sync takes so long that I stopped relying on it. And just add new files to local and the corresponding s3 folder. Kind of annoying. This was still on Django 1.6. Not sure if this has now been fixed.
- ggregoire 5y agoI find it's easier and more convenient to use aws-cli [1] (with subprocess.run) than using Boto. Aws-cli is multithreaded in both download and upload. There is even a setting to tweak the number of parallel requests [2]. [1] https://github.com/aws/aws-cli https://github.com/aws/aws-cli [2] https://docs.aws.amazon.com/credref/latest/refdocs/setting-s3-max_concurrent_requests.html https://docs.aws.amazon.com/credref/latest/refdocs/setting-s... Edit: didn't know about s5cmd (from other commenters). Seems like a faster alternative to aws-cli.
- ramitmittal 5y agoThank you so much for posting this. I was working on something similar and this helped a lot.
- darthShadow 5y agoRClone does this and much more. RClone: https://rclone.org/ https://rclone.org/ S3 Backend: https://rclone.org/s3/ https://rclone.org/s3/
- corby 5y agoMan I had a hard time reading this post because of contrast violations. Light gray on dark gray is VERY hard to read. It was probably a good article tho
- emasquil 5y agoSorry about that! Next time I work on my blog I'll take that into account and find a better theme/color palette. In the meantime you can change it to light mode (clicking on the top right button) and there you have black text on white background.
- dvfjsdhgfv 5y agoI know people lave s5cmd (a successor of s4cmd, which itself is a successor of the popular s3cmd), but the standard aws-cli handles parallel transfers well.