4 ms·
I think the author made a pretty good decision to choose a solution which was proven to meet the requirements by another author. However Amazon does provide a t
by illumen 11y ago
I think the author made a pretty good decision to choose a solution which was proven to meet the requirements by another author. However Amazon does provide a tool to batch download multiple things from S3 though. Let's pretend for a moment those tools don't exist already... Let's also pretend Amazon doesn't suggest another solution for when there are lots of GET requests [0]. Let's pretend that using something like nginx which is a very well tuned web proxy might be way better because of SPDY, SSL connection pooling and such [1]. Let's also pretend that node.js can not use multiple processes with cluster cluster [2].
Why wouldn't one use python for this job? Since the company already does a lot of python, and this job is actually easy to do in python (I've done something similar).
400 processes with python3.5 is way less than the 4gig on one medium instance (less than 2.5MB each). Just farm the work out to a ProcessPoolExecutor, and have a timeout on the S3 get requests. That would let you have enough resources to match the spike of 20000 requests per minute (334 per second).
A lot of an S3 GET request could be the SSL, AWS auth and such. All quite CPU intensive. So using an async framework that doesn't do SSL+AWS auth async is obviously not going to work well once the requests go up.
There's even an example in the concurrent.futures of downloading urls [3].
Made with the beautiful python3.5
import concurrent.futures
import requests
from awsauth import S3Auth
ACCESS_KEY = 'ACCESSKEYXXXXXXXXXXXX'
SECRET_KEY = 'AWSSECRETKEYXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX'
URLS = ['http://www.foxnews.com/', 'http://www.cnn.com/', 'http://europe.wsj.com/',
'http://www.bbc.co.uk/', 'http://some-made-up-domain.com/']
def load_url(url, timeout):
return requests.get(url, timeout=timeout).data
# return requests.get(url, timeout=timeout, auth=S3Auth(ACCESS_KEY, SECRET_KEY)).data
with concurrent.futures.ProcessPoolExecutor(max_workers=400) as executor:
# Start the load operations and mark each future with its URL
future_to_url = {executor.submit(load_url, url, 60): url for url in URLS}
for future in concurrent.futures.as_completed(future_to_url):
url = future_to_url[future]
try:
data = future.result()
except Exception as exc:
print('%r generated an exception: %s' % (url, exc))
else:
print('%r page is %d bytes' % (url, len(data)))
[0] http://docs.aws.amazon.com/AmazonS3/latest/dev/request-rate-perf-considerations.html#get-workload-considerations http://docs.aws.amazon.com/AmazonS3/latest/dev/request-rate-...
[1] https://coderwall.com/p/rlguog/nginx-as-proxy-for-amazon-s3-public-private-files https://coderwall.com/p/rlguog/nginx-as-proxy-for-amazon-s3-...
[2] https://nodejs.org/api/cluster.html#cluster_cluster https://nodejs.org/api/cluster.html#cluster_cluster
[3] https://docs.python.org/dev/library/concurrent.futures.html https://docs.python.org/dev/library/concurrent.futures.html
- mh- 11y ago>However Amazon does provide a tool to batch download multiple things from S3 though. out of curiosity, what tool are you referring to here?
- trimbo 11y agoAnother idea is to use lua in Nginx itself, and keep it all within nginx. Here's an upload example: https://github.com/jamesmarlowe/lua-resty-s3 https://github.com/jamesmarlowe/lua-resty-s3