3 ms·
I use this technique extensively in several production systems. As others have mentioned having an expiration policy is a good idea. Also you can mitigate char
by madmod 10y ago
I use this technique extensively in several production systems.
As others have mentioned having an expiration policy is a good idea. Also you can mitigate charges from malicious activity by using rate limits on the signing endpoint. (API gateway supports this.) Using infrequent access or reduced redundancy storage might also be a good idea if you expect a lot of traffic. It's also good to limit the CORS policy on the bucket to the needed domains and headers.
Signed metadata headers are very useful when combined with S3 event handlers (SQS or straight to Lambda.) using a HEAD request on the uploaded objects. This is a great technique for post processing an upload without requiring client trust or an external data store. (With a separate falliable request which could lead to consistency issues.)
Edit: It is also critically important to have some randomness in each key path so it is ungessable. Otherwise user files would be overwriteable by an attacker. (Many file names are easily guessable and an attacker with many tries could eventually stuff malware in for example.) I used guids for this because they are both URL and S3 key safe. If keeping the original file name is needed I put it in a metadata value and rename the file on download using a Content-Disposition header. Making the S3 headers work with symbols in file names can be tricky but encoding it as a JSON string works around most issues.
In order to overcome the 30 second request limit in API gateway for longer post processing while still offering realtime client feedback you can set up an S3 event handler to trigger the post processing lambda which then updates a DynamoDB record with the S3 key as it's id. A status endpoint lambda is then polled by the client with the S3 key for status events.
For more complex post processing and client side workflows I have used key prefixes (folders) each with seaprate event handlers or CORS configurations. IAM polices with conditions including S3 key prefixes are used to restrict access. Using the S3 API copy command can move large objects quickly between workflow steps.
Also enabling server side encryption is a must imo. Be sure to specify AWS signature version 4 in the S3 constructor so that all parts of the request are signed. (Otherwise some older regions may not sign metadata headers.)
Also the S3 API copy command has an interesting append feature which can be used to build objects iteratively. I once toyed with the idea of using it to create large zip files of many S3 objects efficiently but ended up not needing it. Someday I would like to try that because it could be great for a lot of web apps where users can select a random list of files to download.
Also I (re)implemented most of the above this week using CloudFormation and the newer AWS Serverless Template (not the serverless.com project but the actual AWS feature.) which allows for really easy deployment.