3 ms·
I run a service which writes around 50-80 million small documents into S3 per month and, like you, we spend more on PUTs than storage and retrieval. So reducing
by aclelland 8y ago
I run a service which writes around 50-80 million small documents into S3 per month and, like you, we spend more on PUTs than storage and retrieval. So reducing PUT costs is something we've looked into a few times.
We do make heavy use of the S3 versioning API to let us roll back files, lifecycle rules for past revisions and have never been able to find any service which offers the same amount of functionality with reasonable pricing.
One thing we have done which has saved us quite a lot is adding a buffer between our service and S3 which delays file writing by up to an hour in case an updated version of a file is uploaded. This lets us limit our PUT requests but still gives us hourly revisions of files if we need to roll back. Without this system we'd be looking at over a billion writes a month. Certainly not a strategy suitable for everyone though.
- brianwski 8y agoDisclaimer: I work at Backblaze. > We do make heavy use of the S3 versioning API to let us roll back files, lifecycle rules for past revisions and have never been able to find any service which offers the same amount of functionality with reasonable pricing. Backblaze B2 has both lifecycle rules ( https://www.backblaze.com/b2/docs/lifecycle_rules.html https://www.backblaze.com/b2/docs/lifecycle_rules.html ) and file versioning ( https://www.backblaze.com/b2/docs/file_versions.html https://www.backblaze.com/b2/docs/file_versions.html ). I would LOVE for you to evaluate it and tell us what we are missing to win your business! The Backblaze B2 system is still new and we are actively expanding the APIs all the time, so we're always collecting customer requirements looking for common issues we can tackle next.
- user5994461 8y agoDepends on the size of the file. If you're below 4 KB, the main competing system is databases, probably cassandra or DynamoDB for scaling/availability. You should also look into redis/memcache to provide the caching/buffering. Beyond that, S3 is pretty good and it surely makes sense to model the content as files rather than blob.