2 ms·
I've worked with a S3 + SQL system. It was used for serving data for a reporting dashboard where the stored data was in the 0.1-10 TB range. As the use case was
by bcbrown 8y ago
I've worked with a S3 + SQL system. It was used for serving data for a reporting dashboard where the stored data was in the 0.1-10 TB range. As the use case was only semi-interactive (users didn't mind waiting 1-10 seconds for a report), and all the queries were pre-defined, this solution was a good fit.
I think it makes sense when there's no in-place updates; either querying write-once data like logs or the output of batch data processing roll-ups that replace the previous data. The less you need the relational model (like joins), the better, but some of those needs can be met through careful design of the storage schema and denormalization.
I wouldn't advocate this sort of solution if your requirements include in-place updates of existing data, frequent/granular updates of new data, expressive ad-hoc queries that use the full capability of relational algebra, or tight latency requirements. You also lose the safety net of referential integrity and table-level constraints, as those are now enforced in custom code that can have bugs.
I would say maintaining this system cost about a half-engineer for ongoing maintenance and new functionality.