11 ms·
Do you absolutely have to write the data to files directly? If not, then using a time series database might be the better option. Most of them are pretty much d
by this_user 2y ago
Do you absolutely have to write the data to files directly? If not, then using a time series database might be the better option. Most of them are pretty much designed for workloads with large numbers of append operations. You could always export to individual files later on if you need it.
Another option if you have enough local storage would be to use something like JuiceFS that creates a virtual file system where the files are initially written to the local cache before JuiceFS writes the data to your S3 provider as larger chunks.
SeaweedFS can do something similar if you configure it the right way. But both options require that you have enough storage outside of your object storage.
- Gasp0de 2y agoWe tried some readymade options but they were way more expensive than our custom built S3 solution (by a factor of x10 approximately). I think we tried timescale and AWS Timestream. I haven't heard of SeaweedFS.
- ramses0 2y agohttps://github.com/seaweedfs/seaweedfs?tab=readme-ov-file#quick-start-seaweedfs-s3-on-aws https://github.com/seaweedfs/seaweedfs?tab=readme-ov-file#qu... https://github.com/seaweedfs/seaweedfs/wiki/Cloud-Drive-Benefits https://github.com/seaweedfs/seaweedfs/wiki/Cloud-Drive-Bene... https://github.com/seaweedfs/seaweedfs/wiki/Cloud-Tier https://github.com/seaweedfs/seaweedfs/wiki/Cloud-Tier https://github.com/seaweedfs/seaweedfs/wiki/Benchmarks https://github.com/seaweedfs/seaweedfs/wiki/Benchmarks https://github.com/seaweedfs/seaweedfs/wiki/Words-from-SeaweedFS-Users https://github.com/seaweedfs/seaweedfs/wiki/Words-from-Seawe... https://github.com/seaweedfs/seaweedfs/wiki/Amazon-S3-API https://github.com/seaweedfs/seaweedfs/wiki/Amazon-S3-API ...your true issue is it seems like you're using the filesystem as the "only" storage layer in play, but you also need time and entity querying(!?!). >> we need to be able to query on a per sensor basis and a timespan ...look at the "Cloud-Tier" wiki page. If you're truly in an "everything's hot all the time" situation, you really should be using a database. If you're pulling "usually recent stuff, occasionally old stuff" then fronting with something like SeaweedFS seems like it might "just" transparently reduce your overall costs. Really, I'd nudge towards "write .txt ; compact ... ; SELECT ... && cat .txt". Basically, keep your inbound writes cached to (eg) seaweed as unit files. "Compact them" every hour by appending rows to some appropriate database (I mean: migrate to using litefs, turso, postgres, something like that). When you read, you may need to supplement "tip" data from your incoming files, but the majority should be hitting a "real" remote database, there's plenty to choose from! A nifty note, sqlite can connect to multiple DB's at once: https://www.sqlite.org/lang_attach.html https://www.sqlite.org/lang_attach.html ... https://stackoverflow.com/posts/10020/revisions https://stackoverflow.com/posts/10020/revisions ...something like `select * from raw union (select * from one_hour) union (select * from today) union (select * from historical) ...`