5 ms·
Does anyone have a good alternative for storing large amounts of very small files that need to be individually queriable? We are dealing with a large amount of
by Gasp0de 2y ago
Does anyone have a good alternative for storing large amounts of very small files that need to be individually queriable? We are dealing with a large amount of sensor readings that we need to be able to query on a per sensor basis and a timespan, and we are dealing with the problem mentioned in the article, that storing millions of small files in S3 is expensive.
- paulsutter 2y agoIf you want to keep them in S3, consolidate into sorted parquet files. You get random access to row groups, and only the columns you need are read so it’s very efficient. DuckDB can both build and access these files efficiently. You could compact files hourly/nightly/weekly whatever Of course you could also use Aurora for a clean scalable Postgres that can survive zone failures for a simpler solution
- Gasp0de 2y agoThe problem is that the initial writing is already so expensive, I guess we'd have to write multiple sensors into the same file instead of having one file per sensor per interval. I'll look into parquet access options, if we could write 10k sensors into one file but still read a single sensor from that file that could work.
- spothedog1 2y agoNew S3 Table Buckets [1] do automatic compaction [1] https://docs.aws.amazon.com/AmazonS3/latest/userguide/s3-tables.html https://docs.aws.amazon.com/AmazonS3/latest/userguide/s3-tab...
- necrobrit 2y agoTable buckets are currently quite hard to use for a lot of use cases as they _only_ support primitive types. No nested types. Hopefully this will come at some point. Product looks very cool otherwise.
- bloomingkales 2y agoSomething like Redis instead? [sensorid-timerange] = value. Your key is [sensorid-timerange] to get the values for that sensor and that time range. No more files. You might be able to avoid per usage pricing just by hosting this on a regular vps.
- Gasp0de 2y agoWe use Redis for buffering for a certain timeperiod, and then we write data for one sensor for that period to S3. However we fill up large Redis clusters pretty fast, so we can only buffer for a shortish period.
- hendiatris 2y agoYou may be able to get close with sufficiently small row groups, but you will have to do some tests. You can do this in a few hours of work, by taking some sensor data, sorting it by the identifier and then writing it to parquet with one row group per sensor. You can do this with the ParquetWriter class in PyArrow, or something else that allows you fine grained control of how the file is written. I just checked and saw that you can have around 7 million row groups per file, so you should be fine. Then spin up duckdb and do some performance tests. I’m not sure this will work, there is some overheard with reading parquet, which is why it is discouraged to have small files and row groups.
- 0cf8612b2e1e 2y agoWhy can you not rewrite the initial file into something partitioned by sensor+time? Would the one time job really be that much more additional cost vs the additional complexity of multiple sensors per file? Do you ever go back and reaggregate older data into bigger, sorted files? That is, maybe you originally partitioned by hour, but stale data is so infrequently accessed, you could roll up into partitions per week/month/whatever. Depending on the specifics, you might save some space from less file overhead and better compression statistics.
- Gasp0de 2y agoThe costly thing is the intial writing already. S3 is our cold storage, we don't often read from it. So compaction would only make reading cheaper, but create a writing cost in the process.
- ramses0 2y agoSeaweedFS? https://news.ycombinator.com/item?id=39235593 https://news.ycombinator.com/item?id=39235593
- tobias3 2y agoI guess we need more requirements from OP, such if it should be self-hosted or a cloud service
- this_user 2y agoDo you absolutely have to write the data to files directly? If not, then using a time series database might be the better option. Most of them are pretty much designed for workloads with large numbers of append operations. You could always export to individual files later on if you need it. Another option if you have enough local storage would be to use something like JuiceFS that creates a virtual file system where the files are initially written to the local cache before JuiceFS writes the data to your S3 provider as larger chunks. SeaweedFS can do something similar if you configure it the right way. But both options require that you have enough storage outside of your object storage.
- Gasp0de 2y agoWe tried some readymade options but they were way more expensive than our custom built S3 solution (by a factor of x10 approximately). I think we tried timescale and AWS Timestream. I haven't heard of SeaweedFS.
- ramses0 2y agohttps://github.com/seaweedfs/seaweedfs?tab=readme-ov-file#quick-start-seaweedfs-s3-on-aws https://github.com/seaweedfs/seaweedfs?tab=readme-ov-file#qu... https://github.com/seaweedfs/seaweedfs/wiki/Cloud-Drive-Benefits https://github.com/seaweedfs/seaweedfs/wiki/Cloud-Drive-Bene... https://github.com/seaweedfs/seaweedfs/wiki/Cloud-Tier https://github.com/seaweedfs/seaweedfs/wiki/Cloud-Tier https://github.com/seaweedfs/seaweedfs/wiki/Benchmarks https://github.com/seaweedfs/seaweedfs/wiki/Benchmarks https://github.com/seaweedfs/seaweedfs/wiki/Words-from-SeaweedFS-Users https://github.com/seaweedfs/seaweedfs/wiki/Words-from-Seawe... https://github.com/seaweedfs/seaweedfs/wiki/Amazon-S3-API https://github.com/seaweedfs/seaweedfs/wiki/Amazon-S3-API ...your true issue is it seems like you're using the filesystem as the "only" storage layer in play, but you also need time and entity querying(!?!). >> we need to be able to query on a per sensor basis and a timespan ...look at the "Cloud-Tier" wiki page. If you're truly in an "everything's hot all the time" situation, you really should be using a database. If you're pulling "usually recent stuff, occasionally old stuff" then fronting with something like SeaweedFS seems like it might "just" transparently reduce your overall costs. Really, I'd nudge towards "write .txt ; compact ... ; SELECT ... && cat .txt". Basically, keep your inbound writes cached to (eg) seaweed as unit files. "Compact them" every hour by appending rows to some appropriate database (I mean: migrate to using litefs, turso, postgres, something like that). When you read, you may need to supplement "tip" data from your incoming files, but the majority should be hitting a "real" remote database, there's plenty to choose from! A nifty note, sqlite can connect to multiple DB's at once: https://www.sqlite.org/lang_attach.html https://www.sqlite.org/lang_attach.html ... https://stackoverflow.com/posts/10020/revisions https://stackoverflow.com/posts/10020/revisions ...something like `select * from raw union (select * from one_hour) union (select * from today) union (select * from historical) ...`
- alchemist1e9 2y agohttps://github.com/mxmlnkn/ratarmount https://github.com/mxmlnkn/ratarmount > To use all fsspec features, either install via pip install ratarmount[fsspec] or pip install ratarmount[fsspec]. It should also suffice to simply pip install fsspec if ratarmountcore is already installed.
- themgt 2y agoI've only played with it a bit but Nvidia AIStore project seems underappreciated: "lightweight, built-from-scratch storage stack tailored for AI applications" + S3 compatible https://github.com/NVIDIA/aistore https://github.com/NVIDIA/aistore
- prpl 2y agopartition by sensor first, then timestamp (or the reverse if it makes sense). If they are avro, orc, or parquet, stage and register them directly with Iceberg (see AppendFilesCommit) and compact occasionally. Newer files (uncompacted) will be small, older files will be larger and more optimized. You can also do this with a landing table or even branches+WAP.