4 ms·
I've been looking for the answer to this question for a while w/r/t large datasets with thousands or millions of files. I'm not concerned as much with programma
by web007 6y ago
I've been looking for the answer to this question for a while w/r/t large datasets with thousands or millions of files. I'm not concerned as much with programmatic usability as saving my FS from being overwhelmed both from indexing and allocating so many distinct entities. Doing any bookkeeping on a directory of a million files - even hierarchically organized - tends to be very taxing, versus storing the same amount of data in a single binary format is usually simple to manage.
I'm surprised that ZIP wasn't (isn't?) a contender. Tooling exists everywhere, it lets you mix and match data types, and it seems to hit nearly every point in their comparison. The only point I'm not sure about is "Incremental reads/writes", as it keeps a central directory structure at the end of the file. Incremental reads would need to seek first then could read randomly, and writes are slightly more complicated having to rewrite the entire directory structure to append.
- hikarudo 6y agoI've been looking for a solution to the same problem. So far we've been using a single server for storage, and developers rsync whatever they need locally. A few million images. Training is usually done locally, but when we do use cloud training, we upload just the dataset we need to S3 and use EC2. We're a small team, and currently considering moving to a cloud-first infrastructure. The idea is to store each image in S3, and all metadata (annotations etc) in Postgres or something like that, maybe using Postgres's JSON/JSONB feature. I'd appreciate any thoughts and pointers on handling datasets with a few million images.
- davnn 6y agoI would not see Zip as a file format, it's just compressed version of some other format. You therefore inherit all the pros and benefits of the compressed file, which might be columnar storage (e.g. Parquet) or row-based (e.g. Avro), but those "modern" formats have compression built in, no need for zipping.
- 7952 6y agoWhy not SQLite? You get an efficient file format, and can easily add metadata to each file. That is how the Geopackage format works for storing tile caches for mapping data. There are often millions of small files that don't work well on a file system. I think there is even a standard for storing directory structures on SQLite.
- web007 6y agoSQLite is one of the options evaluated in the article. I haven't considered it because I feel that it's a proprietary format, and also because I haven't read up on its improvements in the past ~decade. My view is still that SQLite corrupts easily because people were treating it like a proper database and doing multithreaded reading/writing, but in reality it's a different thing that doesn't quite work that way.