8 ms·
Just finished installing it on my OpenIndiana NAS to replace Minio. Biggest difference so far is that Minio is just files on disk, Garage chunks all files and
by seized 4y ago
Just finished installing it on my OpenIndiana NAS to replace Minio.
Biggest difference so far is that Minio is just files on disk, Garage chunks all files and has a metadata db.
Minios listing operations were horribly slow, still have to see if Garage resolves that.
- KronisLV 4y ago> Biggest difference so far is that Minio is just files on disk, Garage chunks all files and has a metadata db. I'd kind of expect most blob storage solutions to use abstractions other than just the file system, or at least consider doing so. I recently built a system to handle millions of documents as a proof of concept and when I was testing it with 10 million files, the server ran out of inodes, before I went over to storing the blobs in some attached storage that had XFS: https://blog.kronis.dev/tutorials/3-4-pidgeot-a-system-for-millions-of-documents-deploying-and-scaling https://blog.kronis.dev/tutorials/3-4-pidgeot-a-system-for-m... With abstracted storage (say, files bunches up into X MB large containers or chunked into such when too large, with something else to keep track of what is where) that wouldn't be such an issue, though you might end up with other issues along the way. It's curious that we don't advocate for storing blobs in relational databases anymore, even though I can also understand the reasoning (or at least why having a separate DB for your system data and your blob data would be a good idea, for backups/test data/deciding where to host what and so on).
- alex_sf 4y ago> I'd kind of expect most blob storage solutions to use abstractions other than just the file system, or at least consider doing so. Honestly, I'd expect the exact opposite. Filesystems are really good at storing files. Why not leverage all that work? > I recently built a system to handle millions of documents as a proof of concept and when I was testing it with 10 million files, the server ran out of inodes, before I went over to storing the blobs in some attached storage that had XFS That's a misconfiguration issue though, not a reason to not store blobs as files on disk. Ext4 can handle 2^32 files. ZFS can handle 2^128(?). > With abstracted storage (say, files bunches up into X MB large containers or chunked into such when too large, with something else to keep track of what is where) that wouldn't be such an issue, though you might end up with other issues along the way. A few issues that come to mind for me: * This requires tuning to actually reduce the number of inodes of used for certain datasets. E.g., if I'm storing large media files, that chunking would _increase_ the number of files on disk, not reduce it. At which point, if inode limits are the issue, we're just making it worse. * It adds additional complexity. Now you need to account for these chunks, and, if you care about the data, check it periodically. * You need specific tooling to work with it. Files on a filesystem are.. files on a filesystem. Easy to backup, easy to view. Arbitrary chunking and such requires tooling to perform operations on it. Tooling that may break, or have the wrong versions, or.. etc. > It's curious that we don't advocate for storing blobs in relational databases anymore, even though I can also understand the reasoning In my experience, the popular RDBMS out there just aren't good at it. With the way locking semantics and their transaction queueing works, storing and retrieving lots of blobs just isn't performant. You can get away with it for a long time though, and it can be pretty nice when you can.
- mdaniel 4y ago> Filesystems are really good at storing files. Why not leverage all that work? As an asterisk, the S3 API is key-value pairs, not files; that distinction comes up a lot when interacting with Amazon S3, and I would expect the same with an S3 API clone. For example, ListObjects[1] has a "delimiter" that (AFAIK) defaults to / making it appear to be a filesystem but using "." or "!" would be a perfectly fine delimiter and thus would have no obvious filesystem mapping 1: https://docs.aws.amazon.com/AmazonS3/latest/API/API_ListObjects.html#API_ListObjects_RequestSyntax https://docs.aws.amazon.com/AmazonS3/latest/API/API_ListObje...
- vbezhenar 4y agoWhy is it useful?
- mdaniel 4y agoThat's a complicated question but allows highlighting what I was bringing up: the Key is any unicode character[1] so while it has become conventional to use "/", imagine if you wanted to store the output of exploded jar files in S3, but be able to "list the directory" of a jar's contents: `PutObject("/some-path/my.jar!/META-INF/MANIFEST.MF", "Manifest-Version: 1.0")` Now you can `ListObjects(Prefix="/some-path/my.jar", Delimiter="!")` to get the "interior files" back. I'm sure there are others, that's just one that I could think of off the top of my head. Mapping a URL and its interior resources would be another (`"https://example.com\t/script[1] https://example.com\t/script[1]", "console.log('hello, world')")` Further fun fact that even I didn't know until searching for other examples: "delimiter" is a string and thus can be `Delimiter=unknown` or such: https://github.com/aws/aws-sdk-go/issues/2130 https://github.com/aws/aws-sdk-go/issues/2130 1: see the ListObject page under "encoding-type"
- selfhoster11 4y agoMinio supports virtual ZIP directories for such use cases. In your example, as long as this was enabled and your jar file was properly detected, you could submit a GET for "/some-path/my.jar/META-INF/MANIFEST.MF" and get the contents of that file just fine.
- vbezhenar 4y ago> It's curious that we don't advocate for storing blobs in relational databases anymore That's exactly what I did recently on new work: migrated blobs from DB to S3. It significantly reduced load from the servers (and will reduce more, right now the implementation is primitive - just proxying S3, using URL will allow other services to deal with S3 directly). It solved backup nightmare (those people couldn't do backup because their server run out of space every month). I'll admit that backup issue is more like admin incompetence but I work with what I get. Having database shrink from 200GB to 80MB now allows to backup/restore it in seconds rather than hours. I didn't find any issues with S3 approach. Even transactions solved by a tiny possibility of leaving junk in S3 which is a non-issue. Just upload all data to S3 before commit and delete if commit fails (and if commit fails and delete fails, so be it).
- kilburn 4y ago> Biggest difference so far is that Minio is just files on disk Minio _was_ just files on disk. They don't support that mode anymore since 2022-10-29 (see the big yellow warning box at [1]). [1] https://min.io/docs/minio/linux/operations/install-deploy-manage/deploy-minio-single-node-single-drive.html https://min.io/docs/minio/linux/operations/install-deploy-ma...
- seized 4y agoAh interesting. I found it appealing to always have a way to get at the data natively as a worst case for restores. The whole use case is for Vertical Backup (from the maker of Duplicacy) to back up VMs.
- aftbit 4y agoYeah, I have actually frozen Minio at this older version in my stack, as "just files on disk" was the primary feature that drew me to it. I don't want my data locked into some custom format. I'd be willing to bet that ZFS will still be supported in 20 years, but I would not make the same bet about Minio. For the same reason, it looks like Garage is not an option for my use case.
- soulmachine 4y agoI use this "just files on disk" feature too, I also use ZFS. I have a bunch of crawlers uploading data to AWS S3, but S3 is too expensive, I replaced S3 with MinIO. MinIO stores plain files on disk, which makes it a lot easier to access my data. I can read files without calling MinIO APIs, the speed is super fast. By the way, which old version are you using? I'm using RELEASE.2022-04-26T01-20-24Z
- aftbit 4y agoWell... right now I am using minio version RELEASE.2021-08-31T05-46-54Z. However I really ought to upgrade that to the last supported version.
- semi-extrinsic 4y agoI thought the "stat" command of Minio was supposed to resolve the "listing is horribly slow" issue?
- seized 4y agoMaybe, but that didn't help third party tools that I could see.