4 ms·
We moved over to garage after running minio in production with about ~2PB after about 2 years of headache. Minio does not deal with small files very well, right
by makkesk8 2y ago
We moved over to garage after running minio in production with about ~2PB after about 2 years of headache. Minio does not deal with small files very well, rightfully so, since they don't keep a separate index of the files other than straight on disk. While ssd's can mask this issue to some extent, spinning rust, not so much. And speaking of replication, this just works... Minio's approach even with synchronous mode turned on, tends to fall behind, and again small files will pretty much break it all together.
We saw about 20-30x performance gain overall after moving to garage for our specific use case.
- sandGorgon 2y agoquick question for advice - we have been evaluating minio for a in-house deployed storage for ML data. this is financial data which we have to comply on a crap ton of regulations. so we wanted lots of compliance features - like access logs, access approvals, short lived (time bound) accesses, etc etc. how would you compare garage vs minio on that front ?
- withinboredom 2y agoYou will probably put a proxy in front of it, so do your audit logging there (nginx ingress mirror mode works pretty good for that)
- mdaniel 2y agoAs a competing theory, since both Minio and Garage are open source, if it were my stack I'd patch them to log with the granularity one wished since in my mental model the system of record will always have more information than a simple HTTP proxy in front of them Plus, in the spirit of open source, it's very likely that if one person has this need then others have this need, too, and thus the whole ecosystem grows versus everyone having one more point of failure in the HTTP traversal
- withinboredom 2y agoHmm... maybe??? If you have a central audit log, what is the probability that whatever gets implemented in all the open (and closed) source projects will be compatible?
- Too 2y agoLog scrapers are decoupled from applications. Just log to disk and let the agent of your logging stack pick it up and send to the central location.
- withinboredom 2y agoThat isn't an audit log.
- Too 2y agoWhy not? The application logs who, when and what happened to disk. This is application specific audit events and such patches should be welcome upstream. Log scraper takes care of long time storage, search and indexing. Because you want your audit logs stored in a central location eventually. This is not bound to the application and upstream shouldn’t be concerned with how one does this.
- withinboredom 2y agoThat is assuming the application is aware of “who” is doing it. I can commit to GitHub any name/email address I want, but only GitHub proxy servers know who actually sent the commit.
- Too 2y agoThats a very specific property of git, stemming from its distributed nature. Allowing one to push the history of a repo fetched from elsewhere. The receiver of the push is still considered an application server in this case. Whether or not GitHub solves this with a proxy or by reimplementing the git protocol and solve it in process is an internal detail on their end. GitHub is still “the application”. Other git forges do this type of auth in the same process without any proxies, Gitlab or Gerrit for example, open source and self hosted, making this easy to confirm. In fact, for such a hypothetical proxy to be able to solve this scenario, the proxy must have an implementation of git itself. How else would it know how to extract the commiter email and cross check that it matches the logged in users email? An application almost always has the best view of what a resource is, the permissions set on it and it almost always has awareness of “who” is acting upon said resource.
- zimbatm 2y agoThat's very cool; I didn't expect Garage to scale that well while being so young. Are there other details you are willing/allowed to share, like the number of objects in the store and the number of servers you are balancing them on?