7 ms·
In what way do you think it’s over complicated? Genuinely interested. Not sure about Minio, but most S3 clones don’t clone all API operations, just the ones th
by diroussel 4y ago
In what way do you think it’s over complicated? Genuinely interested.
Not sure about Minio, but most S3 clones don’t clone all API operations, just the ones they and their customers/users need.
- vbezhenar 4y ago1. Buckets is unnecessary concept. Why use it at all, when we have domain and path. That should be enough. 2. Authentication is too convoluted. Signings, etc. Simple `Authorization` header should be enough. 3. Uploading a file should be as simple as PUT /path/to/file. Instead of multipart upload, header `Range: bytes=0-1023` should be used. 4. List is unnecessarily flexible. Why should users have the ability to use any character as path delimiter. `/` should be fixed. List operation should use ordinary REST conventions, like `GET /dir/?from=token` should return JSON with items inside `dir` and continuation token. Basically I want to be able to use object storage with simple `fetch` and few lines of code. Without all that SDK nonsense.
- crgwbr 4y agoEvery one of the things you list as unimportant, someone else lists as a mandatory requirement. And that’s why S3 won. No one uses everything it does, but everyone uses a different 10%.
- stingraycharles 4y agoI don’t think S3 won because of its protocol, but rather being the first in its kind and dominating the market. Then other vendors started competing products, and to easy transition, implemented (parts of) the S3 protocol.
- diroussel 4y agoIt won because it was the first, but also because it’s API was designed to be cloned. Being cloned was a success criteria for the the S3 team. I read that somewhere recently. I don’t have a source to hand. Personally I this the API is ok, but if they did designer it to be cloned, they could have done it slightly better.
- inkyoto 4y ago> 1. Buckets is unnecessary concept. Why use it at all, when we have domain and path. That should be enough. Bucket is a very important concept in S3 buckets. Because a S3 bucket is a container filled with S3 objects. A S3 object can have a fine-grained lifecycle policy defined and attached to them or can have one, all or a combination of the following: 1. S3 objects can be versioned. S3 object versioning is very useful is many scenarios, and the S3 bucket can automatically purge expired object versions for the user after a desired period of time with zero effort – it is defined in the object lifecycle policy. 2. S3 buckets are event-driven. A S3 bucket can be configured to emit an event when something happens to a S3 object in the bucket. There is no need to poll a S3 bucket for changes, for it can notify interested consumers via a new event. S3 buckets offer a rich set of events to choose from that cater for multiple use cases. S3 bucket events are a boon that also greatly simplifies the solution architecture, design and implementation. 3. S3 objects can be locked in governance or in compliance modes for a certain amount of time. Both are useful to prevent the accidental data loss or preclude nefarious attempts to corrupt/delete the important data. The compliance object lock is very useful to satisfy statutory compliance requirements when sharing the data with regulatory and government bodies. Object lock is applied at the S3 object version level and can also be lifted automatically via a defined lifecycle policy wth zero coding effort required. 4. S3 bucket can be set up to automatically replicate the data across multiple S3 buckets, either across different accounts or across regions. S3 buckets can be replicated in full or in part with zero extra coding required – replication rules are defined in the replication configuration. 5. S3 buckets also allow one to choose a storage class which is equivalent to the automatic data archival. > 2. Authentication is too convoluted. Signings, etc. Simple `Authorization` header should be enough. S3 buckets are often useful to share the data with external business partners, therefore the identity of the party accessing accessing the bucket is to be asserted first. Multiple business partners can be granted access to the same bucket either in its entirety, or the S3 bucket can be partitioned and the partner can only be granted access to that partition. Therefore proper authentication is important. > Basically I want to be able to use object storage with simple `fetch` and few lines of code. Without all that SDK nonsense. Then you are probably do not need a S3 bucket and can stand up your own HTTP server with a block storage attached to it or explore alternative solutions. S3 is a complete, feature rich software product that may or may not be suitable for one's needs. In most cases, though, it is very simple and easy to use and having to use the SDK is a non-existing problem.
- vlovich123 4y agoTech lead of Cloudflare R2 if that's helpful as I've thought about this a bit. First things first, a bit of a shameless plug [1] to answer this piece: > Basically I want to be able to use object storage with simple `fetch` and few lines of code. Without all that SDK nonsense. const object = await env.MY_BUCKET.put(objectName, request.body, { httpMetadata: request.headers, }) return new Response(null, { headers: { 'etag': object.httpEtag, } }) It's called `put` instead of `fetch` but basically same concept. > Buckets is unnecessary concept. Why use it at all, when we have domain and path. That should be enough. The utility is because it's important for accounting, organization, ACLs, jurisdiction controls, analytics, different business units needing different properties, etc. If you use a virtual-hosted URL then you have something like https://<bucket>.<account>.r2.cloudflarestorage.com/<path> https://<bucket>.<account>.r2.cloudflarestorage.com/<path>. That's your domain + path. Buckets is still a useful higher-order organizational concept. FWIW Azure Blob storage and S3 were released at roughly the same time. Both had the same organizational concept. The other thing to think about is that folders are tree structures and that just doesn't have the same scaling properties to billions or trillions of files that a flat namespace does. > Authentication is too convoluted. Signings, etc. Simple `Authorization` header should be enough. Yeah probably. 16 years ago, HTTPS was a lot less common. Even still, some users do have legitimate concerns about middleware injecting headers / intercepting signed requests & replaying them. Sigv4 does address that in a comprehensive way. It also ensures some level of protection against corruption that can happen before TLS encryption and after TLS decryption. > Uploading a file should be as simple as PUT /path/to/file. Instead of multipart upload, header `Range: bytes=0-1023` should be used. I agree that multipart has various misdesigns. The one that actually really gets me is ListParts which lets you retrieve the original parts of a multipart file - like wtf that seems legit useless. At the time though (16 years ago now), resumable uploads didn't exist as a standard (unless you count WebDAV which I don't because it failed miserably). There's now an attempt to standardize something using TUS a starting point but I don't see much incentive for Amazon to really update S3 here because that's not where they're investing their R&D budget. It's up to the rest of us (particularly browser vendors) to create a standard and enough demand that they have to update their API. FWIW the range thing has to be careful because aside from resuming the other piece that's important for resumable uploads is concurrent uploads of 1 part + integrity verification. One of the things I'm thinking about is whether to try to retrofit whatever the standard becomes into our S3 endpoint or put it somewhere else. It probably makes the most sense for public buckets (aka static website hosting) but that's a whole other ball of wax (security, etc). To defend multipart as well, the other piece it does really well is the ability to track the state of the upload which regular "PUT /path/to/file" doesn't. Downloads can only publish the object at the path once they complete so then how do you identify ongoing concurrent uploads to the same path? An overwrite upload can be thought of as a strongly consistent "delete + publish" step. > List is unnecessarily flexible. Why should users have the ability to use any character as path delimiter. `/` should be fixed. List operation should use ordinary REST conventions, like `GET /dir/?from=token` should return JSON with items inside `dir` and continuation token. Eh. In terms of complexity, the '/' delimiter really doesn't add any complexity to the API and "users" here are application developers, not end users. It adds a bunch of horrible implementation complexity and it took us a while to get this piece correct, but the complexity was having any delimiter, not around letting the user pick a custom delimiter. And having a delimiter is definitely helpful because people like to view folders. Would I wish that they simplified and require the delimiter to be 1 ASCII character instead of supporting arbitrary UTF-8 strings? Sure, but only from a test coverage perspective. My observation is that a lot of the complexity from the APIs stems from two things. The first is that REST is a terrible way for accessing and mutating a strongly consistent data store. 16 years ago it made sense though and even today if you want interop with the browser, it's all you got. The second is that the more general problem of figuring out a sane API for a distributed strongly consistent database is an unsolved problem. HTTP has stood the test of time on that front but I wonder if something like cap'n'proto might be a better fit. [1] https://developers.cloudflare.com/r2/examples/demo-worker/ https://developers.cloudflare.com/r2/examples/demo-worker/
- thayne 4y ago1. The bucket is the domain. If you have a self hosted, single tenant minio instance, you might not need buckets, it doesn't really hurt you to just set up one bucket, but for people that do need buckets, it would be a lot more difficult to build them on something that didn't have a similar concept. 2. I'm with you on this one. At least as long as it is over https. The signing process might make more sense in the context of plaintext http. And it might make the implementation of signed urls easier, idk. But, yeah definitely one of the rough points of the s3 api 3. An upload can just be a simple PUT in many cases. And for the cases where you do need a multipart upload: The Range header doesn't work the way you suggest, it is for requesting a range for the response, not giving a range for the data in the request. And even if you used a different custom header, you would need some way to indicate when you have uploaded everything. Maybe the API could be simplified a little bit, but I think quite a bit of it is essential complexity for the cases where you need multipart upload. 4. There are use cases where you want to use a different delimiter. And it's worth noting that as an object store, the data isn't really organized into folders, the / is just part of the object key. Definitely agree that it should support a JSON response though. And I do think it would make sense for / to be the default delimiter, because that is usually what people want. > Basically I want to be able to use object storage with simple `fetch` and few lines of code. Without all that SDK nonsense. If it weren't for 2, this would be possible. And that is a big reason why I wish the authentication was simpler. All that said, I do wish there was a standardized specification for an object storage API, that all the different providers adhered to, rather than everyone making an "s3 compatible" API, where exactly what it means to be compatible with S3 varies from product to product.
- jsmith45 4y ago> The Range header doesn't work the way you suggest, it is for requesting a range for the response, not giving a range for the data in the request rfc9110 actually does permit servers to implement range on PUT, with partial update semantics, but warns that user agents can only safely do this if they know the origin server supports this, as if not the server could fallback to replacing the whole resource with just the uploaded range! (To my knowledge such support is not very common. Also due to risks of being misinterpreted by caches and other intermediaries, using the dedicated PATCH verb is probably better.) But this (ranged with PUT) is only semantically appropriate for updating an existing valid resource. For example, if you have database like file format, you can upload only the changed ranges which could be much much smaller than the whole file. Or for something like a log file, one could append new content to the end. For most files, breaking them into many parts for initial upload and range putting each part would be semantic inappropriate, since the resource is likely to only actually be valid once the last part is uploaded.
- jasongi 4y ago> Simple `Authorization` header should be enough. Cool so now you can’t fetch non-public S3 objects in the browser unless you’re fetching over Ajax? What’s in the Authorization header - an API key? There’s a reason services serious about security don’t do it like that anymore.
- klauspost 4y agoYou can use presigned urls for that.