3 ms·
Internally, Ambry stores large objects as a series of chunks. At LinkedIn we've found that 4mb makes a good chunk size. So (for example), your 40mb upload cou
by ambry 3y ago
Internally, Ambry stores large objects as a series of chunks. At LinkedIn we've found that 4mb makes a good chunk size. So (for example), your 40mb upload could be streamed directly where Ambry would handle chunking it, or you could chunk it yourself. Client-side chunking has the advantage of cleaner resumption from a broken connection. After all the chunks are uploaded, Ambry creates (or the client asks Ambry to create[1]) a special metadata blob which contains a listing of all the chunk blob IDs. Clients can use the blob ID of the special metadata blob as a single reference to the large (composite) object without needing to worry about the underlying chunking for GET and DELETE operations.
Depending on context, anything larger than the chunk size could be considered a large object, but using this method Ambry easily handles gigabyte- and larger-sized objects.
[1]: https://github.com/linkedin/ambry/blob/9b7a49ac79b1678fd7fd7a1871ef8f8c0d60c29b/ambry-frontend/src/main/java/com/github/ambry/frontend/PostBlobHandler.java#L58 https://github.com/linkedin/ambry/blob/9b7a49ac79b1678fd7fd7...
- donavanm 3y agoVery similar interface to S3 multi part uploads and (IIRC) google cloud storages equivalent. Relatively easy to put that composition & resumption logic on the client side.
- rad_gruchalski 3y agoThank you. This answers multipart uploads question. Another one, if I may. Can a large object be uploaded to Ambry using chunking? Like s3, gcs, and azure blob storage do? The upload client splits an object into 8KB chunks (configurable) and uploads them in parallel using byte ranges. Each chunk upload has its own retry policy. This prevents a situation where bad egress would interrupt large upload as a whole. Essentially how sftp put operation works. This is different from the multipart upload discussed earlier. Can this be done with Ambry? I don’t see any immediate mention of byte ranges in the source you linked.
- ambry 3y agoShort answer: Yes. However, Ambry doesn't use byte-ranges for the upload. Each chunk can be uploaded separately and the client then requests a _stitch_ operation[1]. In that operation, the client specifies the list of blob IDs for each chunk (in order) and asks Ambry to create a metadata blob listing those IDs. Ambry then returns the blob ID for the metadata blob. [1]: https://github.com/linkedin/ambry/blob/9b7a49ac79b1678fd7fd7a1871ef8f8c0d60c29b/ambry-frontend/src/main/java/com/github/ambry/frontend/PostBlobHandler.java#L58 https://github.com/linkedin/ambry/blob/9b7a49ac79b1678fd7fd7...