3 ms·
This a really dumb question, but since I've never used Glacier what does the workflow for a Glacier application look like? I'm used to the world of immediate a
by physcab 10y ago
This a really dumb question, but since I've never used Glacier what does the workflow for a Glacier application look like? I'm used to the world of immediate access needs, and fast API responses, so I can't imagine sending off a request to an API with a response "Your data will be ready in 1-5 hours, come back later".
- needcaffeine 10y agoIt may be for backing up infrequently accessed data (compliance logs, etc) for example. Hypothetical: you create a logging service for users to send all their log data to you. You promise 365 days of archives, but 30 days of data accessible at any time. You create a lifecycle rule on your S3 bucket to automatically archive data to Glacier 30 days after creation. On the 31st day, your user decides they want to look at an old log. They click the big Download button. You display a message saying they'll get an email from you when that data is ready to download.
- scrollaway 10y agoGlacier is not for data you want readily available. It's for when you care more about storage than access.
- pauloday 10y agoRight, but if you want to retrieve some data through their API, how does it work? Normally you open the connection, ask for the data, then receive it and close the connection, does that change if there's a 5+ hour wait between the ask and the receive? Do you just leave the connection open? Provide them with a webhook to call when it's ready? I don't personally care about the answer but I'm pretty sure that's what they were asking.
- extra88 10y agoAt work our needs are simple, we manually run the aws cli to sync files up to S3 where there's a 1 day lifecycle policy to move them to Glacier. We don't use the API for restores, we do those through the web console and check back in a few hours to see if the files are downloadable. I think through the API you do not leave the connection open, you check with whatever frequency you want and when it's ready, the response will include the temporary location on S3 for the file.
- dgemm 10y agoWell the CLI gives you a job ID that you can use to check on the status and retrieve when it's ready. You can also ask to be notified by SNS. http://docs.aws.amazon.com/cli/latest/reference/glacier/initiate-job.html http://docs.aws.amazon.com/cli/latest/reference/glacier/init...
- Twirrim 10y agoWith Glacier you submit an "InitiateJob" request to say "Fetch me this archive". That returns you a job ID in the response. From there you can submit a "DescribeJob" request, with that Job ID as the parameter, and the Glacier service responds with the state of the job. Once the job is marked as complete, you submit a "GetJobOutput" request with that Job ID. That response is the archive body. (similar to how you'd do a GET request from S3). You've got 24 hours to start the download of the archive before you'll have to repeat the entire InitiateJob->GetJobOutput cycle again.
- int_19h 10y agoWhy not? If you have ever worked with any asynchronous API, you have already been introduced to the "come back later" model. Does it really matter how much later it is?
- pugio 10y agoI work on quality control medical data (MRI images) and have huge data sets from machines going back over a decade. Most of the useful stuff is extracted metrics (stored in a db), but every now and then we need to pull up a data set and run updated analysis algorithms. We'll usually keep the latest couple of years in S3, and the rest in Glacier. The data trove is fairly unique, and valuable in being the only of its kind, but we don't need anywhere near instant access to most of it.