8 ms·
They had the S3 bucket originally set to allow listing - I grabbed it all before I reported it to an engineer there. Not going to post a torrent though - mostl
by rogerhoward 13y ago
They had the S3 bucket originally set to allow listing - I grabbed it all before I reported it to an engineer there.
Not going to post a torrent though - mostly cuz you all can figure it out anyhow. The images are all of the form:
NNNNNN01.JPG on that S3 bucket.
NNNNNN is the object ID (the internal record ID of the work of art) zero-padded to 6 digits. The object ID is the same value that appears, for instance, in the URL on their online collection:
URL for Irises: http://www.getty.edu/art/gettyguide/artObjectDetails?artobj=947 http://www.getty.edu/art/gettyguide/artObjectDetails?artobj=...
URL for the full-rez image on S3: http://gettylargeimages.s3.amazonaws.com/00094701.jpg http://gettylargeimages.s3.amazonaws.com/00094701.jpg
You figure it out :)
- uniclaude 13y agoNice. That would make a convenient fizzbuzz style interview question.
- hobs 13y agoI dont understand... wouldnt offering a torrent be more friendly than us all downloading it again?
- 6cxs2hd6 13y agoS3 buckets can act as torrents. Source: http://docs.aws.amazon.com/AmazonS3/latest/dev/S3Torrent.html http://docs.aws.amazon.com/AmazonS3/latest/dev/S3Torrent.htm...
- toomuchtodo 13y agoThis is exactly how I plan to serve up the compiled archive.
- hobs 13y agoThats pretty sweet actually, definitely +1. I said my original comment because the idea was to have use learn the power of wget by redownloading all files, which seemed pointless to me.
- toomuchtodo 13y ago1) Iterate through all the search results pages (1-94) to get objectids. 2) Iterate through objectids to fetch the following: a) xml data b) thumbnails (from getty servers) c) original work on cloudfront You're assuming the convention holds across all ~4600 objects. Never assume. Follow the links.