4 ms·
Jason Scott, here. Disclaimer: I work for the Internet Archive (although I don't speak for the entire Archive) and I'm vaguely gung-ho on the place: https://twi
by textfiles 12y ago
Jason Scott, here. Disclaimer: I work for the Internet Archive (although I don't speak for the entire Archive) and I'm vaguely gung-ho on the place: https://twitter.com/textfiles/status/527549181175427072 https://twitter.com/textfiles/status/527549181175427072
I am also, as part of my job there, one of the largest individual uploaders of data to archive.org - I've added hundreds of thousands of individual items (texts, movies, music, websites) since I started working there in 2011.
So, moving on.
Welcome to the new user interface beta. I'm glad to see people toying with it, and the commentary and complaints are very, very welcome. As a smallish organization with a lot going on, the responses from people really digging down in the beta are very appreciated.
First, I'll say that the Beta interface is a "true" beta - it's the result of a lot of internal work, arguments, and discussions, but nothing is 100% set in stone. This isn't a beta like Gmail or a FPS trying to determine the rate of firing of the chaingun weapon: this is a lot of best-approach attempts at a whole range of goals. There's bound to be lots of responses from a lot of camps that are now coming forward. (For example, the new site has accessibility issues that need to be addressed.) If the term "beta" has been wrecked, stick with "prototype".
The internal name was V2, so I tend to keep calling it that.
V2 is the first major redesign of the main archive.org site in over a decade. And part of the conditions of this project (done by a handful of people) were to keep the old site (retroactively called V1) running, and mostly unchanged. That was a whole bucket of headache that isn't even obvious when you come into the site. (Anyone who has done this knows how it can be). With over 20 petabytes of data on the site, and millions of items and objects, spanning the whole environment without downtime is a feat in itself. So there's a whole range of philosophies being approached, but just getting the backend into a shape where it could sustain a new interface to it was a lot of non-obvious work.
Moving to the site as it is now.
Definitely slow. Definitely a shock. Definitely some great choices, some which might seem like head-scratchers. There is a designer, with a vision (his name is David) and there's been approaches to all the intended known shortcomings of V1 Internet Archive in this prototype.
One of the issues with Archive.org that's been an issue is non-responsiveness for different platforms - you got one site and that was it. Another was a lack of visual interface as an option. Now there is one.
The tagging and metadata efforts were spotty before now, because you were not really rewarded for doing so. The V2 site uses these tags and metadata extensively, and will continue to. This has been a nightmare for me, frankly - I've had to add logos to the 1,200 collections of items I've been uploading, and I'm doing descriptions as well as tags. But under the new system, the chance for finding things has increased exponentially.
There are definitely cases where I have to swap back to V1 to get kinds of "work done", because as an intense power-user, I do all sorts of grandiose work. But then again, 99% of my interaction with maintaining and adding content to the Internet Archive, I do through the API, and specifically through a python command-line interface we've had a developer working on for over a year:
https://pypi.python.org/pypi/internetarchive https://pypi.python.org/pypi/internetarchive
I've uploaded many thousands of items, analyzed and upgraded their metadata, and done search-and-modify runs by the hundreds with this tool. It's being constantly updated.
In the future, I expect us to see multiple improvements to the interface - one which is much more bandwidth and processor friendly, a version of the "view" (we have image and list right now) that is optimum for researchers, and so on. But I'll stress again:
- This is a prototype which was done with a pretty small team who had to keep the old site running as smoothly as possible, while doing essentially a decade of upgrade in one swoop;
- Now that it's "proven" that it works, refinement by the truckload needs to happen
- Your comments are not just welcome but encouraged
- Increased interest in the archive and the materials, and working together to find ways to access the petabytes of data in a meaningful way is not just a nice side benefit, but a vital core of the Archive's mission
Thanks for reading.
- Curmudgel 12y agoI got the old website when I visited it, but based on the image that another user posted: http://i.imgur.com/m2d1Gf4.png http://i.imgur.com/m2d1Gf4.png the text color of the description for search results is around #808080 on a #FFFFFF, which is a contrast ratio of 3.95 to 1 according to http://webaim.org/resources/contrastchecker/ http://webaim.org/resources/contrastchecker/ This fails both WCAG AA and WCAG AAA recommendations. Please increase the contrast, preferably by leaving the text the way it was (#000000)
- textfiles 12y agoThe team will see this thread. I'll leave it to them to take in the note.
- walterbell 12y agoThanks for the additional detail. > https://pypi.python.org/pypi/internetarchive https://pypi.python.org/pypi/internetarchive Can anyone use the API key, e.g. does it require auth for both upload and download? For upload, is an archive.org userid sufficient, or is a separate API key needed? Will the new metadata be available via the API? Onsite searches are usually less successful than a Google search with site:archive.org. Within the archive, it has been near impossible to create a URL-based search query that will find all editions of a given title/work. Will the new site/tagging help? Thanks to the entire archive team for a precious resource.
- Mithrandir 12y agoYou don't need a key for downloading. For uploading, you use an access/secret key pair from https://archive.org/account/s3.php https://archive.org/account/s3.php You then add that to the ia tool with "ia configure".
- textfiles 12y agoAnyone can generate an S3-like/API key. They have the same rights and restrictions as someone using other methodologies. So, for example, you can upload into the general Audio or Texts collections, but you can't upload, say, right into the Grateful Dead archive or the CD-ROM collections we have. In the future, we hope to have it that accounts will be assigned and de-assigned by some credential different than the current somewhat-binary approach we have now, but that functionality doesn't exist yet. So basically, upload is constricted like before. Download, however, is as unconstricted like before.