3 ms·
> Keep in mind that classifications tend to be married to the notion of physical records stored in a specific location, which is generally not the case for elec
by Mithriil 3y ago
> Keep in mind that classifications tend to be married to the notion of physical records stored in a specific location, which is generally not the case for electronic data. For the latter, useful descriptions of the work or information, information on its provenance, unique item identifier(s), and cross references between related items (e.g., source or derived data) might be more relevant information to capture.
The many replies here on HN have made me come to this conclusion. As long as one can find the datasets they need for a use case. Solid metadata standards for good description and indexing is a must.
> It's quite useful to think of how any system you specify will be used, by whom, and how it will be applied and maintained [...]
These sentences are very good starting points to design said metadata standards internally. I'll keep these questions in mind.
> There are also organisations working with electronic data collections at scale, including the Internet Archive and the Wikimedia Foundation, most particularly Wikidata, which might be of interest or relevance to you.
These are very good suggestions, thank you!
- dredmorbius 3y agoThese sentences are very good starting points to design said metadata standards internally. I'll keep these questions in mind. There's a formulation of the cataloguer's / librarian's goals in creating a catalogue which I tried to find and reference above which I'd failed to do. The "Six Functions of Bibliographic Control" are one version of that. Abbreviating slightly: 1. Identifying the existence of all types of information resources as they are made available. 2. Identifying the works contained within those information resources or as parts of them. 3. Systematically pulling together these information resources into collections in libraries, archives, museums, and Internet communication files, and other such depositories. 4. Producing lists of these information resources prepared according to standard rules for citation. 5. Providing name, title, subject, and other useful access to these information resources. 6. Providing the means of locating each information resource or a copy of it. <https://en.wikipedia.org/wiki/Cataloging_(library_science)#Six_functions_of_bibliographic_control https://en.wikipedia.org/wiki/Cataloging_(library_science)#S...> Citing: - Hagler, Ronald (1997). The Bibliographic Record and Information Technology, 3rd ed. Chicago: American Library Association. - Taylor, Arlene G., & Daniel N. Joudrey (2009). The organization of information. 3rd ed. Englewood: Libraries Unlimited, pp. 5-7 The Wikipedia Cataloguing article raises numerous other points, practices, and history. Keep in mind that your own needs might overlap with these only partially, including some, excluding others, and including other considerations (e.g., provenance and derivative datasets or sources, generating and/or reporting software). On the physical storage/retrieval aspect: idiosyncracies of the US Library of Congress classification become far more understandable when one considers that, say, topics of history and geography, local and proximate topics are not only of greater concern, but historically comprised a far larger portion of the collection. (This is of course changing.) Having gone though the classification in some detail, I'd also been struck by how within the law classification, the degree of detail provided for California and New York State is vastly greater than nearly all other states (a handful of other industrial states also have relatively detailed classifications). The classification, in other words, maps onto the archived works. The history and evolution of the classification is itself a story, though one you may not want to dive into. It's pretty substantively documented in the annual Librarian's Report to Congress, which begin in the 1860s and continue through the present. Recent years (to about 2000) are available through the loc.gov website, older reports, to the first, at the Hathi Trust. Much of the first 50 years or so details the challenges of dealing with a swelling collection in a much-too-insufficient space (either the Old Senate Chambers or the Old Supreme Court Chambers, if memory serves). The current freestanding main structure, the Thomas Jefferson building adjacent to the Congress and Supreme Court both of which it serves though an underground messaging-and-delivery system was opened in 1897. The book delivery system is described on page 7 of the 1897 report, with the five requested books reaching the Capitol in eight to twelve minutes: <https://babel.hathitrust.org/cgi/pt?id=mdp.39015036735044&view=2up&seq=12&size=125 https://babel.hathitrust.org/cgi/pt?id=mdp.39015036735044&vi...> Photo of the old Library (still within the US Capitol) in the 1890s: <https://en.wikipedia.org/wiki/Library_of_Congress#/media/File:Library_of_Congress,_showing_three_levels_crowded_with_stacks_of_books_and_newspapers_LCCN2017646700.jpg https://en.wikipedia.org/wiki/Library_of_Congress#/media/Fil...> Recent LoC annual reports: <https://www.loc.gov/about/reports-and-budgets/annual-reports/ https://www.loc.gov/about/reports-and-budgets/annual-reports...> The Hathi Trust archive of reports 1866--2007: <https://catalog.hathitrust.org/Record/000072049 https://catalog.hathitrust.org/Record/000072049>