4 ms·
I worked for OCLC for about five years, so I have a lot of sympathy for the pains that archivists go through when trying to standardize content metadata. Can't
by xwkd 5y ago
I worked for OCLC for about five years, so I have a lot of sympathy for the pains that archivists go through when trying to standardize content metadata.
Can't be that hard of a problem to solve, right? I mean, we're only talking about coming up with a good solution for cataloging the sum total of all human knowledge.
(Not a small undertaking.)
Libraries are caught between a rock and hard place. On the one hand, they have to be stewards of this gigantic and ever growing mound of paper, and on the other hand, they have to deal with the lofty ideas of the W3C and Tim Berners-Lee, trying to connect all of that paper to his early 2000s vision of the ultimate knowledge graph / the semantic web. No wonder we're sticking to MARC 21.
The only people using that web for research are universities. The rest of us are using WHATWG's spawn-of-satan hacked together platforms because the technology always moves faster than catalogued content. Hell, at this point, GPT-3 is probably a better approach to knowledge processing than trying to piece together something actionable from a half baked information graph born of old programmers' utopian fever dreams.
- sswaner 5y ago“old programmers' utopian fever dreams” - accurate description of me when I first found Dublin Core. Never made it to production. It was unnecessary overhead for describing internal content.
- ghukill 5y ago>> "Hell, at this point, GPT-3 is probably a better approach to knowledge processing than trying to piece together something actionable from a half baked information graph born of old programmers' utopian fever dreams." Greatest thing I've read on HN. As a librarian and developer, can confirm. At least in most cases...(slipping back into fever dream)....
- smitty1e 5y agoSayre's Law: "In any dispute the intensity of feeling is inversely proportional to the value of the issues at stake." https://en.m.wikipedia.org/wiki/Sayre%27s_law https://en.m.wikipedia.org/wiki/Sayre%27s_law Was there ever a less essential fiefdom than citation formats?
- dsr_ 5y agoHave you noticed that search engines suck? Citation formats are an attempt to end up with a search engine that doesn't suck. (They can do lots of other things along the way.) Options: 1. Accept any format and try to turn it into your own internal representation. Problems: (a) your own representation is a citation format; (b) you need to write an infinite number of converters and hire an infinite number of trained people to work out the edge cases. 2. Accept a limited number of common formats. Problem: your search engine will not be useful for a majority of the corpus. 3. Convince everyone that your new citation format is unstoppable. Problems: (a) convincing everyone; (b) actually having that citation format cover, say, 99% of the cases; (c) XKCD#927 (Standards). Dublin Core is/was a terrible attempt at a type 3 solution.
- karaterobot 5y agoSome people catalog information every day, so it matters to them for practical reasons. Same with people who rely on those resources being accurately cataloged.
- smitty1e 5y agoI'm ok with obsessing about the data. Presentation, an order of magnitude less. Substance >>> style.
- karaterobot 5y agoCataloging format is to style as database modeling is to... well, style. It's got nothing to do with aesthetics, it's about describing the data in a way that makes it useful later.
- brazzy 5y agoYou're missing the point. Citation formats are not a matter of style here. It's about making research results easier to find, which directly affects the quality of new research.
- lyaa 5y agoThe problem with models like GPT-3 is that they are unable to differentiate between information sources with different "trustworthiness." They learn conspiracies and wrong claims and repeat them. It's possible to feed GPT-3 prompts that encourage it to respond with conspiracies (i.e. "who really caused 9/11?") but it also randomly responds to normal prompts with conspiracies/misinformation. A recent paper[0] has looked into building a testing dataset for language models ability to distinguish truth from falsehood. [0] https://arxiv.org/abs/2109.07958 https://arxiv.org/abs/2109.07958
- pjc50 5y ago> A recent paper[0] has looked into building a testing dataset for language models ability to distinguish truth from falsehood. Isn't this a massive category error? Truth or falsehood does not reside within any symbol stream but in the interaction of that stream with observable reality. Does nobody in the AI world know Baudrillard?
- rendall 5y agoIt should be at least theoretically possible for an AI to identify contradictions, incoherence and inconsistency in a set of assertions. So, not identifying falsehood per se, but assigning a fairly accurate likelihood score based solely on the internal logic of the symbol stream. In other words, a bullshit detector.
- nimish 5y ago> It should be at least theoretically possible for an AI to identify contradictions, incoherence and inconsistency in a set of assertions Not in the slightest. Likelihood of veracity is opinion -- laundering it as fact to make some people feel better doesn't make it any more subjective, or authoritative.
- rendall 5y agoI think we're not disagreeing, exactly. As a simple example, here is a set of assertions: * The moon is made of green cheese * The moon is crystalline rock surrounding an iron core It wouldn't take an AI to see that both of these can't be true, even if we weren't clear about what a moon is made of, exactly. Some of our common understanding could contain more complicated internal contradictions that might be harder for a human to tease apart, that an AI might be able to identify.
- xg15 5y ago> Hell, at this point, GPT-3 is probably a better approach to knowledge processing than trying to piece together something actionable from a half baked information graph born of old programmers' utopian fever dreams. I mean, at this point, wouldn't it be a lot simpler to go back to the middle ages (or earlier) and have a few humans memorize all that stuff? It's not as if GPT3 would give you any more insight than that approach...
- wvh 5y agoHaving worked with DC, Marc21 and some of its precursors, I feel that one major problem with the approach is that cataloguers try to infinitely cut up metadata into ever smaller pieces until you end up with a impossible large collection of fields with an overly high level of specificity and low level of applicability. For example, the middle name of a married name of a person that contributed to one part of one edition of the work in a certain capacity etc. You end up with so much unwieldily specific metadata that consistent cataloguing and detailed searching become nigh impossible. Ever since search engines, the world has moved into using mostly flat search engines rather than a highly specific facet search for deep fields. Of course, one would want something a bit smarter than full text search, but the real, human world is so complex that trying to categorise any data point quickly becomes an exercise in futility.