4 ms·
Perhaps we should think about establishing a public listing of checksums for important texts? Not sure exactly how it would work, you'd probably only want to r
by dmart 7y ago
Perhaps we should think about establishing a public listing of checksums for important texts?
Not sure exactly how it would work, you'd probably only want to run it against the main text so there would remain some flexibility for chapter headings, forewards, etc.
- caseysoftware 7y agoBut what is the "main text" and who has the definitive copy? At first glance, most people would think of chapter text and headings but even that can vary from edition to edition and printing to printing and country to country and translation to translation. How are typos/corrections handled and who has the right to make them? Then we need the schema/structure for the content. And that schema has to preserve whitespace because while it's not important most of the time, other times it's vital like in poetry. Obviously, we need to be careful of character encodings too. But most artists and writers consider their "work" to not just be the final product itself but the things that go around it like cover art, dedications, etc. At first glance, this is an "easy" problem but gets ugly quickly. It's also fascinating though. * I spent a few years at the Library of Congress working on their digital preservation project so lived and breathed these questions. When we worked with records (as in the musical kind), the album and cover art was just as important as the actual music most of the time.
- sgt101 7y agoOk - your points are all valid, but they are also huge. So if we try and address them we will get nowhere for a long time. What would the MVP of this be? What if we had a register of signed checksums with reputation and community selection (I trust these folks, but not these folks, verified artist wins)?
- holy_city 7y agoImo this is just good old fashioned quality control. When you can't trust your wholesaler you randomly sample the shipment and test it, rejecting the whole shipment (or exact penalties) for product outside an acceptable threshold. The publisher could supply a manuscript. It could be compared against the manuscript database (note: not with a checksum), and if it's too similar to other copyrighted works then it should flagged for manual review, and rejected if identical. That way you leave the door open for new translations/editions from different publishers, which should be treated as separate products. Then when the shipment of physical books arrive, you sample them via OCR + text differences and if it's outside your standards, reject it. But like the article mentions, that's expensive.
- totalperspectiv 7y agoInterestingly I feel like leaning on some work in the field of genomics where comparing different formats, each of which contain potential 'errors' is something done. Search engines also seem to do something like this already as well https://en.wikipedia.org/wiki/MinHash https://en.wikipedia.org/wiki/MinHash. MinHashing is also used in genomics. White space, if handled appropriately are just more characters. But most literature won't be available via flat text files I imagine. Some sort of image -> text converter would be needed, which I bet exists, but may require tweaking to allow more fine grained representation of white spaces. Authors publishing new texts could release some kind of checksum to go with it ... or to venture into waters that I don't know much about ... could blockchain be used in some way to keep a record of edits to text? I'm sure someone out there has put a lot of thought into guaranteeing the authenticity of a text. Edit to add: This is interesting to think about in terms of all media. Wasn't it just last week that there was a headline about Boris Johnson editing some of his old videos? How do you guarantee that the information that you viewed a year ago is the same today as a year ago?
- Tomte 7y ago> But what is the "main text" Right. See this incredible example. Both versions are authentic, in a way. https://www.theguardian.com/books/2016/aug/10/cloud-atlas-astonishingly-different-in-us-and-uk-editions-study-finds https://www.theguardian.com/books/2016/aug/10/cloud-atlas-as... > Mitchell himself explains the reasons for the discrepancies in an interview quoted in Eve’s paper: they occurred because the manuscript of Cloud Atlas sat unedited for around three months in the US, after an editor there left Random House. Meanwhile in the UK, Mitchell and his editor and copy editor worked on the manuscript, but the changes were not passed on to the US.
- _eqet 7y agoOr a signing certificate by a CA?
- dogweather 7y agoI've been brainstorming about this, applied to gov't legal texts, and their republication. Nature, June 6, has a very interesting article on attribution of authorship, that seems relevant and is on my reading list: Credit Data Generators for Data Reuse, pp. 30-32.