3 ms·
Interestingly I feel like leaning on some work in the field of genomics where comparing different formats, each of which contain potential 'errors' is something
by totalperspectiv 7y ago
Interestingly I feel like leaning on some work in the field of genomics where comparing different formats, each of which contain potential 'errors' is something done.
Search engines also seem to do something like this already as well https://en.wikipedia.org/wiki/MinHash https://en.wikipedia.org/wiki/MinHash. MinHashing is also used in genomics. White space, if handled appropriately are just more characters.
But most literature won't be available via flat text files I imagine. Some sort of image -> text converter would be needed, which I bet exists, but may require tweaking to allow more fine grained representation of white spaces.
Authors publishing new texts could release some kind of checksum to go with it ... or to venture into waters that I don't know much about ... could blockchain be used in some way to keep a record of edits to text?
I'm sure someone out there has put a lot of thought into guaranteeing the authenticity of a text.
Edit to add:
This is interesting to think about in terms of all media. Wasn't it just last week that there was a headline about Boris Johnson editing some of his old videos? How do you guarantee that the information that you viewed a year ago is the same today as a year ago?