5 ms·
Voynich Manuscript: word vectors and t-SNE visualization of some patterns
- danharaj 11y agoYou know, I really like this, because it's an example of the kind of structure machine learning finds without my own understanding of the training set clouding my understanding of the machine's understanding.
- acqq 11y agoDo we get any new insight with this?
- Houshalter 11y agoThe author didn't give any novel insights about the text itself. It's just a proof of concept. If someone could translate a few words, it could give strong hints to what the translation should be of the other words.
- vonnik 11y agoI think this approach has a lot of potential, and I wonder what a statistical comparison of character co-occurrences between the Voynich manuscript and other writing systems would reveal. For anyone curious, here is Stephen Bax's video on his 2014 findings. https://m.youtube.com/watch?index=1&v=fpZD_3D8_WQ&list=LLATcCtXq6Eg7iFjmWQ1CNkA https://m.youtube.com/watch?index=1&v=fpZD_3D8_WQ&list=LLATc... He believes he has translated about 10 words in the manuscript, which is huge, and he thinks the script may have been invented to express a language once spoken between the near east and the Himalayas, maybe Turkic or Caucasian...
- benbreen 11y agoIt blows me away how many conflicting but semi-convincing theories there are about this manuscript (I read one fairly recently that argued for a New World origin of many of the plants, for instance). Which gets me thinking, has anyone analyzed the actual paper it's made out of to get any clues about where it originated? I know that spectroscopy can sometimes be used to make a guess at geographic origins of biological material but I haven't heard anything about it being used on the Voynich MS.
- skdfhksdf 11y agoThe vellum was carbon dated not too long ago, indicating that the book dated back to the 15th century: http://phys.org/news/2011-02-experts-age.html http://phys.org/news/2011-02-experts-age.html
- benbreen 11y agoI'd seen that, but I'm wondering if there's any way to use spectroscopy or some other technique to determine roughly what part of the world the vellum came from. It's been shown to be possible with wine: http://www.ncbi.nlm.nih.gov/pubmed/23682581 http://www.ncbi.nlm.nih.gov/pubmed/23682581 And (I think this one is really fascinating) some researchers at the Louvre awhile back even used spectroscopic analysis on a painting by Murillo to determine that the obsidian he painted on began its life in a 14th century Aztec obsidian mine! http://adsabs.harvard.edu/abs/2005NIMPB.240..576C http://adsabs.harvard.edu/abs/2005NIMPB.240..576C
- TillE 11y agoI've seen nothing that persuades me even a little bit away from the most common theory, which is that it's basically a hoax. It's a pseudo-occult book produced by some creative individual for fun and/or profit. There are some contemporary alchemical manuscripts which are written at least partly in code, for example, but the illustrations strongly suggest that Voynich is pure fantasy.
- throwaway2048 11y agoIt obeys statistical models of natural language that "fake text" would not at all. This sort of statistical modelling of language would almost certainly be unknown at the time it was produced. I find the hoax theory to be pretty unsatisfactory. Even if the overall goal was some kind of hoax, the text itself almost certainly carries semantic meaning, and that in itself is fascinating due to its apparent indecipherability.
- Houshalter 11y ago
- perone 11y agoThanks for the feedback, I also believe that this approach has a lot of potential, especially after the work of Stephen Bax, a few word translations can help us to figure out transformations that could allow translation in vector space.
- haddr 11y agoVery interesting approach, but I would say this is just a scratch. There are several factors that might really limit statistical analysis of this manuscript [1]. [1] http://www.ciphermysteries.com/2013/03/09/this-week-a-talk-at-stanford-on-the-voynich-manuscript http://www.ciphermysteries.com/2013/03/09/this-week-a-talk-a...
- cLeEOGPw 11y agoWell, one factor is the sample size is very small.
- Houshalter 11y agoThat blog post is assuming that the manuscript is a cipher. If so then it is unlikely that statistical tools will help much. But I don't think that's been proven. Many seem to believe that it's a real lost language.
- juxtaposicion 11y agoI implemented the new t-SNE in sklearn, so I've got some experience in reading these diagrams. Unfortunately, as wonderful as the algorithm is, it's extremely hard to interpret what it means rigorously. I've seen many diagrams that look like this one -- and they were generated from actual noise. So take the plots with a big grain of salt :) I'd be interested in seeing more direct evidence, like SVD factorizing the PMI matrix (which is what similar to what word2vec is doing) and seeing how much of the variance is explained by the first components. If you want to do this, check out: https://minhlab.wordpress.com/2015/06/08/a-new-proof-for-the-equivalence-of-word2vec-skip-gram-and-shifted-ppmi/ https://minhlab.wordpress.com/2015/06/08/a-new-proof-for-the...
- bearzoo 11y agoi once read that high perplexity can generate embeddings that are very tightly bound on a unit circle...not sure if this is what is going on
- Houshalter 11y agoThe diagram shown is only a visualization. The actual word vectors have many dimensions. To reduce them to 2 dimensions, they use a method which tries to keep vectors that are similar as close to each other as possible, but also unsimilar words apart. This creates the shape seen on the scatter plot. Just looking at the scatter plot by itself doesn't tell you anything about the underlying data.
- andreasvc 11y ago> Just looking at the scatter plot by itself doesn't tell you anything about the underlying data. Well if that were the case it would be perfectly pointless to make such a visualization... The goal of dimensionality reduction is to provide a useful summarization of the data; it is a valid question to ask to what degree it is successful at that.
- bearzoo 11y agoI did not claim anything about the underlying data.. The 2 dimensional embeddings were forced into the unit circle because the 'perplexity' hyper parameter for the t-sne was set too high. From the guy who helped make t-sne: When I run t-SNE, I get a strange ‘ball’ with uniformly distributed points? This usually indicates you set your perplexity way too high. All points now want to be equidistant. The result you got is the closest you can get to equidistant points as is possible in two dimensions. If lowering the perplexity doesn’t help, you might have run into the problem described in the next question. Similar effects may also occur when you use highly non-metric similarities as input.
- lawpoop 11y ago>>> model.most_similar("queen") [(u'princess', 0.519856333732605), (u'latifah', 0.47644317150115967),
- splitbrain 11y agoFirst time I hear about the cultural extinction theory. If that were the case, shouldn't there be more documents using the same script? But assuming the theory is right. Is there any way to decipher it without finding a Rosetta stone?
- alexwebb2 11y agoIf there were more documents with the same script, then it would have been decoded a long time ago, and you probably would never have even heard about it.