9 ms·
Just what would the models be trained on? Machine learning requires you to have a corpus of mappings between peices of texts in two languages, each of which hav
by rutherblood 5y ago
Just what would the models be trained on? Machine learning requires you to have a corpus of mappings between peices of texts in two languages, each of which have been established to be close in meaning -> thereby requiring there to be no decipherment problem withstanding beforehand.
There are a series of problems in decipherment of Minoan scripts from what I understand:
1. The Linear A & Cretan script are undeciphered in the sense that we do not know what phonetics the symbols stand for
2. The Minoan language hypothesized to be written in these scripts is undeciphered in terms of vocabulary (given we do not know the phonetics of the letter symbols) & by extension its semantics. At best a kind of crude syntax can proposed for it based on patterns of the word-symbols in texts of the hypothesized language written in those scripts.
- monkeybutton 5y agoThere are plenty of ML models that are trained unsupervised on text. What you would do next with your dead-language BERT, I don't know. But you could definitely make one.
- melissalobos 5y agoOne issue is that we don't have a lot of text, not even a megabyte of it(represented as unicode characters). So you could get a language model, but how could you judge its output? Maybe it would be really good at generating more similar text, but that text isn't probably super representative of things we would want to be able to read.
- didericis 5y agoI'm not a linguist and haven't done much machine learning, so forgive my naïveté, but aren't there a couple of core features/structures all languages have? That might not be enough to learn anything/might be what you're saying, but I wonder if there's a way to determine what a plausible grammar is and to try to identify what the nouns are. One you have the nouns, you could try just doing a huge substitution game on different texts in that language and try to see what guesses make the most sense for the most sentences.
- labster 5y agoI’m not an electrical engineer, but aren’t there a couple of core features that all electricity has, like ohms law? Surely we can figure out how an undocumented old computer functioned from a couple of electronic components, if we identify which parts are the resistors and which are the transistors.
- jcranmer 5y ago> aren't there a couple of core features/structures all languages have? The answer is 'no', especially in the sense that there's a relatively finite set of possible grammars that a putative natural language could be compared against. In terms of basic parts of speech, I believe that every language does have something that you can describe as a noun, but that's more or less the only "universal" part of speech (there are some languages that essentially don't have verbs--you "do a look" rather than "see", e.g.). The more serious problem, I think, is that the corpus of Linear A is simply too tiny to do any serious study, and I don't know how well the written corpus is at actually reflecting problems like segmenting text into words or morphemes. In essence, the available evidence is so paltry that you could justify just about any grammatical hypothesis, I suspect. If I'm understanding Chomskyian linguistics correctly (that's a really big if), there was originally thought to be an inherent "language organ" that strongly controlled grammar. But over time, and as linguistics documented more languages, the things that are universal in grammars in this subdiscipline has essentially been reduced to 'merge', which is an abstract concept that I'm pretty sure I don't understand.
- catlifeonmars 5y ago> The more serious problem, I think, is that the corpus of Linear A is simply too tiny to do any serious study This is probably a silly question: but if there the corpus is so small, why are we convinced that it has any meaning at all?
- ncmncm 5y ago"There was originally thought" meaning "Chomsky thought". But it was all obvious bollocks, contrary to elementary natural selection. Grammars necessarily have to be compatible with brain organization inherited from our primate ancestors, who obviously could have had no "language organ" carried about waiting to find some sort of use by their future descendants. Brain structures all need to be immediately useful for surviving or reproducing. One theory has been that language runs on a bit of brain hypertrophied as a sort of peacock's tail, not necessarily of any survival value, originally, but needed to impress a potential mate. It could have been used to carry a tune. Others have suggested language originated between mothers and children, growing out of lullabies. The two are not incompatible.
- senorsmile 5y agoThere was some interesting insights they've made on the Indus Valley script here: https://youtu.be/a_-obTZO6pY https://youtu.be/a_-obTZO6pY
- deleted 5y ago[deleted]
- cdrini 5y agoAll human languages have some commonalities. I remember this headline from 2016, when Google had an AI model that trained on some languages was able to somewhat successfully translate between language pairs it had never seen before: https://ai.googleblog.com/2016/11/zero-shot-translation-with-googles.html?m=1 https://ai.googleblog.com/2016/11/zero-shot-translation-with... Not sure what the state of research like this is now, or whether it's been applied to stuff like this. I hope so! EDIT: although it looks like it had seen the languages, but the pair it hadn't translated between.
- YeGoblynQueenne 5y agoAll the "unseen" pairs of languages had parallel texts with English so Google's vaunted "interlingua" was most likely natural English. But you shouldn't be downvoted for repeating Google's claim I think. It should be Google that is shamed for peddling such unmitigated nonsense.
- divbzero 5y agoAre there any commonalities in the type of subject matter considered worthy of recording in written text?
- catlifeonmars 5y ago> Just what would the models be trained on? Machine learning requires you to have a corpus of mappings between peices of texts in two languages, each of which have been established to be close in meaning -> thereby requiring there to be no decipherment problem withstanding beforehand. I was thinking of unsupervised machine translation specifically[1]. [1]: https://paperswithcode.com/task/unsupervised-machine-translation https://paperswithcode.com/task/unsupervised-machine-transla...
- yorwba 5y agoUnsupervised machine translation works by distribution-matching embeddings on the corpora you want to translate between. If the corpora are large enough, their distributions can be estimated robustly, and if they have sufficient overlap in the topics they cover, it's likely that words with similar distribution have similar meanings. So if there were a large amount of undeciphered Linear A inscriptions on a guessable range of topics, unsupervised machine translation might be worth a try. Unfortunately, there aren't that many Linear A inscriptions, and for those where the kind of content was known, the distribution matching has already been carried out by hand. E.g. from the article: "the word AB81-02, or KU-RO if transliterated using Linear B sound-values, is one of the few words whose meaning we do know: it appears at the end of lists next to the sum of all the listed numerals, and so clearly means ‘total’. But we still don’t actually know how to pronounce this word, or what part of speech it is, and we can’t identify it with any similar words in any known languages."