8 ms·
Critical Behavior from Deep Dynamics: A Hidden Dimension in Natural Language
- hacker43 10y agohmm...maybe the way we then encode data digitally is also "wrong" or should I say unnatural.
- jakub_h 10y agoIt seems to me that if we humans are the only ones doing this, it is either trivially natural or trivially unnatural based on whether you include or exclude our creations in/out of the natural world.
- wingcommander 10y ago"... which explains why natural languages are poorly approximated by Markov processes." Is that a joke? It's 2016 and you think you need to explain that a Markov process is a poor approximation of natural language? This has been obvious for computational linguists, and anyone working in the field, from day one.
- w_t_payne 10y agoI wonder what programming languages look like? I'd guess a lot like natural language?
- xtacy 10y agoNot quite, but as the paper says: Corollary: No probabilistic regular grammar exhibits criticality. In the next section, we will show that this statement is not true for context-free grammars (CFGs). That is, there exists CFGs that exhibit criticality. Programming languages are often parsed by CFGs, so it's likely that some programming languages exhibit the same criticality structure as natural languages.
- rntz 10y agoI think that means only that programs written in a language described by a CFG could exhibit "criticality", not that they will. "Exhibiting criticality" is a property of a distribution (e.g. a corpus of human-written programs or an algorithm for generating programs), not of a grammar, IIUC.
- laretluval 10y agoThey say their results from Wikipedia data were influenced by XML tags, so programming languages might look a lot like their Wikipedia data. For reasons like this, I don't trust their empirical results at all.
- mrcactu5 10y agoThe Bach data consists of 5727 notes from Partita No. 2 [11], with all notes mapped into a 12-symbol alphabet consisting of the 12 half-tones {C, C#, D, D#, E, F, F#, G, G#, A, A#, B} with all timing, volume and octave information discarded. I was good until the last part.
- curiousgal 10y agoCan you elaborate?
- deleted 10y ago[deleted]
- eximius 10y agoWas this a sample of text generated by a Markov chain or their 'new' model?
- eximius 10y agoOh. I feel a little silly.
- versteegen 10y agoI misread your comment as saying you discarded the paper. Actually, 'discarded.' is part of the quote.
- mark_l_watson 10y agoThis seems like a very important paper, basically showing that Markov models with exponential decay of influence of tokens by distance are often a poor model, where as deep neural networks with LSTM (long short term memory) has power law decay of influence decay, which performs better for a variety of sequential data. BTW, I went to the North American Association of Computational Linguistics conference in April and it seemed like half the papers used LSTM. Edit: the NAACL 2016 papers are here: http://aclweb.org/anthology/N/N16/ http://aclweb.org/anthology/N/N16/
- osipov 10y ago@mark_l_watson -- the paper also seems important in another respect -- in the conclusion the authors suggest abandoning loss functions as optimization objective functions in machine learning and replacing them with mutual information functions.
- forgotpwtomain 10y ago> This seems like a very important paper, basically showing that Markov models with exponential decay of influence of tokens by distance are often a poor model, where as deep neural networks with LSTM (long short term memory) has power law decay of influence decay, which performs better for a variety of sequential data. As someone outside of this field, it seems to me that this kind of result should have been very obviously foreseeable, hindsight bias and all that of course - but I would never have considered Markov processes to be an adequate predictability model for natural language. Though obviously the formalized results are important. Could someone with more knowledge comment on what the current working assumptions were prior to this paper and what the consequences would be?
- thomasahle 10y ago> I would never have considered Markov processes to be an adequate predictability model for natural language. Would you consider LSTM an adequate model?
- TheOtherHobbes 10y ago
- meeper16 10y agoHere's an implemented Markov + word2vec chatbot http://lexcognition.com/lexi.html http://lexcognition.com/lexi.html
- the_duke 10y agoI am VERY impressed: ----- you: tell me something me: Don't speak for me first time I asked it you: tell me about your mother me: just like my mother is on that AK47 diet you: Tell me about artificial intelligence me: Okay , maybe not intelligence capabilities, etc you: ask me something me: Points bow No one would ask this haha
- deleted 10y ago[deleted]
- BenoitP 10y ago> [...] A Hidden Dimension in Natural Language Mmmh > [...] We show that in many data sequences — from texts in different languages to melodies and genomes Hum, ehrm > [...] natural languages are poorly approximated by Markov processes. Alright, alright > [...] This model class captures the essence of probabilistic context-free grammars Ok, ok > [...] and cosmological inflation Wat. Out of nowhere, Creation of the Univerve. ------------- I'm always baffled by the ability to draw parallels. Did a colleague take at peek at the screen and said, hey I have the same equations?
- ASpring 10y agoOne of the common threads I've noticed between the best researchers I know is their uncanny ability to draw these parallels between the most obscure domains.
- jakub_h 10y agoHeh! For some reason, that reminded me of this famous conversation: "...and that, my liege, is how we know the Earth to be banana-shaped." "This new learning amazes me, Sir Bedevere. Explain again how sheep's bladders may be employed to prevent earthquakes."
- Gargoyle 10y agoWorth noting the authors, Henry Lin and Max Tegmark, are both astrophysicists. Among other things.
- spdustin 10y agoNot to nitpick, Max Tegmark is a cosmologist. I only recently learned the difference when I called a cosmologist friend an astrophysicist. Cosmology deals with the big stuff, almost philosophically, like: "where did the universe come from" and "what is the fate of the universe", while astrophysics deals with the nature of the things within, like: "how do stars form" and "what happens when black holes collide". Max is deeply invested in modeling, analysis and prediction software, and I suspect did the bulk of the work in the paper. Henry Lin is a student who is focused on astrophysics. He gave an interesting TED talk (http://www.ted.com/speakers/henry_lin http://www.ted.com/speakers/henry_lin) a few years back about studying distant galaxy clusters. Henry is energetic and almost viscerally inspired by the beauty of science and mathematics, such a wonderful quality! His voice is definitely in the prose of the paper.
- deleted 10y ago[deleted]