17 ms·
Machine Learning Can't Handle Long-Term Time-Series Data
- skunkworker 7y agoThis article seems out-of-date by 5 years or more even though it was published today, and I am unsure as to why. It calls out long short term memory but doesn't mention recent (last 5 years) improvements like Gated Recurrent Networks (GRUs) or Transformers (GPT-2, huggingface/transformers) which have shown significant improvements over the traditional LSTM model. These can handle time series data much better than older models could.
- joe_the_user 7y agoI think you need to give details and references to support a claim that these innovations make a fundamental difference. I don't doubt that the things you mention involve improvements but are these improvements doing better on the same benchmarks in the same fashion or a fundamental change. I read many claims that recent changes in deep learning represent the former.
- skunkworker 7y agoIt's to the point now where using a basic LSTM is almost discouraged compared to using a Transformer. GPT-2 wouldn't have been possible without these recent innovations. Here are some good guides on Transformers [1] and attention/ multi headed attention [2], as well as the paper that proposed the transformer model "Attention is all you need" [3]. GPT-2 heavily relies upon the advancements that transformers brought [4] [1] http://jalammar.github.io/illustrated-transformer/ http://jalammar.github.io/illustrated-transformer/ [2] https://towardsdatascience.com/attention-for-time-series-classification-and-forecasting-261723e0006d https://towardsdatascience.com/attention-for-time-series-cla... [3] Attention is all you need. https://arxiv.org/abs/1706.03762 https://arxiv.org/abs/1706.03762 [4] http://jalammar.github.io/illustrated-gpt2/ http://jalammar.github.io/illustrated-gpt2/
- curiousgal 7y agoExcept GPT-2 is irrelevant when it comes to time series specifically.
- Erlich_Bachman 7y agoGPT-2 handles natural language which is specifically a time-series set, a sequence of natural word tokens. It is especially relevant because success in NLP tasks requires tracking and learning long-term relationships - not just between words in a phrase or two, but between words and phrases in different paragraphs, in different parts of the text, so that a general coherence of the whole document is kept. This exactly what is done with long-term relationships.
- curiousgal 7y agoI know that NLP sequences are analogous to time series but GPT-2 can't be applied to say financial time series for example. That's what I meant.
- huffmsa 7y agoOr predicting the performance of a sports team over time. From what I've seen, It breaks down at the embedding layer, because while the teams "remain the same" in name / dictionary, their actual relative relationship to each other varies season to season / week to week.
- NegatioN 7y agoMy response here is purely intuition, since I have never worked much with time series. But wouldn't capturing that relationship require periodic retraining or other components to the network regardless? It may suggest that end-to-end training of a transformer is not suitable for these tasks, but that it might still capture the prediction of the long-scale time-series, if provided with extra data at each timestep in addition to the embeddings?
- laingc 7y agoI have no issue with you asking for a reference for this, but when I read your comment I did a double-take. To someone familiar with deep learning, this claim is completely self-evident. The article is indeed almost farcically out of date, and LSTMs haven’t been close to the state of the art for years now.
- hnews_account_1 7y agoAny links to heavy time series based machine learning algorithms? I'm in finance, and while I know how to establish and run a random forest or gradient boost regressor using standard libraries, I've never had a good handle on them.
- deleted 7y ago[deleted]
- cbsmith 7y agoMost everything Eamonn Keogh publishes: https://www.cs.ucr.edu/%7Eeamonn/selected_publications.htm https://www.cs.ucr.edu/%7Eeamonn/selected_publications.htm
- throwawaymath 7y agoLook into matrix profiles and associated algorithms.
- lern_too_spel 7y agoEven weirder, it calls LSTM a "newly invented variant" of RNN. LSTM is 20 years old.
- joe_the_user 7y agoThis claim seems plausible. The reason seems even simpler than the article. Deep learning requires lots of training data - that data naturally needs to more or less be "the same"; follow "the same" logic. A long enough time series is going to involve a change in the logic of the real world, a change that the network won't be trained for.
- dclowd9901 7y agoI think this is where the lack of “imagination” that computers currently cannot replicate becomes a huge problem that will set AI and ML back decades more.
- Erlich_Bachman 7y agoWhat are you even referring to? In which of the tasks that ML is currently applied to do they lack "imagination"? GPT-2 generators have plenty of imagination in generating new phrases and meanings... Alphastar has plenty of imagination of making new moves in SCII that even human players haven't come up with yet... Etc..
- s_Hogg 7y agoThey're designed to appear that way, it doesn't mean they actually are imagining anything. We need to be very careful about accidental anthropomorphism when it comes to models 0 seeing an analogy between the output in front of you and the human mind can blind us to the reality of what's going on. Which is often more prosaic, even if complicated.
- Erlich_Bachman 7y agoThat sounds like you are operating on your own definition which you can apply and shift however you want. This is not what objective science is about. How do you know real humans are not displaying "accidental anthropomorphism"?
- s_Hogg 7y ago
- cbsmith 7y agoI'm feeling like this is entire missing the whole world of Matrix Profiles and Time Series Chains...
- mjburgess 7y agoTime is only a symptom of what's missing: causation. ML operates with associative models of billions of parameters: trying to learn thermodynamics by parameterizing for every molecule in a billion images of them. Animals operate with causal models of a very small number of parameters: these models richly describe how an intervention on one variable causes another to change. These models cannot be inferred from association (hence the last 500 years of science). They require direct causal intervention in the environment to see how it changes (ie., real learning). And a rich background of historical learning to interpret new observation. You need to have lived a human life to guess what a pedestrian is going to do. If you overcome the relevant computational infinities to learn "strategy" you will still only do so in the narrow horizon of a highly regulated game where causation has been eliminated by construction (ie., the space of all possible moves over the total horizon of the game can be known in an instant). The state of all possible (past, current, future) configurations of a physical system cannot be computed -- it's an infinity computational statistics will never bridge. The solution to self-driving cars will be to try and gamify the roads: robotize people so that machines can understand them. This is already happening on the internet: our behaviour made more machine-like so it can be predicted. I'm sceptical real-world behaviour can be so-constrained.
- nextos 7y agoExactly, that's why I think we need to put logic and probability theory back into cutting edge ML. [1,2] are only early approaches that show potential directions to achieve this. Deep learning is very useful, but only one piece of the whole AGI puzzle. Furthermore, many AI problems will benefit the generality of being formulated as a probabilistic program synthesis problem [3]. In this framework, lots of program semantics (~formal methods) concepts like abstract interpretation [4,5] might become very useful. They allow to explore huge program spaces very quickly. Lastly, Pearl's do calculus [6] is a good starting point closely related to [1]. [1] http://probmods.org/ http://probmods.org/ [2] http://pyro.ai/ http://pyro.ai/ [3] https://web.mit.edu/cocosci/Papers/Science-2015-Lake-1332-8.pdf https://web.mit.edu/cocosci/Papers/Science-2015-Lake-1332-8.... [4] http://www.concrete-semantics.org/concrete-semantics.pdf http://www.concrete-semantics.org/concrete-semantics.pdf [5] http://adam.chlipala.net/frap/frap_book.pdf http://adam.chlipala.net/frap/frap_book.pdf [6] http://bayes.cs.ucla.edu/BOOK-2K/causality2-epilogue.pdf http://bayes.cs.ucla.edu/BOOK-2K/causality2-epilogue.pdf
- nabla9 7y agoThis is clever crackpottery type brainstorming from a smart person. The author has extremely grand set of connections he developed. It ties down Buddha, enlightenment, vipassana meditation, artificial intelligence, cybernetics, fractals and neuroscience. Nothing wrong with that, of course. Creative thinker should have these kind of crazy ideas and connections every day or at least once a week. I carry with me a notebook that is full of them. Most ideas die as 'premature babies'. They may be interesting to think and write down, but they are not fully developed and never fit together as well as you initially thought. Filtering and piking some of them to work with is important. Giving them up is the difference between crackpot and non-crackpot. Forcing grand connections prematurely makes this crackpottery type. Sharing the creative brainstorm in an essay that does not try make up connections would be easier to read.
- tgflynn 7y agoThe line between crackpottery and genius is a fine one. If blogs had existed in 1900 and a certain patent clerk had written a post on his ideas about clock synchronization somehow being related to electromagnetism, many would have dismissed him as a crackpot as well. The questions this article relates to are among the most profound and difficult that human reason has ever attempted to confront. I think one should be careful in labeling such ambitious speculation as crackpottery just because it doesn't yet amount to a fully coherent and formally testable theory.
- throwawaymath 7y agoOkay...but Einstein didn't write a blog. He published a paper for peer review. Some of his ideas remained controversial for decades, but he had a sufficiently mature, cogent and well-specified theory that he could at least work through hypotheses and publish results.
- tgflynn 7y agoHe published that paper in 1905. Is it absurd to think that if the Internet existed in his time he might have blogged about his preliminary ideas before publishing a formal paper ?
- deleted 7y ago[deleted]
- scottlocklin 7y agoMachine learning does just fine to extremely well at long term time series data; there are entire branches of machine learning dedicated to this. The fact that this imbecile never heard of these tools is why nobody should be reading his essay. Uber's engineers didn't do this for their human finder because; 1) Image recognition stuff isn't explicitly built to do this (though it easily could be jury rigged to do so) 2) Uber's engineers apparently never heard of the concept of "moving averages" and "threshholds" which would have worked just fine. "More precisely, today's machine learning (ML) systems cannot infer a fractal structure from time series data." -look at this idiot using words he doesn't understand. Muh fractals.
- justapassenger 7y ago“Imbecile”, “idiot”? Please refrain from personal insults as they add nothing to the discussion.
- vitamins 7y agoToo bad this post was flagged before I could read it. It really deserved a Scott Locklin rant. It takes a special kind of disrespect to attack the field of machine learning from the viewpoint of a roleplaying techno mage.
- scottlocklin 7y ago"Imbecile" and "idiot" were measured and reasonable adjectives for the gibbering nonsense published above. The drooling lackwit who wrote this should be tarred and feathered for such frippery and nonsense. As I said above; machine learning does just fine to extremely well at long term time series data; there are entire branches of machine learning dedicated to this.
- longemen3000 7y agoI remembered that neural differential equations are better suited to represent time series data, I saw them being used a lot in pharmacological processes, any additional idea or insight related to this?
- NPMaxwell 7y agoThe article I would like to read is what the challenges are to including a few prior states in navigation. I'm amazed that, when I drive over or under a bridge, my online mapping software changes instructions as if my car were able to levitate 20 feet onto the roadway above or below, even when that roadway is a highway without exit or entrance within a mile.