8 ms·
Time Series Forecasting vs Regression: An informal guide
- taeric 3y agoI always feel this is too close to stochastic versus random. There is a lot of text that pushes an idea that regression is used to understand how well a model fits relationships between variables. But, I start to have major doubts when people push the idea that regression models are not also predictive models.
- VHRanger 3y agoI mean, regression models (GLMs in general) are interpretable. If you are using them to extrapolate (eg. Prediction) that should help you gauge how resilient you expect the model to be in prediction. Obviously, for ARIMA the AR and MA parameters aren't very informative. I use SARIMAX a decent amount, nonetheless.
- mr_toad 3y agoTaken individually the AR and MA terms are an average change and an average difference respectively. When you combine different AR and MA terms it does become less explicable.
- Fomite 3y agoOne of the major time series prediction models my field uses is indeed a regression model.
- rdli 3y ago(Author here) One of the things that confused me is that regression models can be predictive, just like time series forecasting — they just do so in a different way. I tried to make this clear in the article (or maybe I’m not understanding what you’re saying). In a regression model, you’re predicting target variables from feature variables. In a time series, you’re predicting the same variable from its past behavior. This is a subtle but crucial difference. (And then you can do time series with covariates, which combines the two.)
- civilized 3y agoMany of the most important time series prediction models are called "autoregressive", meaning they are regression models predicting the target from (prior values of) itself. This suggests that statisticians don't really share the view that these domains are distinct, or that regression models should only predict with different variables from the target.
- hackerlight 3y agoRight, AR(n) is a regression model, as are models which take only exogenous variables. My question is this. According to definitions, can the latter (f(X_t) = y_t) be a time series model if each row of data is a time step? It doesn't have any autoregressive terms in X, so I don't know if it categorically is a time-series model. Not that this question even matters, it's purely a taxonomy/terminology question.
- nerdponx 3y agoYes, it is. A time series model is any model where the data varies over time; that is, a time series model is any model of time series data. And timeseries data is broadly anything where the data for a single thing/entity varies over time. There are no strict definitions here, just common conventions.
- hackerlight 3y agoOkay. And we can also say that there's some time series models that aren't regression models, right? For example, Kalman Filter is a "model" of a time series but isn't a regression.
- nerdponx 3y agoCorrect. Although the term "regression" is a misnomer anyway, and often when people say "regression" they mean "linear model". And by "linear model", we mean specifically a model in which outputs/predictions are some fixed linear combination of the input. It is however possible to interpret the Kalman filter as a kind of dynamic regression model. Check out here if you want a good math workout on that topic: https://stats.stackexchange.com/q/330696 https://stats.stackexchange.com/q/330696 (Another somewhat distinct meaning of the term "regression" is any model with a "continuous" outcome variable. This is usually in contrast to "classification", which is any model that has a "categorical" or discrete outcome variable.)
- richrichie 3y ago[dead]
- vavooom 3y agoThe author makes a call out to the online book Forecasting: Principles and Practice which is a great reference when conducting time series analyses. https://otexts.com/fpp3/ https://otexts.com/fpp3/
- wardedVibe 3y agoI found it to be way too underspecific; more of cursory overview for a undergraduate seeing it for the first time in a business school than someone interested in digging into time series forecasting in depth. Don't have a better recommendation though, unfortunately.
- foundart 3y ago> At the end of each chapter we provide a list of “further reading”. In general, these lists comprise suggested textbooks that provide a more advanced or detailed treatment of the subject. Where there is no suitable textbook, we suggest journal articles that provide more information.
- disgruntledphd2 3y agoBox and Jenkins is probably the best book, and the older editions are much cheaper (and about the same quality) as the newer ones.
- vmfunction 3y ago> This process is typically called “feature engineering”, and is part art and part science. Choices on including or excluding certain variables, and how they are translated into numerical parameters, can significantly impact the model’s performance. According to this article, to make good predictive/regression model, we need a good artist and a good engineer!
- rdli 3y ago(Author here). Or a lot of trial & error :).
- Vaslo 3y agoMy job is now primarily Time Series Forecasting, and we’ve spent so much time improving our feature selection and engineering. When I started I thought “run correlations against target variables, find the best bunch and as long as we can explain them and their relation to the target we are good” I was wrong.
- eyegor 3y agoWait I still do this, what are your secrets?
- bigger_cheese 3y agoI work mostly with regressions and often it is almost more informative when something you expected to be a significant term isn't. Can help track down interesting behavior. More recently Machine learning has really enhanced what you can do with regression. For example multivariate regressions when there are non-linear (or partially linear) relationships between feature and target variables. For example recent regression problem involved a chemical reaction. It was suspected that a particular feature above a threshold began to display non linear behavior but it was difficult to pinpoint exactly where it began departing from linearity. ML was very helpful analyzing this. Other than regressions and timeseries forecasting I think it's worth knowing about K-means clustering and PCA (Principal Component Analysis)/ PLS (Projection to latent structures) as well. I've found PCA to be pretty unknown but very useful I've had success using it in the past and found it useful to explain the relationship not just between the data features and the target variable but also how the features relate to each other.
- Jakesbeb 3y ago[dead]
- riedel 3y agoTo me time is just one dimension. What is described is just the difference between interpolation and extrapolation. In terms of forecasting state of the art are weather models like graphcast or panguweather. I guess arima won't be much of help in those high dimensional cases. If you consider the univariate case the trick to outperform arima I guess is to detect the context from the time window before to make better contextual predictions: this is much like a regression on a hidden variable.
- mr_toad 3y ago> What is described is just the difference between interpolation and extrapolation. Extrapolation (predicting an unknown future), and interpolation (estimating unknown present/past) are not really that different.
- riedel 3y agoI would mostly agree, this why time series imputation or cleaning is often not that different from time series forecasting (or rather often 'nowcasting'). You would only want to be careful what kind of validation you chose to test the generalisability of the approach. If you take however the example of the weather as an extreme example of time series forecasts, downsizing eddies or forecasting them in a navierstokes surrogate, this can require some different approaches.
- LuciBb 3y ago[flagged]
- arisAlexis 3y agoWith risk of sounding bad: I can ask chatgpt to summarize this for me without any human writing an article since there is ample knowledge already available. What is the future of these kind of articles ?
- iamgopal 3y agoSoon AI will be feeding itself, and gradually will degrade in performance, that time, human generated better content to feed to AI will be in demand.
- blitzar 3y agoPerhaps, one day in the not too distant future there will be a revered old master in a village somewhere, who people travel from miles away to watch as they slowly and carefully write listicles the old way.
- aeonik 3y agoThis will be true until the models can perform science, and update their weights according to their tests.
- enoch2090 3y agoTotally makes sense - I'm already getting used to "ChatGPTing" instead of Googling these questions. I guess the future of these articles is that there are always fields that LLMs hallucinate at, fields that are not very common (explain the products of a small brand, pros &cons of their different models).
- nerdponx 3y agoBecause ChatGPT is somewhere between "subtly" and "totally" wrong on most topics related to statistics and machine learning that I've tested it with. Maybe GPT 4 is better than 3.5, but I don't really trust it on technical subjects. The quality is significantly lower than a good article written by a competent human, but maybe on par with or slightly better than a trashy article written by a content farm. The advantage of the chat interface is that you can ask it clarifying questions. The real benefit of generative AI would be something like Copilot that you can interrogate for clarification as you are working through an article written by another human. That, and the other problem of AI being trained on AI until nobody knows anything anymore.
- t_mann 3y agoARIMA models, seasonal adjustments,... this is still largely based on the Box-Jenkins method (developed in the 70's!). I feel like this stuff has been taught the same way for decades now (maybe similar to undergraduate classical mechanics or other topics that are considered 'solved'). Is this really still the state of the art? Time series analysis seems oddly close to machine learning, which seems to move at break-neck speed all the time, yet it feels completely stuck in time. Can someone unravel that paradox for me?
- lukego 3y agoMaybe of interest: https://github.com/probsys/AutoGP.jl https://github.com/probsys/AutoGP.jl
- disgruntledphd2 3y agoTS models are constrained by the realities of the world they exist in. While you can chase benchmarks in lots of ML problems, forecasting is something that's used by basically every large business with huge consequences to getting it wrong or right. Therefore, people stick with relatively performant & interpretable methods such as ARIMA and friends. Additionally, most TS problems are relatively data constrained (your company/product has only existed for so long) so methods that are sample efficient (which most "modern" ML methods are not) are much more useful. Also, time series/forecasting is a ghetto ;)
- boppo1 3y ago>time series/forecasting is a ghetto ;) What do you mean?
- disgruntledphd2 3y agoIt's very very disconnected from overall modelling, with its own approaches and culture. There's little to no cross pollination from other areas of modelling.
- bradstewart 3y agoNew things are coming out, Meta's "prophet" for one. I've been pretty impressed with it in a "just throw data and don't even think about parameters" sense. But the fact is, ARIMA models work. So people keep using them. And you can see what they're doing, and understand why, and how to tune them.