21 ms·
Ask HN: Data Scientists, what libraries do you use for timeseries forecasting?
I default to Prophet (formerly FBProphet) for my work [which is business-y timeseries data], curious what others are doing.
- jstx1 4y agoProphet, statsmodels, tf.keras for RNNs.
- speedgoose 4y agoI was told to start with XGBoost. Is Prophet better?
- unlikelymordant 4y agoXgboost is a classifier for tabular data, prophet is for time series prediction. They are different use cases, though you can likely massage xgboost to do time series prediction if you really wanted to. So the question of which is better is "it depends"
- istinetz 4y agoI'm not even sure how you can get it to work on time series...
- bayan1234 4y agoCreate features for day of week, day of year, month of year, lagged values of y, lagged values of y for each period (eg: 1, 2, 3 weeks and years ago etc). You then predict forward 1 time step at a time.
- isoprophlex 4y agoCheck out my top level comment in this thread for a (hopefully clear) example. Sometimes you can rephrase a time series problem into boring classical regression. It can make the implementation and maintainability of a codebase better (IMHO), without sacrificing predictive power.
- salty_biscuits 4y agoTime delay embedding (i.e. the observed value at different time lags as a feature) is the usual trick to turn time series data into a tabular form for this sort of regression.
- ollysb 4y agoYou can use XGBoost for both classification and regression. See https://xgboost.readthedocs.io/en/stable/python/python_api.html#xgboost.XGBRegressor https://xgboost.readthedocs.io/en/stable/python/python_api.h...
- ssequeira 4y agoI'll often use tensorflow probability's time series package.
- brylie 4y agoProphet https://facebook.github.io/prophet/ https://facebook.github.io/prophet/
- thegginthesky 4y agoSktime is the best toolkit for time series out there. It provides a sklesrn like API for many models and modules for validating, metrics for evaluating and all that sklearm jazz. Besides that, I also like statsmodels as the docs are pretty good.
- fzliu 4y agoPyTorch for recurrent nets (TensorFlow would work too).
- qsort 4y agoIt really depends on what is the goal. Get some forecasts quick => FB Prophet. It's not as good as they'd have you believe, but it's fast and analysts can play with it to some extent. Outlier detection => Hand-rolled C++ ETS framework. Multilevel predictions and/or more complex tasks => That's where neural models start to have the edge, but at that point it's a costly project. I like simpler stuff to start, moving to the big guns if/when it's needed.
- Imanari 4y agoFor feature engineering check out tsfresh and sktime, especially the minirocket algorithm. https://tsfresh.readthedocs.io/en/latest/ https://tsfresh.readthedocs.io/en/latest/ https://www.sktime.org/en/v0.8.2/api_reference/auto_generated/sktime.transformations.panel.rocket.MiniRocket.html https://www.sktime.org/en/v0.8.2/api_reference/auto_generate...
- isoprophlex 4y agoI've had someone in a team implement feature engineering using tsfresh. It lead to a malignantly under-performing, complicated heap of spaghetti that was a nightmare to get into production. Weird API, slow code, little added value over simple features found in a day of manual exploration. Person doing the implementation wasn't a rock star coder so we couldn't fix the performance and complexity issues in time; it was left it out of the releas. Maybe with more expertise, tsfresh can add value. The experience was pretty off-putting for me, all in all. Maybe others have different experiences?
- esquire_900 4y agoI've tried it a number of times, and had a similar experience. The whole stack is orders of magnitudes slower to compute compared to "simple" features (i.e. rolling averages), without showing real predictive improvements.
- isoprophlex 4y agoThanks for the elaboration. Good to know someone had a similar experience.
- Imanari 4y agoActually, I had a similar off-putting experience with tsfresh. I just thought it was due to me not understanding how to use it properly. The API is indeed pretty weird.
- ollysb 4y agoDarts gives you a lot of options, including newer deep learning approaches like NBEATS and NHiTS. https://unit8co.github.io/darts/ https://unit8co.github.io/darts/
- d4rkp4ttern 4y ago+1 for DARTS, a great PyTorch based TS forecasting library that implements several state of the art algos. Devs are responsive on gitter. A similar pkg is PyTorch-forecasting but overall I prefer DARTS, mainly because of devs responsiveness and better docs
- curiousgal 4y agoTime series analysis is where R shines compared to Python.
- isoprophlex 4y agoCan you substantiate your comment? What would you like to see in a python tool that R does uniquely well?
- RamblingCTO 4y agoNot parent, but easy: https://cran.r-project.org/web/packages/forecast/index.html https://cran.r-project.org/web/packages/forecast/index.html /e: I'm a bit out of this game for 2-3 years now, but Python had nothing comparable. Prophet is suboptimal at best. Some things are implemented here or there, but it's all over the place. So I agree that R really shines here.
- isoprophlex 4y agoWonderful, I'm gonna go check this out in great depth!
- RamblingCTO 4y agoI can also recommend the book and blog by Rob Hyndman! https://robjhyndman.com/hyndsight/ https://robjhyndman.com/hyndsight/
- disgruntledphd2 4y agoTo be fair, Darts looks pretty good relative to forecast: https://github.com/unit8co/darts https://github.com/unit8co/darts I would generally prefer R for this kind of stuff as the experts generally write the code, but Darts seems OK and is well-tested, at the very least (haven't had a chance to use it in anger yet).
- hrzn 4y ago> Python had nothing comparable This is in part why we built Darts. Now I think we can say the situation is quite different. Darts offers many things offered by the R forecast package, and then some (for instance the ability to train ML models on large datasets made of multiple potentially high-dimensional series).
- bayan1234 4y agoForecast package in R is quite useful. Even if you don’t use R, this book by Rob Hyndman is very approachable and easy to follow. https://otexts.com/fpp2/ https://otexts.com/fpp2/
- qsort 4y ago+1 for the book, an excellent reference. As someone who's primarily a developer, it's been extremely useful as a study guide. Also, there's a new version; s/2/3 :)
- nerdponx 4y agoIt's useful as a study guide and reference even for someone who ostensibly learned all this stuff in school. It's a tremendously good book, and it's even more impressive that it's free to read online in a high-quality HTML document.
- hugh-avherald 4y ago`fable` is the successor to `forecast`, according to Hyndman, though the latter is still maintained.
- dxbydt 4y agoThis is a good book and forecast is a very good R package, imo.
- isoprophlex 4y agoCan you reframe the problem to suit a more classical approach - regression using xgboost or lgbm? If so, go for that! As an example, imagine you want to calculate only a single sample into the future. Say furthermore that you have six input timeseries sampled hourly, and you don't expect meaningful correlation beyond 48h old samples. You create 6x48 input features, take the single target value that you want to predict as output, and feed this into your run of the mill gradient boosted tree. The above gives you a less complex approach than reaching for bespoke time-series stuff; I've personally have had success doing something like this. If your regressor does not support multiple outputs, you can always wrap it in sklearns MultiOutputRegressor (or optionally RegressorChain; check it out). This is useful if, in the above example, you are not looking to predict only the next sample, but maybe the next 12 samples. https://scikit-learn.org/stable/modules/generated/sklearn.multioutput.MultiOutputRegressor.html https://scikit-learn.org/stable/modules/generated/sklearn.mu...
- em500 4y agofbprophet is mostly just regression though, with features for trend, yearly and weekly periodicity (smoothed a bit using trigonometric regressors), and holiday features. The only non-standard linear regression part is that it includes a flexible piecewise linear trend, with regularization to select where the trend is allowed to change. Once the change points are selected, it's literally linear regression, that you could fit with anything that can handle regression (statsmodels, sklearn, xgboost, keras, tf, scipy or even just plain numpy). So you could roll your own using lots of different libraries and do the feature engineering, but for plain jane business time series, fbprophet usually works pretty decently with minimum effort. Because most of the predictable variation in business time series is are patterns in human behavior, which are mostly regulated by the weekly and yearly cycle, interrupted a few times per year by holidays. (And the 24h cycle if you forecast intra-day). That said, fbprophet is somewhat tuned for a few years of regular data (ideally not interrupted by pandemics), sampled at the daily frequency. If your data is different (e.g. weekly or intra-day) it can start to break in unexpected ways, and it becomes worthwhile to dive into the model and customize or roll your own.
- nerdponx 4y ago
- siilats 4y agoEasiest is to use cvxpy with your own objective function. You can easily add seasonality regularization etc. other things are too much black box. Also pivot tables. They are free now in online version of google sheets and online excel. Set the time as a row field and it will automatically aggregate. Or if you want irregular spacing you can group by 100 samples.
- nerdponx 4y agoI don't know why you were being downvoted for this, it's certainly a bit idiosyncratic but there's also nothing wrong with it. Edit: I do question why you think this is "easiest". Do you have some really unusual/specific modeling needs?
- Fiahil 4y agoXGboost, LGBM, pmdarima, stanpy (for bayesian modelling). Plus a few others. Don't ask me what they do with all of these, I'm just the guy who make sure the forecast keeps being reproducible.
- chirau 4y agoWhat does that make you titlewise? Data Engineer? ML Engineer?
- Fiahil 4y agoA Software Engineer. I'm just specialised a bit in DevOps, Data Engineering, and (beware buzzword) MLOps
- moralestapia 4y agoWhat's MLops? Is it what I imagine? Maintaining repos of training/data, APIfying the pipeline, deploying an ML processing pipeline with CI/CD, etc?
- nerdponx 4y agoPretty much this. "ML engineering" has come to refer to the somewhat specialized task of implementing models and algorithms, and "ML ops" has come to refer to all of the other stuff that you just mentioned.
- navbaker 4y agoYou got it. It's unbelievably difficult to get model devs out of the mindset of training on their own VM, saving model outputs and metrics dumps to arcanely named file shares, etc. Once you can convince them that using stuff like workflow pipelining tools and centralized model repo servers isn't going to impede their creative process and that it prevents the mad scramble to find artifacts when there's turnover on the team, things become much more efficient.
- Fiahil 4y ago
- crimsoneer 4y agoFor time series, classical methods (ARIMA etc) still continue to perform very well for most problems.
- ackbar03 4y agoI don't think they necessarily perform well, its just that time-series forecasting is notoriously hard and we haven't seemed to progress that far from ARIMA type models (to the best of my knowledge)
- nceasy 4y agoAnyone ever tried pycaret timeseries?
- angrycontrarian 4y agoState of the art is 1D convnets, bleeding edge is transformers.
- jll29 4y agoFormer Reuters Research Director here. When modeling time series, you will want a model that is sensitive both to short term and longer term movements. In other words, a Long Term Short Term Memory (LSTM). Sepp Hochreiter invented this concept in his Master's thesis supervised by Jürgen Schmidhuber in Munich in the 1990s; today, it's the most-cited type of neural network. Here are papers describing it: https://people.idsia.ch/~juergen/rnn.html https://people.idsia.ch/~juergen/rnn.html In Python, you can use TensorFlow's LSTMCell class: https://www.datacamp.com/tutorial/lstm-python-stock-market https://www.datacamp.com/tutorial/lstm-python-stock-market
- sabertoothed 4y agoI don't think LSTMs are state-of-the-art in the domain anymore.
- Tarq0n 4y agoLSTMs never did beat well tuned statistical models anyway. Neural nets are only taking over forecasting in the last few years.
- isoprophlex 4y agoLSTMs have been going the way of the dinosaurs since 2018. If you really need a complex neural network (over 1D convolution approaches), transformers are the current SOTA. Example implementation in "temporal fusion": https://pytorch-forecasting.readthedocs.io/en/stable/tutorials/stallion.html https://pytorch-forecasting.readthedocs.io/en/stable/tutoria... Mind you, in practice I've found these DL approaches overkill for simple problems of the "trends + cyclics + noise" kind.
- nerdponx 4y agoMy impression is that these kinds of models need a lot of data to train properly. I made a comment elsewhere in this thread, musing that tree ensemble models could do the same job, as a kind of low resolution quantized approximation. If you have experience in this area of research, I'd love to hear your thoughts on that.
- d4rti 4y agoStuff I've used: - Prophet - seems to be the current 'standard' choice - ARIMA - Classical choice - Exponential Moving Average - dead simple to implement, works well for stuff that's a time series but not very seasonal - Kalman/Statespace model - used by Splunk's predict[1] command (pretty sure I always used LLP5) I did some anomaly detection work, in business transactions, and found the best way was to create a sort of ensemble model, where we applied all the models, and kept any anomalies, then used simple rules to only alert on 'interesting' anomalies, like: - 2-3 anomalies in a row - high deviation from expected - multiple models all detected anomaly To improve signal vs noise. [1] :https://docs.splunk.com/Documentation/Splunk/9.0.1/SearchReference/Predict https://docs.splunk.com/Documentation/Splunk/9.0.1/SearchRef...
- NeutralForest 4y agoI don't know what you've been using Prophet for but I found it to be very brittle.
- em500 4y agoYup, as per my other comment (https://news.ycombinator.com/context?id=33448802 https://news.ycombinator.com/context?id=33448802), fbprophet is largely tuned for a few years of somewhat regular business data sampled daily (e.g. sales per day). Outside its comfort zone (e.g. if you have monthly, or hourly/minutely data, or step changes / level shifts) it can fall apart pretty quickly. But its comfort zone happens to be very popular in business settings. To be fair, it's pretty hard to create generic models that are can robustly handle any random time series.
- NeutralForest 4y agoFair enough, when I was working with monthly sales, it was pretty bad.
- kqr 4y ago> - 2-3 anomalies in a row > - high deviation from expected > - multiple models all detected anomaly This is basically what statistical process control charts do for you. If you haven't learned about it already, I can recommend looking it up!
- oneoff786 4y agoEverything is either light gbm or an early exploratory experiment.
- nl 4y agoI really like FB's Prophet: https://facebook.github.io/prophet/ https://facebook.github.io/prophet/
- nerdponx 4y agoLots of people have already made good library recommendations, so I will make a non-recommendation for all the data science students out there: stop thinking about libraries, and start thinking about models. "What library do I use?" is the wrong question. "What model do I use?" is the right question. Libraries are just part of the process of answering that question. That said, high quality implementations of interesting times series models seem hard to come by, so it's still a legitimate question to ask about libraries. but consider the goal of asking about libraries: you want to find high-quality implementations of useful models, not a magic black box that you can crank data through.
- krageon 4y agoI wouldn't mind a magic box, honestly.
- Pinegulf 4y agoI can't blame you. However, if I don't understand what the box actually does, potential for huge -oooops- is too high.
- arminiusreturns 4y agoCould you suggest some good resources for learning about timeseries models?
- wheelinsupial 4y agohttps://stats.stackexchange.com/questions/20514/books-for-self-studying-time-series-analysis https://stats.stackexchange.com/questions/20514/books-for-se...
- hrzn 4y agoI would recommend Darts in Python [1]. It's easy to use (think fit()/predict()) and includes * Statistical models (ETS, (V)ARIMA(X), etc) * ML models (sklearn models, LGBM, etc) * Many recent deep learning models (N-BEATS, TFT, etc) * Seamlessly works on multi-dimensional series * Models can be trained on multiple series * Several models support taking in external data (covariates), known either in the past only, or also in the future * Many models offer rich support for probabilistic forecasts * Model evaluation is easy: Darts has many metrics, offers backtest etc * Deep learning scales to large datasets, using GPUs, TPUs, etc * You can do reconciliation of forecasts at different hierarchical levels * There's even now an explainability module for some of the models - showing you what matters for computing the forecasts * (coming soon): an anomaly detection module :) * (also, it even include FB Prophet if you really want to use it) Warning: I'm probably biased because I'm Darts creator. [1] https://github.com/unit8co/darts https://github.com/unit8co/darts
- dxbydt 4y agoPeter Cotton has atleast a dozen very credible studies/results on prophet vs other timeseries libraries. Before committing to prophet, please check out a few of these (all over linkedin). His tone is acerbic because he believes prophet is suboptimal & makes poor forecasts compared to the other contenders. That said, you can ignore the tone, just download the packages & test out the scenarios for yourself. I personally will not use prophet. Like most stat tools in the python ecosystem, it is super easy to deploy & code up, but often inaccurate if you actually care about the results. ofcourse, if its some sales prediction forecast where everything’s pretty much made up & data is sparse/unverifiable, then prophet ftw.
- nerdponx 4y agoI think acerbity is warranted to some extent. We are data scientists. We get paid the big bucks because we have big brains and have the skills and training to use those big brains in order to reason about the work we are doing. Data science has become so easy and accessible nowadays that basically anyone who can write code can fit and use models. That's a great thing in general, but it means that those of us who do this for a living really should hold ourselves to higher standards. Even if you are thoroughly mediocre at your job (like me) and aren't smart enough to come up with something like Prophet on your own, you absolutely do need the ability to reason about the models that you use, and to evaluate their performance correctly.
- boringg 4y agoIll add the better question is what libraries do you have that cleans up all the data before adding the model aka the real heavy lifting ;)
- plutonic 4y agoAs a few other people have mentioned, I find R to be the easiest tool for this job, specifically the forecast package [0]. I had to use this package for an applied econometrics course in college a few years ago, and I have been using it ever since. I find the syntax to be more straightforward than comparable libraries in Python. I also assume that this library (and other libraries in R) offer higher quality models and results than their counterparts in Python, but this is just an assumption. [0] https://github.com/robjhyndman/forecast https://github.com/robjhyndman/forecast
- venk12 4y agoI once built a forecasting framework for a unicorn startup. Revenue and Pipeline predictability was the key as the company was going through the IPO phase. There were three approaches I took and 'ensemble'd them to predict the revenue and pipeline. 1. Time series based forecast based on revenue (the one OP is referring to). All the statistical time-series models come here. I primarily used H2O.ai for this. 2. Conversion based revenue forecast (input -> pipeline, output -> revenue). This proved to be quite tricky as there was a time lag between pipeline creation and revenue conversion 3. Delphi-method: Got the sales/pre-sales folks on-ground to predict a bottom-up number and used that as a forecast. Finally, I combined them by applying weightages to the above approaches - based on how accurate they were on the test dataset. IMHO, Like many of them have pointed out - the model/assumptions are more important than the library. The job of a data scientist is to make the prediction as reliable and explainable as possible.
- juanplusjuan 4y agoUmm this is exactly what I’m trying to build now. Any chance you’d want to connect to talk about it? juan [at] sayprimer.com
- venk12 4y agosure we can. pls drop a note to venkatasubramanian1209@gmail.com
- boredemployee 4y agoWhat about adding external variables like weather/rain to check the impact on sales? What you guys recommend?
- Maro 4y agoMany libraries have this, eg. Prophet. The technical term is 'external regressor'.
- tfehring 4y agoFor cases that Prophet doesn't cover I recommend bsts [0], which is much more flexible and powerful. Anything too complicated for bsts, I'll typically implement in Stan. [0] https://cran.r-project.org/web/packages/bsts/bsts.pdf https://cran.r-project.org/web/packages/bsts/bsts.pdf
- mharig 4y agoI like statsmodels. So far it has all methods I need, and itvis very well documented. But I am just fiddling a little bit with my 'weather station'. No bleeding edge here.
- pruthvishetty 4y agoSktime
- helloded 4y ago[flagged]
- hellded 4y ago[dead]
- helloded 4y ago[dead]
- helloded 4y ago[dead]
- helloded 4y ago[dead]
- dmfolgado 4y agoFor feature extraction check out tsfel: https://github.com/fraunhoferportugal/tsfel https://github.com/fraunhoferportugal/tsfel
- heloitsme22 4y agoDepends on the day. Sometimes may be good sometimes may be shit