4 ms·
Great, another open source tool purporting to solve time series analysis in an "automated way" that my manager will link me tomorrow and ask me to review (as an
by quantperson 5y ago
Great, another open source tool purporting to solve time series analysis in an "automated way" that my manager will link me tomorrow and ask me to review (as an aside, attempting automated statistics of any kind is incredibly dangerous and misguided, but especially so for time series).
Why should I use this over Darts[1] or just Statsmodels[2], if I need more lower level access and diagnostics? Both of these are far more established.
I dislike that Facebook Prophet was chosen as a benchmark; it's not a difficult benchmark to beat for the majority of time series use cases. It signifies to me that this project might targeting cargo cult data science. Prophet is not particularly good at non-daily timeseries and non-seasonal timeseries. The paper itself admits this[3]. Moreover, it's just a generalized additive model that incorporates holidays.
I don't intend to sound demeaning here, really. But I'm trying to understand what the point is. This doesn't look like someone's weekend project, but we already have plenty of established projects which tackle this effectively.
There are three major markets for time series work:
1. You're an analyst with a lot of domain knowledge who needs to analyze daily, seasonal data but you don't have a strong statistical or engineering background. This person should probably just choose Prophet (again, the developers of Prophet explicitly acknowledge that it's designed for scalable good enough models by non-stats people, not for the best model given the data).
2. You're a data scientist with a good statistical background and you need to produce forecasts. You can afford to dig into what the model is doing and select a model based on a series of diagnostics and knowledge about the data itself. This person should probably choose a more complete suite, like Darts. The important thing here is developing good models quickly while being able to do more than just press a button.
3. You're a data scientist (or statistician) which a very strong statistical background who needs to produce the best model they can for answering a specific question. This person is probably going to use R, Stan, Statsmodels or PyMC to come up with something bespoke. They may or may not need to systematize it, but they don't need to produce quantity over quality.
How does this thing improve the state of the art for any of these markets?
--
1 https://github.com/unit8co/darts https://github.com/unit8co/darts
2 https://www.statsmodels.org/stable/index.html https://www.statsmodels.org/stable/index.html
3 https://peerj.com/preprints/3190/ https://peerj.com/preprints/3190/
- fedegr 5y agoThe pipeline we have developed improves the state of the art in the markets you mention in the following aspects: 1. It is a fully automated end-to-end pipeline for forecast generation. The pipeline considers preprocessing such as missing value imputation, feature generation (static and dynamic), forecast generation, and also a module to validate forecasts on important time series competition data sets. 2. Users can deploy the pipeline in their cloud quickly. We use terraform (https://github.com/Nixtla/nixtla/tree/main/iac/terraform/aws https://github.com/Nixtla/nixtla/tree/main/iac/terraform/aws), so it is very simple to deploy the pipeline on AWS. We are working to release versions of terraform on other clouds such as Azure and Google Cloud. 3. Users can use their own models. Just create a fork of the repo and make the appropriate modifications to include any model the user wants to deploy. On our side, we are working to include Deep Learning models with the nixtlats library (https://github.com/nixtla/nixtlats/ https://github.com/nixtla/nixtlats/) that we also developed. About benchmarking using statistical models, we highly recommend using statsforecast (https://github.com/Nixtla/statsforecast https://github.com/Nixtla/statsforecast) that we created. It is designed to be highly efficient in fitting statistical models on millions of time series. More complex models can be built on the results to get a positive Forecast Value Added.