4 ms·
>Great, so it's using historical data to make a prediction about what has already happened. That's ok though because it doesn't know it's acting on old data. R
by giarc 2y ago
>Great, so it's using historical data to make a prediction about what has already happened.
That's ok though because it doesn't know it's acting on old data. Researchers would give it data and ask for the 15 day forecast, the researchers then compare against that real world data. And as noted, "it surpassed the centres forecast in more than 97 percent" of those tests. This is all referred to as backtesting.
- IshKebab 2y agoIt's still not as good as actually new data. The individual model may not overfit the existing data, but the whole system of researchers trying lots of models, choosing the best ones, trying different hyperparameters etc. easily can.
- zactato 2y agoyeah, but weather patterns data isn't different now than 10 years ago, right? right? ...
- staunton 2y agoOverfitting is always bad by definition. The model learns meaningless noise that happens to help getting good results when applied to old data (be it due to trained weights, hyperparameters, whatever) but doesn't help at all on new data.
- gpm 2y agoIn principle they trained on data up to 2017, validated (tried different hyper parameters) on data from 2018, and published results on data from 2019...
- dagw 2y agoI wish these sort of papers would focus more on the 3 percent that it got wrong. Is it wrong by saying that a day would have a slight drizzle but it was actually sunny all day, or was it wrong by missing a catastrophic rain storm that devastated a region? I've worked on several projects trying to use AI to model various CFD calculations. We could trivially get 90+ percent accuracy on a bunch of metrics, the problem is that its almost always the really important and critical cases that end up being wrong in the AI model.
- gyrovagueGeist 2y agoYep, this is also the problem that self-driving cars have with the "accidents per mile" metric.
- genewitch 2y agoi was using scikit-learn or some other scaffolding/boilerplating similar software for doing predictions and i hate massaging data so i just generated primes and trained it on the prime series to like 300 or 1000, then had it go on trying to "guess" future primes or assert if a number was prime or not that it had seen before and hilariously it was completely wrong (like 12% hit rate or something). I complained online and i forget the exact scope of my lack of understanding but suffice to say i did not earn respect that day!