4 ms·
That's not really 100% true. A lot of the data this is trained on is ERA5, which appears to be highly dense in both time and space, but is assimilation data inf
by singhrac 2y ago
That's not really 100% true. A lot of the data this is trained on is ERA5, which appears to be highly dense in both time and space, but is assimilation data inferred from much sparser observations. I wouldn't say it's inaccurate, but I see pretty large deviations between assimilation datasets and private weather observations (I work on this problem).
The results are governed by physics up to some level, but we can't simulate at that fine a level, so there's some inherent aleatoric uncertainty (i.e. noise). And I would generally say that physics-simulation-ML is not moving as fast as say, inference on images or text. For example, if you see a picture of a car, there's very little inherent uncertainty on what the answer is. If you see the world simulator state there's a lot of uncertainty on what happens next.
That all being said, I think is basically the best model out there , and almost certainly the best open model. This is really the culmination of many years of effort getting data and software in place to run such a large scale training job. Very impressive!
- qeternity 2y ago> And I would generally say that physics-simulation-ML is not moving as fast as say, inference on images or text. I’m not sure that a blanket statement like this is a valid argument in an article that perhaps suggests the opposite is true.
- dehrmann 2y ago> if you see a picture of a car, there's very little inherent uncertainty on what the answer is Unless its a captcha.
- wenc 2y ago> For example, if you see a picture of a car, there's very little inherent uncertainty on what the answer is. If you see the world simulator state there's a lot of uncertainty on what happens next. I've been thinking about this a lot. Many ML people work with what is "closed-domain" data -- the data is essentially complete (image, sound, words, or any kind of embeddings) with no unmeasured variables, so the ML algorithm is essentially trying to learn a function that can predict this. Unfortunately a lot of "open-domain" data has tons of unmeasured variables that are contextual. Suppose I were to try to predict how a full a parking lot would be over the course of a week. You can gather lots and lots of data, but still never get to a near-perfect level of accuracy because the co-variates that drive how full a parking lot is (unexpected influencer effects on the demand, competitive forces that happen to shift one day, power outages in the other part of town, other irreducible randomness = "aleatoric uncertainty" in technical parlance) aren't in the data (or at least not completely). Fortunately this isn't a problem in real life because many effects cancel each other out, so we are able to arrive at a good-enough aggregate prediction. But "open-domain" ML problems will never achieve the kind of accuracy that "closed-domain" ML can achieve, even with tons of data. Closed domain ML can assume a degree of regularity that open domain ML can never assume.