5 ms·
> In climate science we do a lot of downscaling. We take temperature and precipitation readings from a coarse scale Global Climate Model grid and map them to a
by roter 6y ago
> In climate science we do a lot of downscaling. We take temperature and precipitation readings from a coarse scale Global Climate Model grid and map them to a fine scale local grid. Let’s say the global grid is 50x25 and the local grid is 1000x500. For each grid cell in the local grid, we want to know to which grid cell in the global grid it corresponds.
Climate scientist here. Colour me a little confused. We have a GCM which is (typically) on a structured grid, i.e. with latitudes & longitudes on a pixel-like grid. Finding the GCM (coarse) grid cell for a given latitude/longitude of the downscaled (fine) grid is an integer operation, i.e. invert lat = lat0 + i x delta_lat and lon = lon0 + j x delta_lon. Must be missing something here. Even if there is a strange projection, you can map it such that the deltas are uniform.
- Ultimatt 6y agoYou're not wrong. This entire article is a classic example of Computer Science failure. It's someone who knows program optimisation in the abstract and nothing about the actual target problem from a domain knowledge perspective. They went about optimising bad code as in quality of thought level, rather than actually getting to grips with the true requirement first. Though perhaps a computer science win, but a software engineer fail.
- mlyle 6y agoIn a lot of ways, the original implementation makes a lot more sense than this "optimized" solution. Something like the original brute force, compare all the points' distances variant is how I'd be inclined to write one of my unit tests for the grid alignment routines. Trying a few dozen points on a few dozen grid variants is a way to make sure I've not completely screwed up math-- whether off by one, rounding problems, etc, while applying sanity checks on the overall grid code base. This binary search "optimized" abomination is... more difficult to read and reason about than either than the integer math or "brute-force" solutions, and slower to boot. All it has going for it is that it's not the absolute slowest choice. It may not be integer math, because we've not been told if the grid is regular... but it certainly is still easier than this.
- projectdelphai 6y agoI mean maybe but if you told me you could take my code that runs in less than a second and make it more readable but it would take 30 minutes to run instead . . . hell no I'm not taking you up on your offer. I'll just comment and document the existing solution better and call it a day.
- mlyle 6y agoYou're missing the point. The solution here is faster than the absolute naive solution of finding all distances, but it is both slower and less readable than integer math grid conversion.. I could see why you would make the choice of the naive, slow solution. I could see why you would choose the fast, integer math grid conversion one. But I don't see why you would choose something slower than the optimum solution that is also much, much less readable. That is, the naive solution and the integer math solution are Pareto-optimal choices... and this abomination ain't.
- dang 6y agoPlease don't post supercilious dismissals of other people or their work, even if they made a mistake and/or were supercilious in their own right. It just contributes to making this a nasty place. The differences between this comment and the GP comment are significant. The GP included specific information, where this comment is just a putdown, and the GP allowed for the possibility that there is missing information, i.e. that the "author is an idiot" interpretation is not the only possible one. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- andi999 6y agoI usually agree, but in this case maybe the author needs to hear exactly that. At least the author might the want to stop boasting about these things (and save his future career etc).
- dang 6y agoI appreciate the sentiment, but that is probably better reserved for personal interactions than internet pile-ons. There are lots of factors we don't know. Also, if we're to be honest with ourselves, internet pile-ons don't arise out of compassionate concern for the other. That's a justification that gets tacked on. It's moot anyhow, because the main reason not to have threads go that way is that it does bad things to the community. Each time it happens, we deepen the pathways towards making HN nastier and more toxic, and since those are the default paths already, we need to consciously cultivate the opposite.
- kolbe 6y agoThis was written in 2015. What's worse logic: his code or HN thinking that shit talking him five years later in a forum we don't even know if he sees is constructive?
- dang 6y agoObviously I agree, but please don't add to the problem by dumping more nastiness into the thread. Note this guideline also: "Please don't sneer, including at the rest of the community." It's worth remembering that HN is a statistical cloud of posts, not a person, and so can't think anything. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- danielvf 6y agoI've upvoted this. Even the "optimized" solution here is so horrible that I almost thought the article was satire when I first read it. Maybe it is. Even if we ignore the fact that it's grid to grid conversion, (which should just let you math the answer, without any searching) the fact that all the data is sorted means you can do the equivalent of a merge sort and only have to look at a two point for each step. And how on earth can you take .117 seconds to line up a half a megapixel worth of grid points, even with their "fast" algorithm. They are using the R language which should be fast, right? Are they doing something expensive with it? Or perhaps there is an insane amount of data elsewhere making these lookups super expensive?
- deleted 6y ago[deleted]
- dang 6y agoPlease post your improved and correct perspectives without name-calling. Doing otherwise makes the community worse, even if you're right. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- danielvf 6y agoMy apologies. While I don't see the name-calling towards an individual, my previous post did mention possible bad faith, and I would not been glad to have my post written about my code. I also slightly misunderstood what the author's code was doing. I can see how it is not conducive to helpful discussion. The author optimized their code in a way that was fast enough, and fast enough to massively change the entire business process around their code. It's very commendable. Further optimizations may have not been worth it.
- dang 6y agoAppreciated!
- forgotpwd16 6y agoR in general and especially loops are slow; slower than even Python. Making a function and applying it on a vector is a simple way to boost performance. Also recursion seems[1] to be faster than iteration. Seeing the code in article, it can be that a part in that speedup is because R is used and not only of a better algorithm. [1]: https://predictivehacks.com/a-comparison-between-iteration-reduce-recursion-memoization-in-r/ https://predictivehacks.com/a-comparison-between-iteration-r...
- Marazan 6y agoYeah, I genuinely don't get the problem statement. As you say as stated the solution to the problem as stated is a single line piece of basic math. I can o lump resume the problem is badly described to miss some subtly. I hope.
- necovek 6y agoI am not a climate specialist, but I was confused about the same thing: I would fully expect for there to be a bijective mapping between a global grid and a local grid. I mean, lat/lon is a 2d space. My other comment was going to be that using binary search is not that much computer science. I mean, it's elementary programming knowledge, you don't need a degree to know that.