2 ms·
More generally, gdal is a raster I/O library. S2 is a point only system. It's not meant to store or work with raster data in any way. (Sure, you can represent r
by jofer 5y ago
More generally, gdal is a raster I/O library. S2 is a point only system. It's not meant to store or work with raster data in any way. (Sure, you can represent raster as points, but it's hideously inefficient to do so.)
Basically, you shouldn't ever be choosing between the two. If you're thinking of using a representation system like S2 for raster data (e.g. images or anything else on a regular grid), rethink things a bit.
- torrey_hoffman 5y agoMinor correction: S2's fundamental geometric types are points, edges, and S2 cells. But your point about raster data is correct. If you're given a raster, like a satellite image, S2 has no built-in type to deal with it. On the other hand, if you can define your own raster, like if you're making a heatmap or something and don't care about it being a lat-long aligned grid, you can use S2 cells like a raster, as they're roughly square and they tile the surface. This is pretty common as it's fast and convenient. It has some advantages over lat-lng grids as well, because S2 cells are roughly equal in size across the whole surface of the sphere, while lat-long aligned pixels obviously get really warped near the poles, etc.
- jofer 5y agoThe issue with using S3 cells as a raster is the inefficiency in storing information. You're basically making a row in a DB for every _pixel_ with the S2 approach. The main point of raster storage is that it avoids the overhead associated with the "row for every datapoint" approach. E.g. the X and Y are implicit and it's simply a big array of Z values, then (optionally but commonly) a sequence of additional downsampled Z arrays for efficient lookup of low-res versions. These are typically tiled for efficiency of extraction of sub-regions. The simplest systems are flat pyramids of raster files in a bucket or on a filesystem. Things like tileDB are essentially a couple of layers on top of this type of idea. Either way, the basic unit is a few thousand pixels instead of 1 pixel, which is generally more efficient, as users are usually requesting millions of pixels. A key advantage is that things are fundamentally in raster format and don't need to be translated back to raster format for the end user. Basically, the requests / etc that wind up being made are for fundamentally pixels that are parallelograms of some sort in regions that are parallelograms of some sort. After all, this has to go into a .tif/.png/.hdf/etc container at the end of the day. There are a lot of advantages to storing data in a form that's close to that.