4 ms·
The issue with using S3 cells as a raster is the inefficiency in storing information. You're basically making a row in a DB for every _pixel_ with the S2 appro
by jofer 5y ago
The issue with using S3 cells as a raster is the inefficiency in storing information.
You're basically making a row in a DB for every _pixel_ with the S2 approach.
The main point of raster storage is that it avoids the overhead associated with the "row for every datapoint" approach. E.g. the X and Y are implicit and it's simply a big array of Z values, then (optionally but commonly) a sequence of additional downsampled Z arrays for efficient lookup of low-res versions. These are typically tiled for efficiency of extraction of sub-regions. The simplest systems are flat pyramids of raster files in a bucket or on a filesystem. Things like tileDB are essentially a couple of layers on top of this type of idea. Either way, the basic unit is a few thousand pixels instead of 1 pixel, which is generally more efficient, as users are usually requesting millions of pixels.
A key advantage is that things are fundamentally in raster format and don't need to be translated back to raster format for the end user.
Basically, the requests / etc that wind up being made are for fundamentally pixels that are parallelograms of some sort in regions that are parallelograms of some sort. After all, this has to go into a .tif/.png/.hdf/etc container at the end of the day. There are a lot of advantages to storing data in a form that's close to that.