4 ms·
Hey, I'm the author of the post and I wish I could simply write I disagree with you and how - but you haven't explained how the integer coordinates help alignin
by bartwr 6y ago
Hey, I'm the author of the post and I wish I could simply write I disagree with you and how - but you haven't explained how the integer coordinates help aligning the grids, so hard for me to disagree with any specifics.
I show how integer coordinates work in my post and what is the problem ("hanging" pixels on the right, problems with resampling). Integer coordinates lead to all those nasty bugs with resampling like the original Tensorflow.
There is a reason why all libraries and GPUs switch to this convention; as it's most sane from the POV of signal processing - and for the most common box reconstruction filter.
I obviously agree with Alvy Ray Smith and his work is super influential. I like his essay (as it's not really a paper!) a lot, and think that programmers don't think about reconstruction filters enough (I have a few blogs posts touching lightly on the topic as well). But this is very different - pixel being a square or not and the optimal reconstruction filters for filtering Monte Carlo renderings don't answer how to align pixel the grids when resampling images and useful conventions.
Finally, I also emphasize at least a few times that while my "default" convention of half pixel offsets is reasonable (i.e. you can multiply UVs by 2 and get decent resampling behavior with any filter - including windowed sincs, Cat-Rom etc), it's just important to understand the one your signal is already represented in and follow it.
- mark-r 6y agoSorry, my reply missed an important point - you did an excellent job of analyzing how GPU shaders work, and why they work that way. And for that I thank you. But it just means that I disagree with the entire GPU industry. My reasoning was in the reply, but perhaps it was too subtle. If you have an infinitesimally thin vector line running through a raster pixel, where should it intersect the pixel? I contend that it should run through the center of the raster pixel, so that it continues to align if you make the vector line thicker. If your raster pixel is at (0,0), you don't want to have to offset your vector line to (0.5,0.5) to make it match. That way lies madness, I know from painful experience. Also when you resize you don't want your input and output coordinates to line up exactly. Why? Because your output should be independent of your input size. For an integer multiple it might not matter so much, but consider for example resizing to 2.5x. If you double the size of your input in both directions, the upper-left corner of the output shouldn't change just because of alignment issues. So for example, if you're doubling the size of a 1-D image with samples at [0, 1] you should be interpolating the points at [-0.25, 0.25, 0.75, 1.25]. That way when you double the input to [0, 1, 2, 3] your output will also double to [-0.25, 0.25, 0.75, 1.25, 1.75, 2.25, 2.75, 3.25]. Of course when you pass that output to the next stage in your pipeline the coordinates are on a new axis and revert to integers again. I've given this subject years of thought, and I back it up with sample code that I use for my own private image resizing. Here's your example downsized to 50% then upsized to 200% using a Lanczos-5 filter: http://marksblog.com/share/LanczosDownUp.png http://marksblog.com/share/LanczosDownUp.png. No half-pixel offsets here.