5 ms·
> This performance optimization doesn't really work for variable-width blurs... An alternative would be to use Summed-area tables [1] where each texel contains
by mrhyperpenguin 15y ago
> This performance optimization doesn't really work for variable-width blurs...
An alternative would be to use Summed-area tables [1] where each texel contains the sum of all texels that are above and to the left of the current texel. This allows variable-width blurs to be computed in constant time, using only four texture samples (!).
And the concept of a variable-width blur is an approximation to a related concept in photography/optics called the "circle of confusion." [2] The circle of confusion, combined with bokeh sprites, are usually what AAA games are doing for depth of field [3].
See [4] for more details.
[1] http://http.developer.nvidia.com/GPUGems3/gpugems3_ch08.html http://http.developer.nvidia.com/GPUGems3/gpugems3_ch08.html
[2] http://en.wikipedia.org/wiki/Circle_of_confusion http://en.wikipedia.org/wiki/Circle_of_confusion
[3] http://udn.epicgames.com/Three/BokehDepthOfField.html http://udn.epicgames.com/Three/BokehDepthOfField.html
[4] http://mynameismjp.wordpress.com/2011/02/28/bokeh/ http://mynameismjp.wordpress.com/2011/02/28/bokeh/
- daenz 15y agoI love this stuff. I haven't had a chance to implement a summed area table yet for anything, but I did do variance shadow maps very recently (did you mean to cite that?). I did read about them though, and one drawback that I remember was that you might have a precision overflow in the bottom/top corner where the summations are the greatest?
- mrhyperpenguin 15y agoThe GPU Gems article talks about summed area tables in the context of variance shadow maps, specifically percentage closer soft shadows variance shadow maps where variable-width blurs are used extensively. Didn't really mean to cite the stuff about VSMs (sorry), more so the SAT stuff toward the latter half of the article as it would apply to DoF. I believe the precision/overflow issues can be mitigated using modern GPU hardware (the authors mention this in section 8.5.2 of the article.) Now-a-days it is pretty much standard to have floating-point render targets (especially with deferred rendering and light-pre pass renderers) so 16/32 bits per component should be adequate especially if you apply the tricks the authors present in the article. Not to mention that now we have DirectCompute/OpenCL/CUDA. Though I don't really see people talking about SATs that much (I myself have never implemented them.) Maybe there is an underlying reason why people don't? Bandwidth maybe? I would imagine it would be pretty taxing on the system to recompute the SAT every frame.