3 ms·
Jumping Flooding Algorithm (JFA) gives very good results but you need to do a multi-pass render. If you want a single pass solution and let the GPU handle most
by rwbt 4y ago
Jumping Flooding Algorithm (JFA) gives very good results but you need to do a multi-pass render. If you want a single pass solution and let the GPU handle most of the work you can render cones[1] using an Orthographic view matrix. It is also super fast thanks to depth buffer optimizations.
[1] - https://nullprogram.com/blog/2014/06/01/ https://nullprogram.com/blog/2014/06/01/
- whizzter 4y agoJFA might be multi-pass but the pass-count is logarithmically resolution bound and doesn't depend on the number of "seeds", this cone-rendering stuff creates overdraw (basically extra passes), as you can read in the article you linked he becomes draw-bound at around 500 seeds. I rewrote my diffusion curve[1] rendering code to GPU the other year and rendering diffusion curves is basically analogous to voronoi rendering (your first pass should be seeding a canvas with the closest source color, then an varying amount of blurring depending on how the distance to the closest source, the max being canvas resolution so it can be done in log(res) passes). Since the curves (ie "voronoi seeds") are splines you want to render them at canvas resolution, so for just a simple circle on a 2048x2048 canvas you're looking at about 4-8k of seeds just for this simple case, but reducing the size of the cone isn't an option (the center of the canvas will have a 2k pixel distance to the closest pixel), so you'd basically be rendering a "large" amount of huge cones. Anyhow I found a solution to the nearest pixel+distance problem by running a couple of repeated axis aligned "run" passes(kind of like flood-filling), had I know of this JFA algo I probably would've used that. I did stumble upon the cone rendering but the overdraw issue concerned me. If I get some time over I should revisit that code and see if JFA would be faster.. the number of passes should be in the same ballpark but JFA should be more paralell since my compute shader had loops and was parallelism bound by each axis dimension. [1] - https://maverick.inria.fr/Publications/2008/OBWBTS08/ https://maverick.inria.fr/Publications/2008/OBWBTS08/
- rwbt 4y agoNot sure if you read the article till the end but > On the same machine, the instancing version can do a few thousand seed vertices (an order of magnitude more) at 60 frames-per-second, after which it becomes bandwidth saturated. If you have a powerful GPU you can get even faster results. I have a project that does Lloyd's relaxation using voronoi and for me the using cones with a depth buffer gave the best result.
- whizzter 4y agoRight, in my case the "few thousands" was for a single fairly simple shape (diffusion curves are basically continious line/splines for each shape border, so a single horizontal spline over the screen would generate almost 2k or 4k seeds depending on type at 1080p ). Real diffusion curve images would multiply that by an order or two at the least depending on complexity of the image (the algo I wrote as well as JFA might have a higher fixed number of iterations but it's fixed so the number of seeds don't matter). Besides, GPU's have turned more and more general so even if there's optimizations for the triangle rendering the compute case will probably gain even more from more recent GPUs. Optimizing for GPU's can be funny, and I wouldn't be surprised if the break-point on where the tri-fans are faster is higher than I expect, but at the same time unless it's a tile-based renderer I'm pretty sure there will be a break-point (also tilebased renderers can sometimes fall of performance cliffs in special cases).
- rwbt 4y agoYes- the drawing cones performance will eventually bottleneck depending on the num of sites, in that case your proposed solution makes much more sense.
- Lichtso 4y agoVery interesting approach, but I wonder why use geometry to get cones? Why not just fill quads and do distance (sqrt) to center in a fragment shader? That would not even need an approximation but be exact.
- rwbt 4y agoIf you fill quads, you won't get to use the depth buffer to occlude fragments easily. GPUs can do a lot of depth buffer optimizations before they run the fragment shader.
- yarg 4y agoThat's neat - makes me think of 2D light-cones. The author doesn't mention it, so I'm not sure if he noticed - but this method is trivially extensible to weighted Voronoi graphs. All you need to do is change the gradients of the cone such that d(radius)/d(z) = weight (and that makes things significantly easier when you need to deal with curved inter-cell boundaries). This potentially could be used to generate 3D voxels, but you'd need depth buffer support in 4 dimensions.