4 ms·
Yep. A uniform mesh of flat quads with a vertex shader doing the displacement will perform better.
by kuzehanka 7y ago
Yep. A uniform mesh of flat quads with a vertex shader doing the displacement will perform better.
- mourner 7y agoWould love to learn more, since it seems counter-intuitive — are there any studies / benchmarks on this for modern hardware? Why would a uniform mesh with, say, 10x more triangles perform better if it's not updated often (e.g. a few times a second)?
- kuzehanka 7y agoIt'll come down to how frequently you're doing updates. A few times a second will certainly cost more than a prefab mesh with a texture fetch shader. If you can precompute various LOD chunks for each tile and cache them, under some circumstances that's optimal, but I don't think any of them are webgl. A typical heightmap based terrain renderer today will consist of several (3-5) prefab LOD meshes deformed by height map textures in shader, like[1]. Swapping a textures in and out of vram is cheaper than swapping geometries, even if those geometries are of lower resolution. Modern GPUs are almost never bound by the number of triangles on screen. The bottleneck will be the fillrate and bandwidth and pipeline stalls, especially in webgl. If you can throw down more triangles to do less of the other things, it's almost always worth it. I have a toy mapbox tile renderer based on this approach. I'd love to know if you guys have plans to implement full 3d like google maps any time soon. [1]: https://imgur.com/a/51g2bB4 https://imgur.com/a/51g2bB4
- mourner 7y agoI would really love to see your toy Mapbox renderer — please share it, even if unpolished! We'd benefit a lot from fresh ideas on how to improve Mapbox GL performance.
- pastrami_panda 7y agoMainly because GPU throughput is increasing much faster than CPU throughput. Spending cycles removing triangles doesn't matter if the GPU can just churn through them like nothing. In the context of a moving camera you want to reuse as much data as possible between terrain updates. With regular nested grids a ton of data can be reused, only the height needs updating as you move the grids around (and height can be reused between adjacent grids half the time if they have a power of two resolution relationship). But baking irregular meshes into a tree-structure could be powerful too I suppose, I'm not that familiar with algorithms involving them. All of this is highly dependant on actual size of terrain and LOD requirements, terrain rendering is a deep rabbit hole.
- mourner 7y agoYeah, it's a thin balance, and to me as someone without years of experience in graphics programming, GPU feels like such a blackbox — hard to predict what will be the actual bottleneck until you implement an approach fully, with real-world loads, and even then it's hard to assess, especially with such variety of GPU performance characteristics on different devices. So what's simpler / easier to implement in a given system will play an important role, and so are any additional things you want to do with the height data like collision detection, querying, data analysis etc. Anyway, that's a super-exciting topic to learn more about! Thanks a lot for all the thoughtful comments.
- taneq 7y agoGPUs are effectively “magic” and can do particular things hundreds or thousands of times faster than CPUs. Offloading a per-vertex or per-pixel task to the GPU, especially if doing so reduces bandwidth requirements, is almost always a big win.
- vardump 7y agoI think reality is closer to this: GPU rasterization is 20-50 times faster than CPUs and throughput computing maybe 5-20 times faster. For example consumer Ryzen 3 with 12 cores peaks at 32 FLOPS/core/cycle. So at 3.8 GHz, peak ~1.4 SP TFLOPs (or 0.7 DP TFLOPs). But just 50-60 GB/s memory bandwidth limits it somewhat. Quick googling says consumer Nvidia RTX 2080 peaks at 10 SP TFLOPs (or 0.314 DP TFLOPs, yes, less than half than the CPU example). Memory bandwidth being at 448 GB/s. GPUs win massively at rasterization, because they have huge memory bandwidth, a large array of texture samplers with hardware cache locality optimizations (like HW swizzling), texture compression, specialized hardware for z-buffer tests and compression, a ton of latency hiding hardware threads, etc. But they're definitely not thousands or even hundreds of times faster.
- taneq 7y agoOr even better, couldn’t you use a single quad and a geometry shader?
- pastrami_panda 7y agoGeometry shaders seems sort of the bastard child of the graphics pipeline. It was a mess to implement in hardware and the throughput isn't that great when creating tons of triangles. The tesselation stages of the pipeline addresses this and is more suitable for arbitrary subdivision of triangles.
- vardump 7y agoWhy not a single quad and a fragment (=pixel) shader? A bit of ray marching, something like this [0]. [0]: http://www.iquilezles.org/www/articles/terrainmarching/terrainmarching.htm http://www.iquilezles.org/www/articles/terrainmarching/terra...
- nineteen999 7y agoNot very helpful if you need to do collision on the displaced mesh though, I would have thought, since a vertex shader only make it "appear" as if the vertices are not in their original position? I'm thinking of tracing for foot placement in conjunction with inverse kinematics for walking characters, or vehicle wheels traversing the terrain. I've run into this problem with UE4.
- manfredo 7y agoWould it be possible to either compute or approximate the fragment shader's vertex results on segments of the terrain that require collision detection? I bet it wouldn't be feasible to do to check collisions for gun projectiles in a video game, but it might be feasible for something like making sure the player's model doesn't clip beneath the tessellated terrain.
- nineteen999 7y agoUE4 has adaptive tessellation anyway, and you can scale the tessellation multiplier based on distance with a material function so that the landscape has more triangles closer to the camera and less further away. Beyond a certain distance (where the player wouldn't notice the detail anyway, so there's no point drawing it) you can scale the multiplier to zero to improve performance. I'm not an expert but honestly in a game engine like UE4 where collision is important, it seems like the adaptive geometry tessellation approach is the way to go, as least from my initial experiments.
- manfredo 7y agoI was thinking of a somewhat different use case. I was thinking of a flight simulator where tessellation only occurred close to the camera, to provide the appearance of a more detailed mesh when viewed up close. The problem is, with a quad mesh where each vertex is 20 or 30 meters apart a plane flying nap of the earth could end up clipping through the tessellated mesh but not actually collide with the non-tessellated terrain quad mesh being used for collision detection. Or worse, the player could collide with the terrain quad mesh but the tessellation makes it visually appear to have not collided with the terrain. To only two ways I see to solve this are to either not use tessellation up close, or to compute with the CPU the vertices produced by the tessellation and use that mesh for collision.