6 ms·
GPU-Driven Clustered Forward Renderer
- zeristor 1y agoApostrophe as a number separator? Where’s that from?
- deleted 1y ago[deleted]
- dahart 1y agoSwitzerland and Italy for two. https://en.wikipedia.org/wiki/Decimal_separator# https://en.wikipedia.org/wiki/Decimal_separator# Also note C++14 introduced the apostrophe in numeric literals! https://en.cppreference.com/w/cpp/language/integer_literal https://en.cppreference.com/w/cpp/language/integer_literal
- lacoolj 1y agoLearn somethin new every day. And I would never have known this existed without hackernews
- logdahl 1y agoInteresting that Sweden explicitly do NOT use it... Not sure where i picked it up! :-)
- qingcharles 1y agoI've started using the underscore in my code since that is becoming the (non-localized) standard and trendy: https://en.wikipedia.org/wiki/Integer_literal#Digit_separators https://en.wikipedia.org/wiki/Integer_literal#Digit_separato...
- m-schuetz 1y agoApostrophe are nice because they are not ambiguous. Started using them myself after getting used to them from C++ and learning that they are used in switzerland.
- unclad5968 1y agoThis is awesome! At the end you mention the 27k dragons and 10k lights just barely fits in 16ms. Do you see any paths to improve performance? I've seen some demos on with tens/hundreds of thousands of moving lights, but hard to tell if they're legit or highly constrained. I'm not a graphics programmer by trade. I need a renderer for a personal project and after some research decided I'll implement a forward clustered renderer as well.
- logdahl 1y agoWell, the core issue is still drawing. I took another look at some profiles again and seems like its not the renderer limiting this to 27k! I still had some stupid scene-graph traversal... But clustering and culling is 53us and 33us respectively, but the draw is 7ms. So a frame (on the GPU-side) is like 7ms, and some 100-200 us on the CPU side. Should really dive deeper and update the measurements for final results...
- godelski 1y agoI haven't look at the post in the detail it deserves, but given your graphs the workload looks pretty bursty. I'd suspect there are some good I/O optimizations or some predication. Definitely that last void main block looks ripe for that. But I'd listen to Knuth, premature optimization and all, so grab a profiler. I wouldn't be surprised if you're nearing peak performance. Also NVIDIA GPUs have a lot of special tricks that can be exploited but are buried in documentation... if you haven't already seen it (I suspect you have), you'd be interested in "GPU Gems". Gems 2 has some good stuff on predication. But also, really good work! You should be proud of this! Squeezing that much out of that hardware is no easy feat.
- gmueckl 1y agoThis seems fairly well optimized. There's probably room to squeeze out some more perf, but not dramatic improvements. Maybe preventing overdraw of shaded pixels by doing a depth prepass would help. Without digging into the detailed breakdown, I would assume that the sheer amount of teeny tiny triangles is the main bottleneck in this benchmark scene. When triangles become smaller than about 4x4 pixels, GPU utilization for raterization starts to diminish. And with the scaled down dragons, there's a lot of then in the frame.
- fabiensanglard 1y agoThis website has a beautiful layout ;) !
- logdahl 1y agoFun to see you ;) Love your site!
- deleted 1y ago[deleted]
- curtisszmania 1y ago[dead]
- rezmason 1y agoTen thousand lights! Your utility bill must be enormous
- Flex247A 1y agoLights in games use real electricity :)
- amelius 1y agoEven the stars use real electricity.
- cluckindan 1y agoNot really, nuclear fusion doesn’t run on electrons.
- DiabloD3 1y agoSo where does the magnetic field come from? ;) ;) ;)
- cluckindan 1y agoNuclear fusion produces a million times more energy from proton and neutron collisions than is produced by electron shells during the same event.
- amelius 1y agoThe energy leaves the star in the form of EM energy. This is also the energy that is responsible for electricity.
- monster_truck 1y agoAm I missing a link somewhere or is there no way to build/run this myself? Interested to see what a modern flagship gpu is good for
- wizzwizz4 1y ago> As some other renderers do, we share a single GPU buffer for all vertex data. Instead, we use a simple allocator which manages this contigous buffer automatically. I'm not sure what this part is supposed to say, but it doesn't look right. "Instead" usually follows differences, not similarities.