4 ms·
Project author here. Hello :) Let me just clarify a few bits. The GL wrapper is deliberately this low-level, because the library provides a baseline for a much
by mosra 4y ago
Project author here. Hello :)
Let me just clarify a few bits. The GL wrapper is deliberately this low-level, because the library provides a baseline for a much broader range of uses than just games themselves. For an application that performs real-time image / video processing on a GPU, draw call or transparency sorting would be extra overhead. An _actual game engine_ that implements its own custom draw batching would have to work around the implicit batching, if it was there. The wrapper is just lower-level than you expected, that's all, and same is for the currently unadvertised in-progress Vulkan wrapper.
Regarding uniforms, other people already pointed out there is a UBO interface [1]. But there are also users who _rely_ on WebGL 1.0 or GL 2.1 compatibility (imagine that!), so I can't just drop the classic uniform path (and other related suffering). On the other hand, not all hardware and drivers have the UBO path faster in all cases, so providing an UBO iterface and failing back to classic uniforms only if UBOs are not available doesn't make sense either. In my measurements, the only case where UBOs actually improved performance across all platforms was in a multi-draw scenario, with UBOs being split by frequency of update, containing data for 100+ draws and being uploaded at most once per frame. The engine supports such use case [2], but doing that efficiently means having your data and logic ready upfront. If the library would try to attempt that automagically implicitly, it wouldn't end well and the gains would be minimal. That being said, I don't rule out a possibility for an opt-in high-level "render graph" or "pass management" API that does fancy things on its own, but it needs a lower-level layer built upon. Not all or nothing, layers. Layers so projects with specific needs can ignore that and use the layer below directly.
There's no global state anymore except for the Renderer, and I'm painfully aware of that last piece, yes. It's getting fixed, but because on certain targets (WebGL) every call matters, it's not as simple as packing all that state into a single struct and then resetting it in a bulk before a draw call.
"SceneGraph"? Yes, that part is in a maintenance mode, and I recommend new users to use entt, flecs or other ECS frameworks instead, until I build a replacement for this ancient technology that's there since 2010 (yes). But a lot of projects still want to use it because it's _damn convenient_, so I can't really remove it yet either to save the project from immediately raising a "red flag" to certain people ;)
Finally, the project has been under frantic development since its last version annoucement in 2020. Not everything is done and documented yet, the website shows slightly outdated info and the examples don't really use the fresh APIs yet. The Vulkan parts you saw as "unrelated" or "unused", those are still getting finished. Plus there's of course a roadmap to satisfy the needs of long-running Magnum-based projects, so the focus is often in different areas than the "yikes"-triggering ones. Such as asset management and conversion [3].
Have a nice day.
[1] https://doc.magnum.graphics/magnum/shaders.html#shaders-usage-ubo https://doc.magnum.graphics/magnum/shaders.html#shaders-usag...
[2] https://doc.magnum.graphics/magnum/shaders.html#shaders-usage-multidraw https://doc.magnum.graphics/magnum/shaders.html#shaders-usag...
[3] https://blog.magnum.graphics/announcements/new-geometry-pipeline/ https://blog.magnum.graphics/announcements/new-geometry-pipe...
- Jasper_ 4y ago> In my measurements, the only case where UBOs actually improved performance across all platforms was in a multi-draw scenario, with UBOs being split by frequency of update, containing data for 100+ draws and being uploaded at most once per frame. That's interesting -- back when I was working on mobile platforms, I found that UBOs were a win from call overhead alone -- replacing N different glUniform* calls with one UBO upload per object was a win on both the Adreno and Mali drivers, though we did have to work around some driver bugs (as is Android tradition). I haven't done the measurements as well, but I suspect it's also a win on WebGL 2 over the old strategy, because calling the WebGL functions is just more expensive than it should be (I've been told this is getting fixed in the browsers eventually, but I was also told that nearly 3 years ago, and it hasn't happened yet...) But these days I definitely advocate for a single giant UBO for your scene, uploaded at the start of your frame. I prepare all draw calls at the start of the frame in structures and allocate a single UBO for the frame at the same time, and add them to the relevant passes, sort the passes, and draw them using something that makes as few state changes as possible between the draw calls. > it's not as simple as packing all that state into a single struct and then resetting it in a bulk before a draw call. No, it's definitely not, you need to track your current state and the draw call's state, but I think it's a lot easier to program against, and more readily adapts to Vulkan, Direct3D, Metal, etc., and is the pattern adopted by wgpu/sokol/bgfx. https://github.com/magcius/noclip.website/blob/master/src/gfx/platform/GfxPlatformWebGL2.ts https://github.com/magcius/noclip.website/blob/master/src/gf...
- mosra 4y ago> on mobile platforms, I found that UBOs were a win from call overhead alone That was what I got from it also. Be it a phone or WebGL translated using ANGLE to something else, every call mattered. But on desktop drivers that were relatively low-overhead to begin with, going to UBOs was not always a win, especially on Intel GPUs with dynamic indexing into a UBO array or if I didn't use the GL 4.4 immutable buffer storage. And it didn't matter if I was on Linux with Mesa or on an Intel Mac. Of course no such problem on NV or AMD, but it made me approach UBOs with caution. I still need to dive deeper, because it'd be silly if such implicit overhead carried over to Vulkan. So what I found ideal is submitting a multi-draw call with as many draws as possible, trading the weird extra UBO overhead with less time spent in the driver. That's limited by max bound UBO size, so depending on how well I pack the data I might only get as low as 256 draws in a single submit, and then I'd have to rebind another UBO range and submit another multi-draw call. > you need to track your current state and the draw call's state Yup, and that's the thing -- once I ship it, it has to be completely bug-free to not break existing apps in boundary conditions. And since a large portion of users is doing crazy advanced stuff like sharing the GL context with Qt or Gtk, or using it together with 3rd party GL libraries for drawing vector graphics and whatnot, the state tracker needs to be able to reset itself (or, worse, reset the GL state to not make a buggy 3rd party lib misbehave). Altogether it's a lot of tiny things to get right, and somehow there was always something more important that users needed so this fell on the bottom of the priority list. If I would be creating a new GL wrapper from scratch, I'd definitely go this way from the start, it's harder to retrofit a codebase with over a decade of history and a ton of depending projects that learned to expect smooth upgrades :)