14 ms·
WebGPU – All of the cores, none of the canvas
- xyhhx 3y agocan't wait to see what exciting new exploits are in store for us with this
- dezmou 3y agoBitcoin mining gro brrrr https://github.com/dezmou/SHA256-WebGPU https://github.com/dezmou/SHA256-WebGPU
- ReactiveJelly 3y ago> The most popular of the next-gen GPU APIs are Vulkan by the Khronos Group, Metal by Apple and DirectX 12 by Microsoft. ... (WebGPU) introduces its own abstractions and doesn’t directly mirror any of these native APIs. Huh. I was wondering about that. Until now I just figured every "Web*" thing was browsers exposing (to JS alone) something that they already compiled in: - WebRTC is ffmpeg - Canvas is Skia - WebGL is ANGLE - WebCodecs is also ffmpeg - WebTransport is QUIC - WebSockets are TCP I might be wrong on some of those.
- nl 3y agoI think all of these are wrong. > WebRTC is ffmpeg No. WebRTC is a transport protocol for media communications. > Canvas is Skia Skia is a graphics engine you can build a Canvas implementation on top of > WebGL is ANGLE I don't know what ANGLE is, but WebGL is based on OpenGL. As this article says "WebGL’s API is really just OpenGL ES 2.0" > WebCodecs is also ffmpeg Both allow conceptually similar things (low level access to specific parts of a media stream). But the APIs are dramatically different. > WebTransport is QUIC No it is an API to expose lower level parts of HTTP/3 to developers. HTTP/3 uses QUIC as a transport protocol, but it is very wrong to say it "is" QUIC. > WebSockets are TCP Well WebSockets is built on top of TCP. As is HTTP/1 and HTTP/2. (HTTP/3 uses UDP via QUIC)
- dagenix 3y ago> I don't know what ANGLE is, but WebGL is based on OpenGL. As this article says "WebGL’s API is really just OpenGL ES 2.0" I believe that ANGLE is a library that is widely used to implement WebGL by translating OpenGL ES calls into Direct 3D calls: https://en.wikipedia.org/wiki/ANGLE_(software) https://en.wikipedia.org/wiki/ANGLE_(software)
- nl 3y agoThe OP seems to be confusing the concept of API and library layering with the idea that something is something else. To be clear: in software almost everything is built on top of other things using libraries. This doesn't mean the new thing that is built is that thing at all, and indeed that new thing may be able to switch out the lower level library for a different implementation.
- still_grokking 3y ago> > WebTransport is QUIC > No it is an API to expose lower level parts of HTTP/3 to developers. HTTP/3 uses QUIC as a transport protocol, but it is very wrong to say it "is" QUIC. Well, that's the only thing the parent got almost right. (The rest was obvious nonsense, though. I agree.) WebTransport is of course not QUIC. But it allows to use QUIC streams almost directly. There are no "lower parts" of HTTP/3 other than QUIC. HTTP/3 is a quite thin layer directly atop of QUIC. With WebTransport you send a CONNECT request with some special flags / headers to the web server and given a correct response you can start using raw QUIC streams over your HTTP/3 QUIC connection. The overhead to get at your raw QUIC streams is quite low and a one time thing. From there you can directly use all the capabilities QUIC gives you (client or server initiated reliable unidirectional and bidirectional data streams or unreliable datagrams transporting arbitrary binary messages over a kind of "virtual" connection).
- CMCDragonkai 3y agoI was looking into this but are you sure web transport will expose bidirectional binary quic streams and datagrams to the browser? If so please link so I can start hacking!
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- kalleboo 3y ago> Canvas is Skia Canvas is Apple Quartz. They implemented it in WebKit for their dashboard widgets (which were implemented what we called "HTML5" back then) which leaked into Safari, and it turned out to be so useful that it got adopted in other browsers as a WHATWG standard.
- Jasper_ 3y agoI'm a graphics programmer who has quite a bit of experience with WebGL, and (disclaimer) I've also contributed to the WebGPU spec. > Quite honestly, I have no idea how ThreeJS manages to be so robust, but it does manage somehow. > To be clear, me not being able to internalize WebGL is probably a shortcoming of my own. People smarter than me have been able to build amazing stuff with WebGL (and OpenGL outside the web), but it just never really clicked for me. WebGL (and OpenGL) are awful APIs that can give you a very backwards impression about how to use them, and are very state-sensitive. It is not your fault for getting stuck here. Basically one of the first things everybody does is build a sane layer on top of OpenGL; if you are using gl.enable(gl.BLEND) in your core render loop, you have basically already failed. The first thing everybody does when they start working with WebGL is basically build a little helper on top that makes it easier to control its state logic and do draws all in one go. You can find this helper in three.js here: https://github.com/mrdoob/three.js/blob/master/src/renderers/webgl/WebGLState.js https://github.com/mrdoob/three.js/blob/master/src/renderers... > Luckily, accessing an array is safe-guarded by an implicit clamp, so every write past the end of the array will end up writing to the last element of the array This article might be a bit out of date (mind putting a publish date on these articles?), but these days, the language has been a bit relaxed. From https://gpuweb.github.io/gpuweb/#security-shader https://gpuweb.github.io/gpuweb/#security-shader : > If the shader attempts to write data outside of physical resource bounds, the implementation is allowed to: > * write the value to a different location within the resource bounds > * discard the write operation > * partially discard the draw or dispatch call The rest seems accurate.
- 29athrowaway 3y agoCan you elaborate more on this? It seems interesting > if you are using gl.enable(gl.BLEND) in your core render loop, you have basically already failed.
- tim1994 3y agoI'd also be interested in details on this but I assume the gl.enable() API changes fundamental things about the rendering pipeline. It allows enabling things like depth testing and stencil (both involve an extra buffer) and face culling (additional tests after vertex shader). For blending in particular I think it requires the fragment shader to first read the previous value from the frame buffer. Changes this stuff is probably not a trivial operation and requires a lot of communication with the GPU which is slow (just a guess). If you want to change blending for each draw call you can change the blending function or just return suitable alpha values from the fragment shader.
- nevi-me 3y agoMy disappointment with WebGPU has been limited data type support. I wanted to write some compute stuff with it, but the limitation of not supporting a lot of integer sizes made it undesirable. Does anyone know if the spec is likely to be revised to add more support over time?
- dezmou 3y agohttps://github.com/gpuweb/gpuweb/issues/3620 https://github.com/gpuweb/gpuweb/issues/3620
- pjmlp 3y agoEventually, but expect a progression rate measured in years.
- TazeTSchnitzel 3y agoWhat kinds of integer sizes? Depending on the target GPU, 64-bit integers are likely to not be available at all, or be quite slow. If you need 8-bit or 16-bit integers, on the other hand, those can be trivially emulated with 32-bit operations.
- flohofwoe 3y agoThe current WebGPU spec is basically the common feature subset across desktop and mobile GPUs, and what of those hardware features are actually exposed by D3D12, Metal and Vulkan. If something is missing then it's most likely the fault of some random mobile GPU. Such missing features might be added later via optional extensions, but the focus was to get the thing out of the door first.
- atgctg 3y agoAs an example, INT8 support in WebGPU would enable running quantized models, allowing larger LLMs to run locally in the browser. See Limitations section here: https://fleetwood.dev/posts/running-llms-in-the-browser https://fleetwood.dev/posts/running-llms-in-the-browser
- hutzlibu 3y agoFollowing the article, you build a simple 2D physic simulation (only for balls). Did by chance anyone expand on that to include boxes, or know of a different approach to build a physic engine in WebGPU? I experiemented a bit with it and implemented raycasting, but it is really not trivial getting the data in and out. (Limiting it to boxes and circles would satisfy my use case and seems doable, but getting polygons would be very hard, as then you have a dynamic size of their edges to account for and that gives me headache) A 3D physic engine on the GPU would be the obvious dream goal to get maximum performance for more advanced stuff, but that is really not an easy thing to do. Right now I am using a Box2D for wasm and it has good performance, but it could be better. https://github.com/Birch-san/box2d-wasm https://github.com/Birch-san/box2d-wasm The main problem with all this is the overhead of getting data into the gpu and back. Once it is on the gpu it is amazingly fast. But the back and forth can really make your framerates drop - so to make it worth it, most of the simulation data has to remain on the gpu and you only put small chanks of data that have changed in and out. And ideally render it all on the gpu in the next step. (The performance bottleneck of this simulation is exactly that, it gets simulated on the gpu, then retrieved and drawn with the normal canvasAPI which is slow)
- tomsmeding 3y ago> But the back and forth can really make your framerates drop - so to make it worth it, most of the simulation data has to remain on the gpu and you only put small chanks of data that have changed in and out. And ideally render it all on the gpu in the next step. In my (limited, cuda so not webgpu) experience, memory transfers are fast and computation is fast, the thing that is slow is memory transfer _latency_. Doing a memory transfer takes a long time, but if you're doing one anyway, might as well transfer the world. Is my recollection correct?
- hutzlibu 3y ago"Doing a memory transfer takes a long time, but if you're doing one anyway, might as well transfer the world." Not in my experience and experiments. But I am pretty much a beginner with WebGPU and might be missing a lot. Otherwise yes, latency is the big issue as well. Sometimes all is well, sometimes nothing happens for 20+ms.
- andrewstuart 3y agoWill there be better typography in WEbGPU?
- flohofwoe 3y agoWebGPU is completely separate from the browser's text rendering engine (unfortunately).
- andrewstuart 3y agoI ask because WebGL almost never has typography. I assume there are technical reasons.
- flohofwoe 3y agoBecause text rendering is (very) hard, and most 3D rendering people are not text rendering experts, but just want to get some text on screen even if it looks ugly. Unfortunately the browser lacks a proper layered API design with low level APIs at the bottom (like WebGL and WebGPU), medium level APIs in the middle (like text rendering and layout), and high level APIs at the top (like the DOM and CSS) - so if you want to render text in WebGL or WebGPU, you are entirely on your own.
- paulgb 3y agoIt won't live in WebGPU itself, but I do expect to start to see more third-party libraries for text. There’s already wgpu_glyph (https://github.com/hecrj/wgpu_glyph/tree/master https://github.com/hecrj/wgpu_glyph/tree/master) which uses a glyph atlas (CPU-rendered sprite map of characters), but techniques for signed-distance field fonts have come a long way too.
- Const-me 3y agoGood article, but couple remarks. > most hardware seemingly just runs workgroups in a serial order The hardware runs them in parallel, but it’s complicated. The nVidia GPU I’m currently using has 32-wide SIMD, which means groups of 32 threads run in parallel, exactly in lockstep. Different GPU APIs call such group of threads wavefronts or warps. Each core (my particular GPU has 28 of these) can run 4 of such wavefronts = 128 threads in parallel. When a shader has more than 128 threads, or when the GPU core is multi-tasking running multiple workgroups of the same or different shaders, different wavefronts will run sequentially. And one more thing, the entire workgroup runs within a single GPU core, even when the shader pushes workgroup size to the limit with 1024 threads per workgroup. “Sequentially” doesn’t mean the order of execution is fixed, or predefined, or fair. Instead, the GPU is doing rather complicated scheduling trying to hide latency of computations and memory transactions. While some wavefront is waiting for data to arrive from memory, instead of sleeping the GPU will typically switch to another active wavefront. Many modern CPUs do that too because hyperthreading, but CPUs only have 2 threads per core, they are visible to OS as two distinct virtual cores. For GPUs the number is way higher, only limited by amount of in-core memory, and amount of that memory required by the running shaders. > as the difference between running a shader with @workgroup_size(64) or @workgroup_size(8, 8) is negligible. So this concept is considered somewhat legacy. I think it’s convenience, not legacy. When a shader handles 2D data like a matrix or an image, it’s natural to have 2D workgroup sizes like 8x8. Similarly, when a shader processes 3D data like a field defined on elements or nodes of 3D Cartesian grid, it can be slightly easier to write compute shaders with workgroups of 4x4x4 or 8x8x8 threads.