4 ms·
Zero-Copy GPU Inference from WebAssembly on Apple Silicon
- pjmlp 6mo agoGoodbye WebAssembly "security". Also, these folks should be amazed by 8 and 16 bit games development, or games consoles in general.
- wmf 6mo agoThis works in wasmtime not browsers.
- thrill 6mo agoWhy would it not work in a browser?
- trueno 6mo ago> on Apple Silicon, a WebAssembly module's linear memory can be shared directly with the GPU: no copies, no serialization, no intermediate buffers enhance > no copies, no serialization, no intermediate buffers would it kill people to write their own stuff why are we doing this. out of all the things people immediately cede to AI they cede their human ability to communicate and convey/share ideas. this timeline is bonkers.
- rvz 6mo agoThis sort of obvious pattern is an instant AI dead give-away that I keep on seeing in hundreds of blogs and code posted on this site: "Here is X - it makes Y" "That's not X, it's Y." "...no this, no that, no X, no Y." Another way of telling via code is by deducing the experience of the author if they became an expert of a different language since...yesterday. There will be a time where it will be problematic for those who over-rely on AI and will struggle on on-site interviews with whiteboard tests.
- bensyverson 6mo agoI think the days of on-site interviews with whiteboard tests may be drawing to a close faster than you suspect
- m00dy 6mo agoI also think we will never go back to good old days.
- dylan604 6mo agoIt'll put the "everything old becomes new again" idea to the test.
- JSR_FDED 6mo agoHuh, I’m 100% going to interview this way the next time I have to hire an engineer. I can’t think of a better way to get a sense of how a candidate reasons about things, and of their values - do they have a sense of responsibility, conscientiousness, team fit. All other things that could be LLM-mediated have no more signal.
- andsoitis 6mo ago> I can’t think of a better way to get a sense of how a candidate reasons about things Some ideas to help you: ask the candidate something underspecified and watch what they do first. Do they ask clarifying questions, make their assumptions explicit? After they answer ask what would change their mind, where does that break down? Pick a topic they know and ask them to explain it to a smart non-engineer. Make them estimate something they can’t look up (forces them to decompose, bound, and calibrate). Once they’ve proposed a solution to a question, change the constraints to see if they can adapt or whether they’re stuck. What you want to evaluate is dynamic reasoning, adaptability.
- saagarjha 6mo agoI'm curious what this offers over just building the host side code to be native?
- jsomedon 6mo agoMy quick guess is that this approach offers near zero overhead for gpu to access data inside sandbox with all the security/privacy benefit of sandbox.
- swiftcoder 6mo agoFor one thing, it's a lot easier to distribute a webpage than a native app
- saagarjha 6mo agoThis doesn't work with webpages though
- swiftcoder 6mo agoI somehow missed that tidbit
- agambrahma 6mo agoYes, simply for local inference -- not much, native is the obvious choice. The value would be in actor processes, where you can delegate inference without paying the 'copy tax' for crossing the sandbox boundary. So, less "inference engine" and more "Tmux for AI agents" Think pausing, moving, resuming, swapping model backend. I scoped the post to memory architecture, since it was the least obvious part ... will follow up with one about the actor model aspect.
- saagarjha 6mo agoI'm a little confused what an actor process is. To me a process is inherently local?
- EthanFrostHI 6mo ago[flagged]
- nl 6mo agoI'm pretty sure this is just "yes (parts of), memory control in WASM works"[1]. The whole Apple Silicon thing is (in this case) just added details that don't actually matter. [1] https://github.com/WebAssembly/memory-control/blob/main/proposals/memory-control/Overview.md https://github.com/WebAssembly/memory-control/blob/main/prop...
- eis 6mo agoApple Silicon uses unified memory where the CPU and GPU use the exact same memory and no copies from RAM to VRAM are needed. The article opens with mentioning just that and indeed it is the whole point of the article.
- fho 6mo agoI am always a bit baffled why Apple gets credited with this. Unified memory has been a thing for decades. I can still load the biggest models on my 10th gen Intel Core CPU and the integrated GPU can run inference. The difference being that modern integrated GPU are just that much faster and can run inference at tolerable speeds. (Plus NPUs being a thing now, but that also started much earlier. Thr 10th gen Intel Core architecture already had instructions to deal with "AI" workloads... just very preliminary)
- mirekrusin 6mo agoThat’s shared, not unified, it’s partitioned where cpu and gpu copies are managed by driver. Lunar lake (2024) is getting closer but still not as tightly integrated as apple and capped to 32GB only (Apple has up to 512GB). AMD ryzen ai max is closer to Apple but still 3 times slower memory.
- fc417fc802 6mo agoShared vs unified is merely a driver implementation detail. Regardless, in practice (IIUC) data is still going to be copied if you perform a transfer using a graphics API because the driver has no way of knowing what the host might do with the pointed-to memory after the transfer. If you make use of host pointers and run on an iGPU no copy will take place.
- fulafel 6mo ago> Apple Silicon changes the physics. The CPU and GPU share the same physical memory (Apple's Unified Memory Architecture) ... no bus! Beware the reality distortion field: This is of course how it's worked on most x86 machines for a long time. And also on most Macs when they were using Intel chips.
- littlecranky67 6mo agoWhy did all my x86 onboard iGPU reserve a fixed amount of RAM on boot, inaccessible to the OS? Why do dGPU bring their own VRAM and how to directly manipulate it from the CPU without copying?
- fulafel 6mo agoTo the first question: blame Windows I guess. But even on older chips, GPU code could access memory allocated on the CPU side so this didn't cap the amount of data your GPGPU code could crunch.
- littlecranky67 6mo agoI remember this was mostly a BIOS setting how much memory to allocate for iGPU - and once set in the BIOS, that memory was not accessible to the underlying OS (besides GPU I/O).
- fulafel 6mo agoYes, but this was to appease Windows, probably older versions and/or 32 bit versions of it.
- ben-schaaf 6mo agoCorrect me if I'm wrong, but that reserved memory is for the framebuffer? The iBoot bootloader also reserves some memory for the framebuffer. dGPUs bring their own VRAM because it's a different type of memory, allowing them to get higher performance than they could with DDR. The M4 Max requires 128GB of LPDDR5X to reach its ~500GB/s bandwidth. The RX Vega 64 had that same bandwidth in 2017 with just 8GB of HBM2.
- itamos 6mo agoOn one side it sounds promising to exploit shared memory properties to speed up inference. But on the other hand, the well established inference engines are perhaps already well optimized to overlap compute and communication efficiently. In this case the host-device copies are likely not a problem to tackle.
- adamsilvacons 6mo ago[dead]
- jedisct1 6mo agoDoesn't work on web browsers, only with one headless runtime, one one CPU architecture. What's even the point of using webassembly here?
- tancop 6mo agoloading third party agents in a sandbox with full custom model support. right now you need to either run that code directly (super dangerous) use a vm/container (slow and complicated) or a interpreter like lua (language bound, slow and weak security). wasm is perfect for this, its almost native speed, built for security and language neutral. onnx and coreml are secure but they can only do the actual model not all the code around it.
- agambrahma 6mo agoYes, that's the right idea. It's less about browsers, and more about server/edge/local-agent runtimes. Wasm lets you have - sandboxing (untrusted actor code) - clean snapshot/restore - portability of actor across machines If you don’t need those properties, then yes ... native is obviously the better choice