4 ms·
imagining a simple x86/64 machine that also has a simple GPU... what does the amd64 assembly look like for "telling the GPU to do something"? Is there an amd6
by 2bitencryption 6y ago
imagining a simple x86/64 machine that also has a simple GPU...
what does the amd64 assembly look like for "telling the GPU to do something"?
Is there an amd64 instruction that says "Tell the GPU to start executing here?"
Or is it at the OS layer instead of the assembly layer, with an OS syscall that defers to the GPU?
- cylon13 6y agoCPU code calls in to drivers which talk to the GPU over the PCIe bus. So assembly code that does stuff with the GPU just looks like loading a shared library, passing in some shader code, and then passing in some drawing commands. This project looks like it aims to compile rust to a shader intermediate format so you can use it instead of more common shading languages like GLSL or HLSL. Basically your HLSL/GLSL/(now Rust) code ends up as an intermediate format, which the CPU passes in to a graphics driver before asking that driver to draw things.
- MayeulC 6y agoCPU peripherals are memory, so it just looks like copying a GPU "kernel" (program) from a memory area (from disk) to another (GPU, via PCIe). I don't know enough about PCIe to be an authoritative source, but I'll just go along and describe the way most microcontrollers and peripherals work. I doubt x86 has instructions dedicated to PCIe. There will likely be a PCIe controller memory-mapped to some hardware address, that the OS will write to as if it was regular memory (that address is physically wired to the controller, be it on the CPU die or on the motherboard). The address is generally specified by the BIOS/UEFI (devicetrees on most embedded platforms). Of course, what you will write will depend on the controller itself, but the OS will have a driver that contains that information. You have a back-and-forth to list peripherals on the bus, identify the one you want to communicate to, set-up access mode (direct memory access where you map an area of your memory to the peripheral, or just sequential access trough the controller). Once you can communicate with the GPU, you start the same dance to give it a program you compiled ahead of time, have it run, and fetch the result. I don't have low-level experience with that specific part of the system, but that kind of stuff is generic enough that you find the same patterns everywhere, at least for turing machines, which x86 CPUs are.
- olalonde 6y ago> for turing machines I suppose you meant "von Neumann machines" here?
- MayeulC 6y agoRight, sorry for the confusion. Though "machines" makes me think of the self-replicating machines concept, so "von neumann architecture" would be a better fit. https://en.wikipedia.org/wiki/Von_Neumann_architecture https://en.wikipedia.org/wiki/Von_Neumann_architecture
- wmf 6y agoIt's not (currently) an instruction. The key concepts at the hardware level are buffers, descriptors, and doorbell registers. These are accessed using mov instructions. https://compas.cs.stonybrook.edu/~nhonarmand/courses/sp17/cse506/slides/13-devices.pdf https://compas.cs.stonybrook.edu/~nhonarmand/courses/sp17/cs...
- dragontamer 6y agoThe OS level is purely so that multiple programs can talk to the same hardware at the same time. For example, your Firefox Browser may talk to the GPU to render something (H.264 video), or to send audio over the HDMI cable. But at the same time, you're playing some video game (Civ-6), and the OS needs to prioritize which gets access. ------------ When it comes to I/O, there are two strategies: memory mapped I/O, and I/O ports (the "in" and "out" assembly instructions). I'm pretty sure most modern hardware is memory-mapped type, where you just read to (or write from) certain "memory" locations. Its not really memory (it doesn't go to RAM), but instead to some other component, like PCIe root complex for further processing. ------ Something to note: the whole OSI level exists on the motherboard. You've got physical connections, you've got link-layer connections, and even a network similar to TCP/IP. In fact, there are multiple networks: SATA networks, PCIe networks, and USB Networks, that have their own addressing rules and physical protocols. Different USB Ports, different PCIe addresses, and more! From the perspective of the CPU, I expect any "GPU Command" to simply be a link-level call to the PCIe root complex. You send data to the PCIe controller, which then forwards the data to the GPU. Or NVMe RAM, or your network card. To fully make a GPU Command get to the GPU, you need to know the PCIe address, and route the message appropriately (there could be PCIe switches in between). Finally, there's a protocol, similar to application-level code (HTTP) where applications talk to the GPU. Given the similarities between DirectX, OpenCL, Vulkan, Metal, and CUDA/HIP, I expect that GPUs all have roughly the same "application" interface. Shaders are compiled into machine code. Machine code is loaded into GPU VRAM. Command Queues issue remote-function calls with an event-driven dependency graph that the GPU selects kernels from. Etc. etc.
- orbifold 6y agoThis is a bit of an oversimplification, because I don't actually know the details, but it works roughly like this: GPUs are devices on the PCI Express bus, without an operating system getting in the way everything on the PCI Express bus is organised into something called ECAM (https://en.wikipedia.org/wiki/PCI_configuration_space https://en.wikipedia.org/wiki/PCI_configuration_space). Each device there specifies memory regions it supports. These memory regions map to a large number of configuration registers and on-device memory the GPU supports. Since GPUs are really complex most of its behaviour is actually not directly controlled but instead predetermined by its firmware (basically an operating system in its own right). Finally GPUs really execute multiple instruction streams with proprietary instruction formats and can independently load data from main memory with their build in DMA engines. These operations are typically kicked off by pointing various registers on the GPU to memory regions in main memory to tell it where to fetch instructions and data from. Everything else happens for the most part as consequences of instructions executing on the GPU.
- Jasper_ 6y agoThis specific component is a SPIR-V backend for Rust. In practice, how you use this is you use a native API (Vulkan), give it your generated SPIR-V program, and it "runs it" using its own semantics. This is a combination of a userspace API talking to a kernel driver, and the kernel driver using a combination of IN/OUT instructions (unlikely) and memory-mapped I/O (much more likely) to talk to the GPU. The actual details of the communication is documented by some IHVs, either through PDFs or source code. Normally, the CPU doesn't wait for programs to complete, but instead, the GPU driver has a list of tasks to run (a "command buffer") that it itself can schedule across many different pieces of hardware, and a combination of the driver and GPU itself determine scheduling, execution, and so on. Note that there's still a lot of work for the driver to do, including managing the device's memory (textures, buffers, so on), compiling any programs into machine language (similar to a JIT), and coordinating multiple different programs trying to use the GPU.
- Const-me 6y ago> Or is it at the OS layer instead of the assembly layer, with an OS syscall that defers to the GPU? Yes. Modern OSes don’t allow userspace code to directly access peripheral devices, GPUs included. Due to the enormous complexity of modern GPUs (GPUs typically have more transistors, consume more electricity, and have comparable amount of memory), the API surface is huge. On Windows, the native API surface is Direct3D. It has two large pieces, user-mode and kernel mode. The vendor-provided GPU driver is similarly split in two. User-mode parts implement HLSL compiler (Microsoft), another downstream compiler to compile DXBC into proprietary byte code GPUs actually run (vendors), implement other higher-level stuff like mipmap generation (vendors, at least for nVidia), and expose user-facing APIs, i.e. multiple versions of D3D (Microsoft). Kernel mode driver interfaces with actual hardware, the Microsoft’s vendor-agnostic part of that is in dxgkrnl.sys. Various GPU drivers often expose extra user-facing APIs in addition to Direct3D (CUDA, Vulkan, OpenGL, OpenCL), but all of them are optional.
- nomel 6y ago> Modern OSes don’t allow userspace code to directly access peripheral devices, GPUs included I was under the impression that user-space drivers (like Linux UIO) were given an address that they can mmap for direct reads/writes from/to the peripherals address space. Is this not "direct"?
- Const-me 6y ago> were given an address that they can mmap for direct reads/writes from/to the peripherals address space Kernel-mode drivers indeed do that under the hood, but the majority of GPU I/O bandwidth doesn’t go that way. With DMA, GPUs have full access to system memory. They have specialized piece of hardware (exposed to programmers as copy command queue in D3D12, or transfer queue in Vulcan) to move large blocks of data in either direction between system RAM and VRAM.
- rrss 6y ago> Modern OSes don’t allow userspace code to directly access peripheral devices, GPUs included This depends a lot on the OS and on the peripheral, but this is certainly not true across the board. RDMA verbs, for example, definitely involve userspace directly communicating with the NIC, with no kernel involvement. At least some of the AMD GPU graphics stacks support userspace submission to the GPU also, see http://www.hsafoundation.com/html/Content/Runtime/Topics/02_Core/queues.htm http://www.hsafoundation.com/html/Content/Runtime/Topics/02_.... This requires hardware support, like an MMU and ensuring bad submissions from userspace can't do bad things to other processes, but lots of hardware supports these things. The windows graphics stack does require going into the kernel to submit work to the GPU, but that isn't the case for all operating systems and all peripherals. somewhat related: GPUs started getting support for virtual memory in ~2006 or so [1] but the windows graphics stack was improved to actually use that capability in Windows 10, ~10 years later [2]. [1] https://www.anandtech.com/show/2116/2 https://www.anandtech.com/show/2116/2 [2] https://en.wikipedia.org/wiki/Windows_Display_Driver_Model#WDDM_2.0 https://en.wikipedia.org/wiki/Windows_Display_Driver_Model#W...
- cesarb 6y agoSimplifying a lot, both the CPU and the GPU have access to a few regions of shared memory, which both can read from and write to. Some of it is on the GPU (device local memory) and some of it is on the CPU (host local memory). The CPU writes a list of commands to one of these shared memory regions, then writes to a special register on the GPU (these GPU registers are visible to the CPU as yet another region of shared memory) telling it where that list of commands is. That is, from a x86-64 assembly point of view, it's just writing to memory. Usually, only the final write to the special register is privileged, so that write will need an OS syscall into the graphics driver; the rest can be written directly by the userspace driver (which used beforehand another OS syscall to get access to some of that memory shared with the GPU). The userspace driver also compiles the SPIR-V programs into native GPU code, which can be loaded and executed as instructed by the list of commands.
- deleted 6y ago[deleted]