12 ms·
Ask HN: How does a CPU communicate with a GPU?
I've been learning about computer architecture [1] and I've become comfortable with my understanding of how a processor communicates with main memory - be it directly, with the presence of caches or even virtual memory - and I/O peripherals.
But something that seems weirdly absent from the courses I took and what I have found online is how the CPU communicates with other processing units, such as GPUs - not only that, but an in-depth description of interconnecting different systems with buses (by in-depth I mean an RTL example/description).
I understand that as you add more hardware to a machine, complexity increases and software must intervene - so a generalistic answer won't exist and the answer will depend on the implementation being talked about. That's fine by me.
What I'm looking for is a description of how a CPU tells a GPU to start executing a program. Through what means do they communicate - a bus? How does such a communication instance look like?
I'd love get pointers to resources such as books and lectures that are more hands-on/implementation aware.
[1] Just so that my background knowledge is clear: I've concluded NAND2TETRIS, watched and concluded Berkeley's 2020 CS61C and have read a good chunk of H&P (both Computer Architecture: A Quantitative Approach and Computer Organization and Design: RISC-V edition), and now am moving on to Onur Mutlu's lectures on advanced computer architecture.
- roschdal 5y agoThrough the electrical wires in the PCI express port.
- danielmarkbruce 5y agoI could be misunderstanding the context of the question, but I think OP is imagining some sophisticated communication logic involved at the chip level. The CPU doesn't know anything much about the GPU other than it's there and data can be sent back and forth to it. It doesn't know what any of the data means. I think the logic OP imagines does exist, but it's actually in the compiler (eg the cuda compiler), figuring exactly what bytes to send which will start a program etc.
- coolspot 5y agoNot in the compiler but in GPU driver. A graphic program (or compute) just calls APIs (DirectX/Vulkan/CUDA) of a driver, which then knows how to do that on a low-level writing to particular regions of RAM mapped to GPU registers.
- danielmarkbruce 5y agoYes! This is correct. My bad, it's been too long. I guess either way the point is that it's done in software, not hardware.
- lxgr 5y agoThere's also odd/interesting architectures like one of the earlier Raspberry Pis, where the GPU was actually running its own operating system that would take care of things like shader compilation. In that case, what's actually being written to shared/mapped memory is very high level instructions that are then compiled or interpreted on the GPU (which is really an entire computer, CPU and all) itself.
- alberth 5y agoNit pick… Technically it’s not “through” the electrical wires, it’s actually through the electrical field created around the electrical wires. Veritasium explains https://youtu.be/bHIhgxav9LY https://youtu.be/bHIhgxav9LY
- tux3 5y agoNitpicking the nitpick: the energy is what's in the fields, but the electrical wires aren't just for show, the electrons do need to be able to move in the wire for there to be a current, and the physical properties of the wire have a big impact on the signal. So things get very complicated and unintuitive, especially at high frequencies, but it's okay to say through the wire!
- a9h74j 5y agoAnd as you might be alluding, particularly high frequencies: in the skin (via skin effect) of the wire! I'll confess I have never seen a plot of actual rms current density vs radius related to skin effect.
- rayiner 5y agoTypically CPU and GPU communicate over the PCI Express bus. (It’s not technically a bus but a point to point connection.) From the perspective of software running on the CPU, these days, that communication is typically in the form of memory-mapped IO. The GPU has registers and memory mapped into the CPU address space using PCIE. A write to a particular address generates a message on the PCIE bus that’s received by the GPU and produces a write to a GPU register or GPU memory. The GPU also has access to system memory through the PCIE bus. Typically, the CPU will construct buffers in memory with data (textures, vertices), commands, and GPU code. It will then store the buffer address in a GPU register and ring some sort of “doorbell” by writing to another GPU register. The GPU (specifically, the GPU command processor) will then read the buffers from system memory, and start executing the commands. Those commands can include, for example, loading GPU shader programs into shader memory and triggering the shaders to execute those shaders.
- divbzero 5y agoGoing one deeper, how does the communication work on a physical level? I’m guessing the wires of the PCI Express bus passively propagate the voltage and the CPU and GPU do “something” with that voltage?
- _3u10 5y agoIt’s signaled similar to QAM. Far more complicated than GPIO type stuff. Think FM radio / spread spectrum rather than bitbanging / old school serial / parallel ports. Similar to old school modems if the line is noisy it can drop to lower “baud” rates. You can manually try to recover higher rates if the noise is gone but it’s simpler to just reboot.
- tux3 5y agoOh, that is several levels deeper! PCIe is a big standard with several layers of abstraction, and it's far from passive. The different versions of PCIe use a different encoding, so it's hard to sum it all up in a couple sentences in terms of what the voltage does.
- throw82473751 5y ago
- ar_te 5y agoAnd I you looking for some strange architecture forgoten by time:). https://www.copetti.org/writings/consoles/sega-saturn/ https://www.copetti.org/writings/consoles/sega-saturn/
- throwra620 5y ago
- throwmeariver1 5y agoEveryone in tech should read the book "Understanding the Digital World" by Brian W. Kernighan.
- dyingkneepad 5y agoIs this before or after they read Knuth?
- ncmncm 5y agoMmm, do you know anybody who has read Knuth, really?
- dyingkneepad 5y agoI own the books and I read parts of it. But regardless, my point was that there are many many things that people say every programmer should read, and that sometimes includes Knuth. In fact, so many things that even if I stop everything I'm doing and just keep reading everything "every programmer should read", I will die before I finish. tl;dr: it was a (bad) joke.
- arduinomancer 5y agoIs it very in-depth or more for layman readers?
- throwmeariver1 5y agoMost normal people and junior devs would get a red head when reading it, techies and seniors would nod along and sometimes say "uh... so that's how it really works". It's in between but a good primer on the essentials.
- pedrolins 5y agoIt doesn't really answer my question, but from what I've seen in the TOC I'd say it's equivalent to an introductory course on computer architecture + computer systems and some cryptography as well. Kind of an introduction (don't get me wrong with the word 'introduction', it covers a decent amount of material) to the most important concepts and technologies that guide computers and the internet.
- aliasaria 5y agoThere is some good information on how PCI-Express works here: https://blog.ovhcloud.com/how-pci-express-works-and-why-you-should-care-gpu/ https://blog.ovhcloud.com/how-pci-express-works-and-why-you-...
- thunkshift1 5y agoThis was good intro
- dyingkneepad 5y agoOn my system, the CPU sees the GPU as a PCI device. The "PCI config space" [0] is a standard thing and so the CPU can read it and figure out its device ID, vendor ID, revision, class, etc. From that, the OS looks at its PCI drivers and tries to find which one claims to drive that specific PCI device_id/vendor_id combination (or class in case there's some kind of generic universal driver for a certain class). From there, the driver pretty much knows what to do. But primarily the driver will map the registers to memory addresses, so accessing offset 0xF0 from that map is equivalent as accessing register 0xF0. The definition of what each register does is something that the HW developers provide to the SW developers [1]. Setting modes (screen resolution) and a lot of other stuff is done directly by reading and writing to these registers. At some point they also have to talk about memory (and virtual addresses) and there's quite a complicated dance to map GPU virtual memory to CPU virtual memory. On discrete GPUs the data is actually "sent" to the memory somehow through the PCI bus (I suppose the GPU can read directly from the memory without going through the CPU?), but in the driver this is usually abstracted to "this is another memory map". On integrated systems both the CPU and GPU read directly from the system memory, but they may not share all caches so extra care is required here. In fact, caches may also mess the communication on discrete graphics, so extra care is always required. This paragraph is mostly done by the Kernel driver in Linux. At some point the CPU will tell the GPU that a certain region of memory is the framebuffer to be displayed. And then the CPU will formulate binary programs that are written in the GPU's machine code, and the CPU will submit those programs (batches) and the GPU will execute them. These programs are generally in the form of "I'm using textures from these addresses, this memory holds the fragment shader, this other holds the geometry shader, the configuration of threading and execution units is described in this structure as you specified, SSBO index 0 is at this address, now go and run everything". After everything is done the CPU may even get an interrupt from the GPU saying things are done, so they can notify user space. This paragraph describes mostly the work done by the user space driver (in Linux, this is Mesa), which implements OpenGL/Vulkan/etc abstractions. [0]: https://en.wikipedia.org/wiki/PCI_configuration_space https://en.wikipedia.org/wiki/PCI_configuration_space [1]: https://01.org/linuxgraphics/documentation/hardware-specification-prms https://01.org/linuxgraphics/documentation/hardware-specific...
- simne 5y agoLot of things happen there. But most important, PCIe bus is serial bus, which have virtualized interface, so there is no physical process of communication, what happen more similar to Ethernet network, mean on each device exists few endpoints, each has it's own controller with its own address and few registers to store state and transitions, and memory buffer(s). Videocards usually have many behaviors. In simplest modes, they behave just as RAM mapped to large chunk of system RAM space, plus video registers to control video output, and to control address mapping of video ram, and to switch modes. In more complex modes, Videocards generate interrupts (just special type of message on PCIe). In 3D modes, which are most complex, Videocontroller take data from its own memory (which mapped to system space), there are stored tree of graphic primitives, some draw directly from videoram, but for others used bus master option of PCIe, in which videocontroller read additional data (textures) from predefined chunks of system RAM. About GPU operation, usually, CPU copy data to Videoram directly, than ask videocontroller to run program in videoram, and when complete, GPU issue interrupt, and than CPU copied result from videoram. Recent additions where, add GPU possibility to read data from system disks, using mentioned before bus master, but those additions are not already wide implemented.
- simne 5y agoFor beginner, I think the best to begin read about Atari consoles, Atari-65/130, NES, as their ideas where later implemented in all commodity videocards, just slightly extended. BTW all modern videos use bank-switching.
- xg15 5y ago> Recent additions where, add GPU possibility to read data from system disks, using mentioned before bus master, but those additions are not already wide implemented. My impression is that high-end graphics cards (Nvidia RTX 30x and professional equivalents) more and more replicate parts of the PC architecture and become sort of mini-computers within a computer. Following that logic, I wonder when we'll see the first card with its own dedicated flash memory - or why not a PCIe controller, so you can hook op an SSD...
- simne 5y ago
- dragontamer 5y agoI'm no expert on PCIe, but its been described to me as a network. PCIe has switches, addresses, and so forth. Very much like IP-addresses, except PCIe operates on a significantly faster level. At its lowest-level, PCIe x1 is a single "lane", a singular stream of zeros-and-ones (with various framing / error correction on top). PCIe x2, x4, x8, and x16 are simply 2x, 4x, 8x, or 16 lanes running in parallel and independently. ------- PCIe is a very large and complex protocol however. This "serial" comms can become abstracted into Memory-mapped I/O. Instead of programming at the "packet" level, most PCIe operations are seen as just RAM. > even virtual memory So you understand virtual memory? PCIe abstractions go up to and include the virtual memory system. When your OS sets aside some virtual-memory for PCIe devices, when programs read/write to those memory-addresses, the OS (and PCIe bridge) will translate those RAM reads/writes into PCIe messages. -------- I now handwave a few details and note: GPUs do the same thing on their end. GPUs can also have a "virtual memory" that they read/write to, and translates into PCIe messages. This leads to a system called "Shared Virtual Memory" which has become very popular in a lot of GPGPU programming circles. When the CPU (or GPU) read/write to a memory address, it is then automatically copied over to the other device as needed. Caching layers are layered on top to improve the efficiency (Some SVM may exist on the CPU-side, so the GPU will fetch the data and store it in its own local memory / caches, but always rely upon the CPU as the "main owner" of the data. The reverse, GPU-side shared memory, also exists, where the CPU will communicate with the GPU). To coordinate access to RAM properly, the entire set of atomic operations + memory barriers have been added to PCIe 3.0+. So you can perform "compare-and-swap" to shared virtual memory, and read/write to these virtual memory locations in a standardized way across all PCIe devices. PCIe 4.0 and PCIe 5.0 are adding more and more features, making PCIe feel more-and-more like a "shared memory system", akin to cache-coherence strategies that multi-CPU / multi-socket CPUs use to share RAM with each other. In the long term, I expect Future PCIe standards to push the interface even further in this "like a dual-CPU-socket" memory-sharing paradigm. This is great because you can have 2-CPUs + 4 GPUs on one system, and when GPU#2 writes to Address#0xF1235122, the shared-virtual-memory system automatically translates that to its "physical" location (wherever it is), and the lower-level protocols pass the data to the correct location without any assistance from the programmer. This means that a GPU can do things like perform a linked-list traversal (or tree traversal), even if all of the nodes of the tree/list are in CPU#1, CPU#2, GPU#4, and GPU#1. The shared-virtual-memory paradigm just handwaves the details and lets PCIe 3.0 / 4.0 / 5.0 protocols handle the details automatically.
- justsomehnguy 5y agoTL;DR: bi-directional memory access with some means to notify the other part about "something has changed". It's not that different for any other PIC/E device, be it a network card or a disk/HBA/RAID controller. If you want to understand how it came to this - look at the history of ISA, PCI/PCI-X, a short stint for AGP and finally PCI-E. Other comments provides a good ELI15 for the topic. A minor note about "bus" - for PCEe it is mostly a historic term, because it's a serial, P2P connection, though the process of enumerating and qurying the devices is still very akin to what you would do on some bus-based system, e.g.: SAS is a serial "bus", compared to SCSI, but still you operate with it as some "logical" bus, because it is easier for humans to grok it this way.
- pedrolins 5y agoI find it very interesting that you mention looking at the history of ISA's first in order to understand the current iteration of the technology. I was reading the RISC-V privileged ISA recently and the amount of seemingly arbitrary registers and behaviours that must be implemented to support a UNIX-like OS is crazy, and that got me thinking about the history behind all of these things that the hardware must support in order to support the OS. But thank you for the pointers, I'll definitely use this.
- justsomehnguy 5y agoHa! I definitely meant ISA bus, as others had mentioned. Kudos to swetland and phendrenad2!
- phendrenad2 5y agoNot that ISA: :) https://en.wikipedia.org/wiki/Industry_Standard_Architecture https://en.wikipedia.org/wiki/Industry_Standard_Architecture
- swetland 5y agoThe "ISA" mentioned above is the "Industry Standard Architecture", the 8/16bit bus used by PCs and PC clones back in the day, not "Instruction Set Architecture (x86, ARM, RISC-V, etc): https://en.wikipedia.org/wiki/Industry_Standard_Architecture https://en.wikipedia.org/wiki/Industry_Standard_Architecture
- derekzhouzhen 5y agoOther has mentioned MMIO. MMIO has several kinds: 1. CPU accessing GPU hw with uncache-able MMIO, such as lower level register access 2. GPU accessing CPU memory with cache-able MMIO, or DMA. such as command and data stream 3. CPU accessing GPU memory with cache-able MMIO, such as textures They all happen on the bus with different latency and bandwidth.
- pizza234 5y agoYou'll find a very good introduction in the comparch book "Write Great Code, Volume 1", chapter 12 ("Input and Output"), which also explains the history of system buses (therefore, you'll find an explanation of how ISA works). Interestingly, there is a footnote explaining that "Computer Architecture: A Quantitative Approach provided a good chapter on I/O devices and buses; sadly, as it covered very old peripheral devices, the authors dropped the chapter rather than updating it in subsequent revisions."
- pedrolins 5y agoWow. Just skimmed across that chapter and that looks like a great resource. No wonder I couldn't find it in any of my searching sessions, I'd never think a book titled like that would cover hardware concepts so extensively. This will definitely help me in understanding buses better. Thank you.
- zoenolan 5y agoOther are not wrong in saying Memory mapped IO. taking a look at the Amiga hardware Reference manual [1] and a simple example [2] or a NES programming guide [3] would be a good way to see this in operation. A more modern CPU/GPU setup is likely to use a ring buffer. The buffer will be in CPU memory. That memory is also mapped into the GPU address space. The Driver on the CPU will write commands into the buffer which the GPU will execute. These will be different to the shader unit instruction set. Commands would be setting some internal GPU register to a value. Allowing the setting resolution, framebuffer base pointer, set up the output resolution, setting the mouse pointer position, reference a texture from system memory, load a shader, execute a shader, set a fence value (Useful for seeing when a resource, texture, shader is no longer in use). Hierarchical DMA buffers are a useful feature of some DMA engines. You can think of them as similar to sub routines. The command buffer can contain an instruction to switch execution to another chunk of memory. This allows the driver to reuse operations or expensive to generate sequences. OpenGL's display list commonly compiled down to separate buffer. [1] https://archive.org/details/amiga-hardware-reference-manual-3rd-edition https://archive.org/details/amiga-hardware-reference-manual-... [2] https://www.reaktor.com/blog/crash-course-to-amiga-assembly-programming/ https://www.reaktor.com/blog/crash-course-to-amiga-assembly-... [3] https://www.nesdev.org/wiki/Programming_guide https://www.nesdev.org/wiki/Programming_guide
- melenaboija 5y agoIt is old and I am not sure everything still applies but I found this course useful to understand how GPUs work: Intro to Parallel Programming: https://classroom.udacity.com/courses/cs344 https://classroom.udacity.com/courses/cs344 https://developer.nvidia.com/udacity-cs344-intro-parallel-programming https://developer.nvidia.com/udacity-cs344-intro-parallel-pr...
- brooksbp 5y agoWoah there, my dude. Let's try to understand a simple model first. A CPU can access memory. When a CPU performs loads & stores it initiates transactions containing the address of the memory. Therefore, it is a bus master--it initiates transactions. A slave accepts transactions and services them. The interconnect routes those transactions to the appropriate hardware, e.g. the DDR controller, based on the system address map. Let's add a CPU, interconnect, and 2GB of DRAM memory: +-------+ | CPU | +---m---+ | +---s--------------------+ | Interconnect | +-------m----------------+ | +----s-----------+ | DDR controller | +----------------+ System Address Map: 0x8000_0000 - 0x0000_0000 DDR controller So, a memory access to 0x0004_0000 is going to DRAM memory storage. Let's add a GPU. +-------+ +-------+ | CPU | | GPU | +---m---+ +---s---+ | | +---s------------m-------+ | Interconnect | +-------m----------------+ | +----s-----------+ | DDR controller | +----------------+ System Address Map: 0x9000_0000 - 0x8000_0000 GPU 0x8000_0000 - 0x0000_0000 DDR controller Now the CPU can perform loads & stores from/to the GPU. The CPU can read/write registers in the GPU. But that's only one-way communication. Let's make the GPU a bus master as well: +-------+ +-------+ | CPU | | GPU | +---m---+ +--s-m--+ | | | +---s-----------m-s-----+ | Interconnect | +-------m----------------+ | +----s-----------+ | DDR controller | +----------------+ System Address Map: 0x9000_0000 - 0x8000_0000 GPU 0x8000_0000 - 0x0000_0000 DDR controller Now, the GPU can not only receive transactions, but it can also initiate transactions. Which also means it has access to DRAM memory too. But this is still only one-way communication (CPU->GPU). How can the GPU communicate to the CPU? Well, both have access to DRAM memory. The CPU can store information in DRAM memory (0x8000_0000 - 0x0000_0000) and then write to a register in the GPU (0x9000_0000 - 0x8000_0000) to inform the GPU that the information is ready. The GPU then reads that information from DRAM memory. In the other direction, the GPU can store information in DRAM memory, and then send an interrupt to the CPU to inform the CPU that the information is ready. The CPU then reads that information from DRAM memory. An alternative to using interrupts is to have the CPU poll. The GPU stores information in DRAM memory and then sets some bit in DRAM memory. The CPU polls on this bit in DRAM memory, and when it changes, the CPU knows that it can read the information in DRAM memory that was previously written by the GPU. Hope this helps. It's very fun stuff!
- chubot 5y agoBTW I believe memory maps are set up by the ioctl() system call on Unix (including OS X), which is kind of a "catch all" hole poked through the kernel. Not sure about Windows. I didn't understand that for a long time ... I would like to see a "hello world GPU" example. I think you open() the device and the ioctl() it ... But what happens when things go wrong? Similar to this "Hello JIT", where it shows you have to call mmap() to change permissions on the memory to execute dynamically generated code. https://blog.reverberate.org/2012/12/hello-jit-world-joy-of-simple-jits.html https://blog.reverberate.org/2012/12/hello-jit-world-joy-of-... I guess one problem is that this may be typically done in vendor code and they don't necessarily commit to an interface? They make you link their huge SDK
- kllrnohj 5y agoThe OSDev Wiki is a great resource on how this all works from the perspective of actually programming it at least on x86 For example here's the page on talking PCI-E https://wiki.osdev.org/PCI_Express https://wiki.osdev.org/PCI_Express
- phendrenad2 5y agoAt a high level, it's actually really simple. Your PCIe devices are each given a region of the address space, say, 0x8428000000000000-0x8428000000000fff. Just write to that region from kernel mode. But what do you write? Well, that isn't standardized. It's not even really documented. The best documentation is the source code to the GPU drivers in the Linux kernel, which are usually added to by engineers working at GPU vendors, and they don't discuss it much.
- account42 5y agoAMD does have some GPU register documentation for GCN at the bottom of https://developer.amd.com/resources/developer-guides-manuals/ https://developer.amd.com/resources/developer-guides-manuals... but not for RDNA / RDNA2.
- ncmncm 5y agoWhile we're here: is there any reasonable prospect of keeping one's GPU from being able to read and write to literally anywhere in physical memory? I.e., a practical way a kernel and driver might be able to forward to the GPU only commands and shaders that can access only your process memory, and nobody else's, and your process's pixels, and no other process's pixels, when they live in GPU RAM? For all I know, this is the norm for all GPUs, but I wonder why it is hard, then, for VMs to share a GPU.
- kllrnohj 5y ago> I wonder why it is hard, then, for VMs to share a GPU. It isn't. Nvidia & AMD just charge a massive premium for the privilege. Nvidia calls it vGPU https://docs.nvidia.com/grid/13.0/grid-vgpu-user-guide/index.html#vgpu-use-introduction https://docs.nvidia.com/grid/13.0/grid-vgpu-user-guide/index... and AMD calls it MXGPU https://www.amd.com/en/graphics/workstation-virtual-graphics https://www.amd.com/en/graphics/workstation-virtual-graphics Both have been around for a while now, and both refuse to bring it to their consumer cards.
- brooksbp 5y ago> is there any reasonable prospect of keeping one's GPU from being able to read and write to literally anywhere in physical memory? This is the purpose of an IOMMU. > I.e., a practical way a kernel and driver might be able to forward to the GPU only commands and shaders that can access only your process memory, and nobody else's, and your process's pixels, and no other process's pixels, when they live in GPU RAM? So IOMMU and the GPU's MMU. > For all I know, this is the norm for all GPUs What? > but I wonder why it is hard, then, for VMs to share a GPU. Engineering is hard.
- cesarb 5y ago> What I'm looking for is a description of how a CPU tells a GPU to start executing a program. Through what means do they communicate - a bus? How does such a communication instance look like? For most modern computers, through the PCI Express bus. Take a look at the output of "lspci -v" and you'll see something like: 00:02.0 VGA compatible controller: [...] [...] Flags: bus master, fast devsel, latency 0, IRQ 128 Memory at ee000000 (64-bit, non-prefetchable) [size=16M] Memory at d0000000 (64-bit, prefetchable) [size=256M] I/O ports at f000 [size=64] Expansion ROM at 000c0000 [virtual] [disabled] [size=128K] That is, the GPU on this particular laptop makes available a region of memory sized 16 megabytes at physical address 0xee000000, and another region of memory sized 256 megabytes at physical address 0xd0000000. Whenever the CPU writes to or reads from these memory regions, it is writing to memory on the GPU, not on the normal RAM chips. And not all of that "memory" on the GPU is real memory; some of it are registers, which are used to control the GPU. The same happens on the opposite direction: for code running on the GPU, some regions of memory are actually the RAM normally used by the CPU. In either case, the memory read and/or write transactions go through the PCI Express bus to the other device. The exact details of what is written to (and read from) that memory vary depending on the device. For most GPUs, the driver sets up a list of commands in memory (either "host" memory, which is the RAM on the CPU, or "device" memory, which is the RAM on the GPU accessible through these PCI Express "memory windows"), and writes the address of that command list to a register on the GPU; the GPU then reads the list and executes the commands found in it. These commands can include things like "start N threads of the program found at X with Y as the input" (GPU programs are commonly called "shaders", and they are highly parallel), but also things like "wait for event W to happen before doing Z".
- rasz 5y agoOn the PC side start by reading some basics like https://archive.org/details/URP_8th_edition/ https://archive.org/details/URP_8th_edition/ (never editions require logging in and borrowing) >What I'm looking for is a description of how a CPU tells a GPU to start executing a program. Through what means do they communicate - a bus? How does such a communication instance look like? Long time ago you would memory map the framebuffer and just write directly to it. Then first 2D acceleration showed up in 1987 in form of IBM 8514 (later cloned by ATI/Matrox/S3/Tseng and others). You wrote commands one at a time using I/O port access to FIFO with pooling for idle/full, no direct access to the framebuffer http://www.os2museum.com/wp/the-8514a-graphics-accelerator/ http://www.os2museum.com/wp/the-8514a-graphics-accelerator/ Next evolution was MMIO - memory mapped IO. You no longer executed dedicated CPU IO instruction (assembler IN/OUT), IO ports were simply addresses in memory. You still had FIFOs and wrote one command at a time http://www.o3one.org/hwdocs/video/voodoo_graphics.pdf http://www.o3one.org/hwdocs/video/voodoo_graphics.pdf Then someone threw DMA into the mix. Now you could DMA contents of a circular buffer filled with your commands http://www.bitsavers.org/components/s3/DB019-B_ViRGE_Integrated_3D_Accelerator_Aug1996.pdf http://www.bitsavers.org/components/s3/DB019-B_ViRGE_Integra... We finally got command list/command buffer/bundle copied directly to the GPU. Nowadays you have multiple command lists/command buffers/bundles going in parallel https://developer.nvidia.com/blog/advanced-api-performance-command-buffers/ https://developer.nvidia.com/blog/advanced-api-performance-c... On a hardware side 8/16 bit ISA bus was a shared parallel connection to CPU bus at fixed clock (4.77-10MHz, 4 clocks per transfer, ~5MB/s max speed). It took us up to 1992 to get the next commonly used solution, a "rogue" consortium of companies tired of IBM shit designed VESA Local Bus (a true hack) in form of slapping expansion cards direct on the raw 32bit CPU bus of 486 processors. Cheap, no licensing fees, extremely fast (40MHz x 32bit = potentially faster than later PCI), easy to implement. This got replaced with the advent of Pentium (64bit external CPU data bus) and introduction of PCI. PCI is still a shared parallel bus, but this time 32bits at 33MHz with packetized transactions. AGP was "just" a faster PCI on its own dedicated separate controller (no contention with other PCI devices) and optimized addressing (sideband). 32bit at 66MHz, then x2 DDR, x4 QDR, x8 ODR. Last one means there are 8 transfers taking place between one clock cycle for a nice 2GB/s. PCI-E is faster bidirectional serial point-to-point PCI with ability to combine links into bundles (x1-x16). PCI-E devices live on a network switch and dont block each other from talking simultaneously. You could think of PCI-E as every PCI device getting its own dedicated dual direction AGP connector. Some vintage hands on coding examples: 2D Tseng Labs ET4000 coding https://www.youtube.com/watch?v=K8kZ4BFxOtc https://www.youtube.com/watch?v=K8kZ4BFxOtc 2D Cirrus Logic https://www.youtube.com/watch?v=WoAE7x-u1g0 https://www.youtube.com/watch?v=WoAE7x-u1g0 "How 3D acceleration started 20 years ago: S3/Virge register level programming" https://www.youtube.com/watch?v=fXJ11_wG_0U https://www.youtube.com/watch?v=fXJ11_wG_0U "Acceleration code working on real S3 Virge/DX" https://www.youtube.com/watch?v=Hsg1N4IqXac https://www.youtube.com/watch?v=Hsg1N4IqXac "Direct hardware accelerated 3d in 20kB code" https://www.youtube.com/watch?v=n509_wN02u8 https://www.youtube.com/watch?v=n509_wN02u8 "Bare metal hardware 3d texturing in 23kb of code w/ S3/Virge" https://www.youtube.com/watch?v=UgvBGXiw6LY https://www.youtube.com/watch?v=UgvBGXiw6LY "Testing our latest low-level hardware 3d code on real S3/Virge hardware" https://www.youtube.com/watch?v=px--LWdRoYA https://www.youtube.com/watch?v=px--LWdRoYA "Live coding and testing more low-level 3D w/ S3/Virge" https://www.youtube.com/watch?v=l3lH0cIZUSA https://www.youtube.com/watch?v=l3lH0cIZUSA "Finishing low-level hardware S3/Virge acceleration demo" https://www.youtube.com/watch?v=JmfeB2LEDbc https://www.youtube.com/watch?v=JmfeB2LEDbc "3dfx Voodoo: Low-level & bare-metal driver-less code" https://www.youtube.com/watch?v=LDT6KlfOG2k https://www.youtube.com/watch?v=LDT6KlfOG2k "Finally 3dfx Voodoo triangles" https://www.youtube.com/watch?v=ZWaDqY4gqhw https://www.youtube.com/watch?v=ZWaDqY4gqhw "More GPU programming Voodoo case study" https://www.youtube.com/watch?v=AYZvNyxFHqk https://www.youtube.com/watch?v=AYZvNyxFHqk "Quite final 3dfx Voodo low-level code working" https://www.youtube.com/watch?v=2ADQgIEWrx4 https://www.youtube.com/watch?v=2ADQgIEWrx4
- Randolf_Scott 5y agoDrivers make all hardware communicate.