5 ms·
If you add multiple GPUs to the same PCIe switch, their bandwidth to main memory of the SOC will be shared. For some use cases that bandwidth was already the bo
by rerx 3y ago
If you add multiple GPUs to the same PCIe switch, their bandwidth to main memory of the SOC will be shared. For some use cases that bandwidth was already the bottle neck.
- detourdog 3y agoUsing the older PCI bus every three or 4 slots represented a dedicated controller. We would install the high bandwidth cards on seperate controllers. PCI-E exists in a faster timing bus which may make a difference. The visionOS demo apparently shows sub 20 nanosecond processing and graphics blitting to appear imperceptibly. Finally Woz’s design for the Apple II color graphics was based on unified memory. Dramatic shifts in technology sometimes obviates all our past learning and experience.
- c00lio 3y agoThe defining characteristic of PCIe isn't faster bus timings. It is that PCIe isn't even a "bus" in the traditional sense. With a bus, you have a number of shared physical wires for addresses and data, and an arbitration system that assigns those shared wires to one device at a time. The available bandwidth is just the bandwidth of the bus, divided into the time you can reserve it for yourself. PCIe is a tree-shaped network (root is the root hub in the CPU, inner nodes are switches, leafs are devices) with point-to-point links (called lanes) that transmit packets. Bandwidth is determined by the number of lanes that a device's path up to the root hub has at the narrowest node.
- detourdog 3y agoThat sounds like a good description but do you think the term serial should be worked in there?
- c00lio 3y agoOn some level yes, on some level no. The physical interface for each point-to-point link is serial, but one device can have several of those links and use them (send/receive packets on them) in parallel.
- detourdog 3y agoDoes the lane count correspond to the quantity of parallel links? would the switches be single transistors?
- c00lio 3y agoYes, lane count is the number of parallel links. Switches are far more complex. The protocol specifies addressing, bandwidth reservations, flow control, priorities and message ordering requirements, which the switches have to handle. Not a switch thing per se, but also relevant for "single transistors won't do": On the physical layer, each port/link is also relatively complicated to initialize and operate. The transmission encoding, due to the relatively high speeds involved, requires a learning phase where both ends measure the link characteristics and adjust their (relatively complex) signal processing accordingly. The physical layer also handles different link speeds for down-/upward compatibility.
- detourdog 3y agoThis sounds like an ideal config for unified memory. I suppose that one can count all the lanes to figure out the actually total throughput. Does every vendor come up with their own implementation of a "busmaster"? This may explains how they achieve the graphics blitting in the Vision Pro.