15 ms·
So, at 64-bit bus width, we're looking at 1 and 2 Tbps for PCIe 4 and 5, respectively. Looking forward to THAT!
by Keyframe 9y ago
So, at 64-bit bus width, we're looking at 1 and 2 Tbps for PCIe 4 and 5, respectively. Looking forward to THAT!
- cjensen 9y agoWhat do you mean by "64-bit bus width"? The most lanes you can get in PCIe is 16.
- spamizbad 9y agoMaybe a traditional "PCI-E card" formfactor, but I don't think there's anything in the specification preventing you from building a 32 or 64-lane device.
- kabdib 9y agoPCIe 3.x supports a power-of-two number of lanes up to 16; it handles all the gory stuff needed to aggregate those lanes. Going to more than 16 lanes means having multiple channels, and you'll have to worry about clock skew, delayed (retried) and failed transactions and other wonderful things you get with multiple independent pipes. I've never done it, but I'm interested in how hard it is. I know a guy who's done detailed design of PCIe hardware; I'll have to ask him.
- mechagodzilla 9y agoI’ve done a lot of PCIe hardware development. I think there was some notion of a 32x connection early on, spec wise, but x16 is as wide as it’s ever gone in practice. All high performance PCIe traffic is controlled by DMA hardware on the endpoint, so having an endpoint with multiple logical x16 connections back to the host cpu would be quite straightforward to implement, and you wouldn’t need to worry about any of the gory details you listed. Some FPGAs are available with multiple endpoints embedded in them, for instance.
- luckydude 9y agoCould you explain it like I'm five (or really like I'm a software geek and don't know this stuff) how the signals in the multiple lanes are synchronized? Is there some centralized clock that everyone is using?
- p_l 9y agoThe signals are skewed in predictable way, which allows recovering the ordering of transmission. IIRC, there's no common clock at all, each lane is "independent" and the ordering is applied on top.
- crzwdjk 9y agoFrom what I understand, at the physical level PCI-E uses an 8b10b encoding, which allows the clock to be recovered from the data. Each lane's clock might still be skewed with respect to other lanes, but there's some deskew logic that detects and corrects for this skew. I'm not 100% sure when the skew detection happens. It might be negotiated as part of link negotiation, or possibly it gets detected at the start of a packet, which has a special 8b10b symbol that's not used for data so it's easy to spot.
- kabdib 9y agoPCIe 3.0 actually uses a 128b to 130b encoding (I guess clock recovery got a lot better, but they also use different whitening patterns). Data is distributed across the lanes, bit by bit, and reassembled by the receiver. There's a bunch of training and link configuration stuff going on. Once you have multiple channels of these (because 16 lanes wasn't enough bandwidth) you have the possibility that one channel gets transaction errors while the other doesn't, resulting in skewed arrival or even just missing data. So you need a software layer to deal with this, it's no longer "plug something in and you get bus level reliability," instead it's more like networking, where you have to do some reassembly as data arrives, and probably embed sequence number or other information in the data so you can identify bad and missing transfers. (Or you could say, "screw it, it's reliable enough" and wait for customers to complain; if they never do, you either made the right engineering decision or your business isn't big enough yet :-) ).
- Keyframe 9y agoEach lane is a 8b/10b full duplex. 8x is 64 bits. 16x is 128 bits... PCIe goes up to 32x. If anything, I was conservative with going 8x.
- cjensen 9y agoThat's not how bus width works :-). Bus width is the number of physical wires in a bus. PCIe 3 is 128b/130b, and it would be wildly inaccurate to refer to an x16 as a "2048-bit bus."