3 ms·
> It is interesting that IP addresses are only 32 bits so you can provide a route for every individual IP address in a 16GB chunk of memory ... 130ns. 1) you'd
by waps 12y ago
> It is interesting that IP addresses are only 32 bits so you can provide a route for every individual IP address in a 16GB chunk of memory ... 130ns.
1) you'd need at least 1 byte to identify a next-hop, right. So we're talking 8 bytes * 2^32 = 32 gigabyte, not 16
2) Luckily IPv6 fixed that. For IPv6 you'd need millions of terabytes to do this trick.
3) At, say, 64 10Gbit ports and 130ms lookups you'd need a mere 215 of these memories to actually be able to forward line rate.
4) Network devices are currently moving up to 40G and 100G ports in the top models. 8*100G forwarding cards exist today.
- ChuckMcM 12y agoThe way I was imagining how you would do this is that you would take the 32 bit IP address, shift it left by two, adding it to the base address of your 16GB of memory block and reading the 32 bits that were there. Those 32 bits being the IP address of the next hop. I used a similar 'the address is an operand' trick to create a very fast 8x8 multiply on an 8 bit Z80, arg 0 was the upper 8 bits arg 1 was the lower 8 bits which was applied to A0-15 of a 64K x 16 EPROM, the contents of the eprom at any given location was the product of arg0 and arg1. That allowed me to save enough time to respond to packets coming in at 38400 baud over an FM side band at the time.
- waps 12y agoAnd if you had a tree based lookup that requires 4 lookups, but the first 2 levels are in L1 cache (suppose you do a byte trie, that's only 64k + 256 bytes), then have the next levels in small, "cheap", separate memory ? If you had 100M of L2 cache, or 100M of SRAM, you could fit the third level in what amounts to L2 cache. You'd still have (at most) 1 memory lookup for any route (1 lookup in SRAM, 1 in RAM, worst case), but you'd need something like 256M of ram instead of 16 gigs. Vast majority of packets (going to < /24 routes) would be routed in < 10 ns, and the 130 ms becomes a 140ms worst-case. With average < 10ns routing suddenly having one of these processors for, let's say, 160 Gbit worth of ports becomes possible. And I'm sure your customers would immediately ask you to put 320G of ports on that processor. (Principle of a trie is that IP address is a.b.c.d. So you build an array so that route_table[a][b][c][d] yields the next hop. route_table[a][b][c] happens entirely in cache + SRAM, since it only requires 100M)