4 ms·
And if you had a tree based lookup that requires 4 lookups, but the first 2 levels are in L1 cache (suppose you do a byte trie, that's only 64k + 256 bytes), th
by waps 12y ago
And if you had a tree based lookup that requires 4 lookups, but the first 2 levels are in L1 cache (suppose you do a byte trie, that's only 64k + 256 bytes), then have the next levels in small, "cheap", separate memory ?
If you had 100M of L2 cache, or 100M of SRAM, you could fit the third level in what amounts to L2 cache. You'd still have (at most) 1 memory lookup for any route (1 lookup in SRAM, 1 in RAM, worst case), but you'd need something like 256M of ram instead of 16 gigs. Vast majority of packets (going to < /24 routes) would be routed in < 10 ns, and the 130 ms becomes a 140ms worst-case. With average < 10ns routing suddenly having one of these processors for, let's say, 160 Gbit worth of ports becomes possible. And I'm sure your customers would immediately ask you to put 320G of ports on that processor.
(Principle of a trie is that IP address is a.b.c.d. So you build an array so that route_table[a][b][c][d] yields the next hop. route_table[a][b][c] happens entirely in cache + SRAM, since it only requires 100M)