3 ms·
The default partitioning of CAM space on Cisco gear is the obvious issue, but the root cause is the massive deaggregation of announced IPv4 routes on the Intern
by rwg 12y ago
The default partitioning of CAM space on Cisco gear is the obvious issue, but the root cause is the massive deaggregation of announced IPv4 routes on the Internet. You can see various statistics about this problem at http://www.cidr-report.org/ http://www.cidr-report.org/, but the short of it is that if the top 30 networks (based on announced route savings) completely aggregated their announcements as much as possible, ~41,000 routes (~8% of the routing table) would be eliminated.
And that's just the top 30 networks — if every network cleaned up their announcements, it would eliminate ~232,000 routes (~45% of the table).
Adding to the deaggregation problem is the inability to easily filter out route announcements based on RIR minimum allocations without having to add tons of exceptions for CDNs that operate as islands of connectivity and carve out IP space for each island from a single address space allocation. (There's no covering route for the islands of connectivity since these CDNs have no "backbone" connecting the islands, so if you filter out those smaller announcements, you lose connectivity to those islands.)
There are many people who think this problem will just magically go away as IPv6 adoption increases, but all increased IPv6 adoption will do is make limited CAM space even more limited as network engineers have to balance dividing precious CAM space between a ballooning-quickly IPv4 route table and a ballooning-slightly-less-quickly IPv6 route table.
(To be clear: I think ubiquitous, functioning, end-to-end native IPv6 connectivity needs to happen sooner than later, but it's not a magic bullet for the Internet's technical problems.)
- sp332 12y agoThis is one of the reasons IPv6 addresses are being given out in such massive allotments. ICANN doesn't want anyone to have to get multiple "chunks" and blow up the routing table later.
- MichaelGG 12y agoSo people multi-homing (which is the only reason I end up getting and announcing /24s) is not a significant amount of the route table? Cause IPv6 can't really fix that part (multihoming) can it?
- cnvogel 12y agoBut it solves the problem of you, after some growth, having received several distinct v4/24s, instead of one huge v6/64...
- waps 12y agoYou're assuming that just getting a massive opportunity for more de-aggregation won't cause de-aggregation out of shear laziness and convenience. It's much easier to de-aggregate than it is to aggregate. The maximum IPv4 de-aggregation possible is 2^24 - 2 ^ 21. The maximum IPv6 de-aggregation with what's currently being handed out (mostly between /32 and /48's) is way, way, way more. So IPv6 has more addresses, yes. Just so long as you don't actually use them. The problem of course is that the memory requirements for IPv4 assignments were going up linearly. If you bought gear with a good amount of memory, you could therefore expect it to last a few years. Clearly the network vendors that designed IPv6 saw this as a problem ... how can we make it explode ? Well IPv6 was the answer. Problem is that they went completely batshit insane overboard. "If" IPv6 deaggregates we'll need routers with about 2^(48-24) TIMES more memory. Storing a deaggregated IPv6 routing table (which has to be in memory) requires 524288 terabytes of memory. The rest of IPv6 isn't much better. It's some academics wish-list, with total disregard for real-world concerns. It is not possible to use 1/10th of the IPv6 features. Multicast ? Won't work (on any public or even just somewhat large network). Encryption ? Won't work (too many devices don't support it). Anycast ? Won't work. Site-local anycast ? Won't work. Larger address space ? Won't work (in half the world). NAT-avoidance ? Won't work (didn't get past security engineers). Autoconfiguration ? Actually kinda handy in some scenarios, but again, won't work on most networks, where you fall back to the IPv4 mechanism. Faster routing due to "smarter" headers ? Doesn't work according to my load stats. Better Qos ? Doesn't work in either of the major routing gear vendors. Mobility ? Let's not go there. Instead of implementing sane "live" re-addressing they went with ... Aargh. Let's not go there. Automatic network renumbering ... riiiight. Heh. I wonder if this was put in as a joke. There are known solutions to all these problems (well, except the multicast and anycast ones), but of course, the IPv6 designers knew better. Welcome to the "solution". On the other hand, solving the solution pays rather well.
- MichaelGG 12y agoIs CAM space a truly constrained resource, as-in, is Cisco really on the cutting edge to deliver whatever amounts of CAM they ship? Or is it similar to their "software" routers, where it's nearly entirely marketing/product differentiation driven? Is adding more CAM as hard as Intel shipping 64GB of L2 cache, for instance?
- ChuckMcM 12y agoIt is interesting that IP addresses are only 32 bits so you can provide a route for every individual IP address in a 16GB chunk of memory. You can build a 'next hop' appliance with an FPGA and a couple of DDR3 chips such that any 32 bit number input would return the next hop in 130ns after loading. It would be straight forward to create command port that took a network + netmask and filled in all the next hop info for every host on that network. And if you wanted to be cheesy you could re-use the 10/8 and the 192.168/16 spaces for various control changes. But it is still a lot of network activity and IPV6 is just around the corner so probably not worth building today.
- waps 12y ago> It is interesting that IP addresses are only 32 bits so you can provide a route for every individual IP address in a 16GB chunk of memory ... 130ns. 1) you'd need at least 1 byte to identify a next-hop, right. So we're talking 8 bytes * 2^32 = 32 gigabyte, not 16 2) Luckily IPv6 fixed that. For IPv6 you'd need millions of terabytes to do this trick. 3) At, say, 64 10Gbit ports and 130ms lookups you'd need a mere 215 of these memories to actually be able to forward line rate. 4) Network devices are currently moving up to 40G and 100G ports in the top models. 8*100G forwarding cards exist today.
- ChuckMcM 12y agoThe way I was imagining how you would do this is that you would take the 32 bit IP address, shift it left by two, adding it to the base address of your 16GB of memory block and reading the 32 bits that were there. Those 32 bits being the IP address of the next hop. I used a similar 'the address is an operand' trick to create a very fast 8x8 multiply on an 8 bit Z80, arg 0 was the upper 8 bits arg 1 was the lower 8 bits which was applied to A0-15 of a 64K x 16 EPROM, the contents of the eprom at any given location was the product of arg0 and arg1. That allowed me to save enough time to respond to packets coming in at 38400 baud over an FM side band at the time.