7 ms·
TL;DR: Fastly needs full routing tables on all CDN nodes in order to determine which is the best transit path to push out content through. In order to save mone
by amazon_not 10y ago
TL;DR: Fastly needs full routing tables on all CDN nodes in order to determine which is the best transit path to push out content through. In order to save money, they used a programmable Arista switch instead of a traditional router. Their solution is to reflect BGP routes via the switch to the nodes and and fake direct connectivity between the nodes and the transit provides, so that nodes can directly push out content to whichever transit provider they determine is best on a per packet basis.
Please correct me if I'm wrong.
Maybe I'm just obtuse, but I found the blog post confusing about what it really is about and long winded, taking a very long time to come to the point. There is a severe lack of context in the beginning, it would have tremendously benefited from the what and the why of what they are trying to do.
- edwhitesell 10y agoI'd bet the best transit path isn't the primary reason. Having the routes in the application means it choose a path based on cost, customer preference or any number of other business rules.
- scurvy 10y agoRoutes aren't exposed to customers as a configurable option. At least not at my pay grade with them.
- bogomipz 10y agoBut it is, you can probe and route around brownouts in certain AS's in the path. I would say that a CDNs ability to get to the content to the eyeballs is the most important and trumps all else.
- edwhitesell 10y agoI think cost is most important. I know other CDNs use routes for managing costs too. Being able to modify traffic flows means you control the flow of money between peers and/or customer=>transit relationships.
- amazon_not 10y agoCost is unlikely to be most important. If it were the CDN could just buy the cheapest transit, accept a default route and call it a day. In the article Fastly has 13 transit links at one site and the ability to ping each route to determine the best path. They would hardly go to all this trouble if cost control was the main objective.
- Swannie 10y agoAs it's a CDN, they peer. Look up the AS in the interconnect DBs. Lots of peering.
- bogomipz 10y agoCost is important to any business, but CDNs are all about speed, the biggest metric is "time to first byte." This is always going to take precedence over cost for a CDN. Differences in transit costs for a large CDN with high commits is negligible.
- edwhitesell 10y agoThere are more costs involved than just those of the CDN. Let's say I'm a network operator. I have some number of POPs/facilities, peering relationships, transit customers maybe even some relationships where I pay for transit. If I decide to partner with or become a customer of a CDN who can control the path traffic takes to reach the CDN on my network, I/they/we can use that negate, or elminate, my transit costs AND drive revenue to me. (very simple example) Lets say one of my transit customers also has a non-paid peering relationship with one of my upstreams. Customer may choose to take the longer path to reach the CDN on my network because it's free to them, rather than paying me. Now, not only are they not paying me for transit, but I may be incurring additional cost to pay my transit. If this routing data is in the CDN application, a logical step is to report on costs of the traffic. I would see this traffic taking a more costly path and be able to act on it. Maybe that means pre-pending my ASN, maybe that means modifying the DNS requests made from my customer's IP range(s) to use a block not routed across the other ASN. Maybe it means I simply don't want that traffic to my "nodes" anymore. Maybe all anyone needs is "fast" today, I agree its not the only factor, but being able to manage cost was a discussion point in the past. EDIT: last line for clarity
- Swannie 10y agoYou're not alone in finding the post confusing. For me the TL;DR would be: Our edge nodes originally had full BGP tables from our upstreams. This was good as our nodes care about path selection. When we needed to scale, we didn't want to use a pair of (expensive) routers, because we didn't actually need routing, but our upstream were not happy about having to peer to more than a pair of devices. We bought some switches, reflected BGP routes to the nodes, and had to hack L2 connectivity between the switch and the nodes. Also, using FIB to describe RIB irks me.
- bogomipz 10y agoI think he was talking about FIB on the edge device not on the hosts, so that would be correct usage.
- scurvy 10y agoI'm also really confused as to why they insisted on taking full tables at their edge if they also built their own route optimization technology. It seems like it would be much simpler and faster if they just took a default route from each provider and then did route probing on their top 30% prefixes in each PoP. You don't need full edge routes for that. You need a single route reflector (or two for redundancy) to tell you what routes are in the table. Then probe the busiest prefixes (as determined by flow data) via each provider and take the best route. Localpref the best route higher in the route reflector and send it to the Linux hosts. No need for a huge FIB on the edge. This is basically what Noction IRP/Internap FCP platforms already do. This reminds me of the scene in Primer where they ask if they're doing it for a reason or just showing off. FWIW, the Netflix networks just take a default route from a pair of transit networks in each PoP. At their scale, they found it better to negotiate better worldwide transit deals than optimize networking at the PoP level. Different game than Fastly, but another data point.
- bogomipz 10y agoI remember those FCP boxes, its been a while. I think its maybe what they call Miro now, do you know? The Avaya route science boxes did the same thing. It seems both of these have fallen out of usage. I always thought it was odd that there wasn't an opensource project for this. I know Internap originally did it in PERL.
- scurvy 10y agoYep, FCP's successor is Miro. I haven't had a chance to look into it, but their sales people claim it's all new. Noction has pretty much taken over the entire market. Overall their product is good, but the Internet game has changed a lot since the original FCP days (lots more direct peers, lots more exchanges, lots more route servers, lots more paid peering), and I don't think that things like FCP/Miro/IRP do well once you move past paid peering. Things just become too complex. IRP exposes some of this to you, but the implementation is kludgey (routing instance per peer, DSCP bits to identify peers on exchanges, and more).
- 10y ago