6 ms·
You're thinking of a very narrow use-case for VxLAN. What happens when your VM needs to talk to legacy networking gear? What happens when your VM(s) need to t
by Nrsolis 11y ago
You're thinking of a very narrow use-case for VxLAN.
What happens when your VM needs to talk to legacy networking gear? What happens when your VM(s) need to talk to a legacy security infrastructure that's mandated because of regulatory concerns? What happens when you need to route between two VxLAN domains?
You're going to need to provide the VxLAN VTEP functionality someplace and that requires hardware support in the network someplace. Waving your hands and saying "that's not a concern" won't cut it. I have a customer that's facing this issue RIGHT NOW.
And no, MPLS isn't a proprietary standard. It's open and it's available on oodles of networking hardware. It's interoperable and it's a proven technology that scales.
Developing a tunneling protocol is the easy part. Developing a scalable control plane to handle RIB and FIB state is another matter entirely. Getting all of that to operate on a chipset that can do 2-4Tbps per slot is another thing altogether.
Go check out some of the HUGE cloud infrastructure guys. They aren't using BrandX networking hardware to build their data-centers. They're using the Big-3 players.
- jpgvm 11y agoNo, they are not using the Big-3 players. ISPs are using Alcatel and Juniper still sure, big enterprise shops are still swearing by Cisco. Big cloud infrastructure has all moved to Linux on merchant silicon, either stuff like Quanta, Pico8 or even more DIY like OCP. As for terminating VxLAN on legacy gear, yeah that is probably a bad idea and someone probably made a bad decision to end up with an architecture that requires that. As for control plane. Most people using VxLAN at scale have their own control planes. They are actually stupidly easy to build because the edge is so easy to work with. I built my own that integrates with the Linux native VxLAN implementation using netlink to program the forwarding table. At the end of the day VxLAN isn't a great replacement for MPLS but it's a great encapsulation system for fully software defined datacenters that have already invested into a full control plane for all compute and networking (think Mesos, Kubernetes). It's also a good choice for smaller scale stuff because it scales down nicely. Multicast forwarding might suck in a real DC situation but for a lot of newbies playing around it's a good way to get into doing L2oL3. I don't think the WAN cases matter that much. At the point where it's a problem you have the people around to make said problem go away. Specifically trying to do cross DC IP address mobility is dumb in the first place. Most models that actually work well cross datacenter simply terminate it in one place and bring it up the workload on an entirely new set of resources in the new DC. This is much easier and is shared nothing usually (except maybe the dataset, but that is usually stored in S3/HDFS/Ceph/other distributed store here). Long story short, it does it's job fine. Use it for something it wasn't built for and yes it will hurt you.
- Nrsolis 11y agoYou're downvoting me but you're so wrong on so many points I don't know where to begin. 1. Cisco, Juniper, Alcatel, Arista all have merchant silicon platforms. They marry those to their own control planes because the customers want that kind of continuity and support. Pico8 and others are a FRACTION of the market place for switching (including datacenter switching). My own company has so many racks full of gear at AMZN and GOOG that it's hard to understand why you feel like you know that infrastructure better. 2. Writing a scalable control plane isn't as easy as you are representing. GOOG took many years to write theirs and it's still problematic for them in anything in lots of their use cases. I don't know why you think that quite literally the entire Internet along with all of the protocol work that has built it is a drag on novel architectures. 3. Mesos/Kubernetes are vanishingly small parts of the larger ecosystems out there. You can't just wave you hands and disregard quite literally the billions of dollars of infrastructure that are deployed each and every year by major service providers and corporations. Tiny cloud providers are not the largest slice of the pie when it comes to dollars spent on networking gear right now. 4. "At the point where it's a problem you have the people around to make said problem go away." HUH???? I can think of maybe 8 or 9 use-cases off hand that absolutely REQUIRE this kind of functionality because billions of dollars of transactions are handled by the infrastructure within and between those datacenters. Sorry buddy. You're way wrong here. Either you've never built an infrastructure that anyone cares about losing for an hour or two while you figure out what went wrong or you're so tanked up on "cloud" kool-aid that you've forgotten how we got to this point and have failed to understand how large systems scale up from small ones. Show me a large multinational bank that's storing their transaction data on S3 or Ceph and I'll show you a bank that's not "systemically important". Most large enterprises don't have the luxury of simply discarding 100% of their working, proven architectures on the promise of a few small startups hoping to cash in big. Things like operational stability, redundancy, fault isolation, and monitoring are not "nice to haves", they are mission critical requirements. They aren't up to the whim of some bright-eyed CTO, but the watchful eye of umpteen nations of regulators.
- parasubvert 11y agoGOOG took many years to write theirs and it's still problematic for them in anything in lots of their use cases. I think that's a creative interpretation of the facts. GOOG is operating at enormous scales pushing the limits of operational knowledge. I think it's quite acceptable and natural that vendors are packaging up their learnings from 2007 into products now for most enterprises (which are starting to really eat them up). "Mesos/Kubernetes are vanishingly small parts of the larger ecosystems out there. You can't just wave you hands and disregard quite literally the billions of dollars of infrastructure that are deployed each and every year by major service providers and corporations. Tiny cloud providers are not the largest slice of the pie when it comes to dollars spent on networking gear right now." While I agree with your broader point about VxLAN v. MPLS (I think), the above isn't even wrong. Mesos/Kube aren't vanishingly small, that would imply they're shrinking, rather than small startups/projects that are growing at astonishing rates. You're also confusing billions of dollars of low-mid margin hardware with potential billions of dollars of mid-high margin software that's aiming at IBM, HP, CA, Oracle, and Microsoft's application servers and management tooling. Secondly, Mesos/Kube aren't cloud providers, they're the startups that represent the next generation (along with Cloud Foundry, OpenShift, and whatever Docker comes up with) of data center operating systems that are going to run the bulk of enterprise systems the way VMware does today. That said, there's a belief that all of this requires SDN/Overlay Networking like NSX or VxLAN that will magically fix network problems by bundling it with the app platform and waving a wand. Here I agree ... they won't. The secret behind good software defined networking is solid hardware defined networking ;) Sorry buddy. You're way wrong here. Either you've never built an infrastructure that anyone cares about losing for an hour or two while you figure out what went wrong or you're so tanked up on "cloud" kool-aid that you've forgotten how we got to this point and have failed to understand how large systems scale up from small ones. I dunno. Stepping back, I remember James Hamilton from Amazon at re:invent clearly was aiming directly at network vendors as the last bastion of costly proprietary mainframe thinking that will be commodified by software-defined cloud services on commodity hardware. It will take time. But they're pretty jazzed about it. "Show me a large multinational bank that's storing their transaction data on S3 or Ceph and I'll show you a bank that's not "systemically important". Ceph, I agree. Amazon OTOH has won the object storage game. S3 is the de facto API for all object storage now, whether it's from EMC, NetApp, etc. And mission critical banks are definitely using it, at humungous scale. I have no idea why you'd think S3 is appropriate for transactional data, it's an object store.