6 ms·
You're downvoting me but you're so wrong on so many points I don't know where to begin. 1. Cisco, Juniper, Alcatel, Arista all have merchant silicon platforms.
by Nrsolis 11y ago
You're downvoting me but you're so wrong on so many points I don't know where to begin.
1. Cisco, Juniper, Alcatel, Arista all have merchant silicon platforms. They marry those to their own control planes because the customers want that kind of continuity and support. Pico8 and others are a FRACTION of the market place for switching (including datacenter switching). My own company has so many racks full of gear at AMZN and GOOG that it's hard to understand why you feel like you know that infrastructure better.
2. Writing a scalable control plane isn't as easy as you are representing. GOOG took many years to write theirs and it's still problematic for them in anything in lots of their use cases. I don't know why you think that quite literally the entire Internet along with all of the protocol work that has built it is a drag on novel architectures.
3. Mesos/Kubernetes are vanishingly small parts of the larger ecosystems out there. You can't just wave you hands and disregard quite literally the billions of dollars of infrastructure that are deployed each and every year by major service providers and corporations. Tiny cloud providers are not the largest slice of the pie when it comes to dollars spent on networking gear right now.
4. "At the point where it's a problem you have the people around to make said problem go away." HUH???? I can think of maybe 8 or 9 use-cases off hand that absolutely REQUIRE this kind of functionality because billions of dollars of transactions are handled by the infrastructure within and between those datacenters.
Sorry buddy. You're way wrong here. Either you've never built an infrastructure that anyone cares about losing for an hour or two while you figure out what went wrong or you're so tanked up on "cloud" kool-aid that you've forgotten how we got to this point and have failed to understand how large systems scale up from small ones.
Show me a large multinational bank that's storing their transaction data on S3 or Ceph and I'll show you a bank that's not "systemically important". Most large enterprises don't have the luxury of simply discarding 100% of their working, proven architectures on the promise of a few small startups hoping to cash in big.
Things like operational stability, redundancy, fault isolation, and monitoring are not "nice to haves", they are mission critical requirements. They aren't up to the whim of some bright-eyed CTO, but the watchful eye of umpteen nations of regulators.
- parasubvert 11y agoGOOG took many years to write theirs and it's still problematic for them in anything in lots of their use cases. I think that's a creative interpretation of the facts. GOOG is operating at enormous scales pushing the limits of operational knowledge. I think it's quite acceptable and natural that vendors are packaging up their learnings from 2007 into products now for most enterprises (which are starting to really eat them up). "Mesos/Kubernetes are vanishingly small parts of the larger ecosystems out there. You can't just wave you hands and disregard quite literally the billions of dollars of infrastructure that are deployed each and every year by major service providers and corporations. Tiny cloud providers are not the largest slice of the pie when it comes to dollars spent on networking gear right now." While I agree with your broader point about VxLAN v. MPLS (I think), the above isn't even wrong. Mesos/Kube aren't vanishingly small, that would imply they're shrinking, rather than small startups/projects that are growing at astonishing rates. You're also confusing billions of dollars of low-mid margin hardware with potential billions of dollars of mid-high margin software that's aiming at IBM, HP, CA, Oracle, and Microsoft's application servers and management tooling. Secondly, Mesos/Kube aren't cloud providers, they're the startups that represent the next generation (along with Cloud Foundry, OpenShift, and whatever Docker comes up with) of data center operating systems that are going to run the bulk of enterprise systems the way VMware does today. That said, there's a belief that all of this requires SDN/Overlay Networking like NSX or VxLAN that will magically fix network problems by bundling it with the app platform and waving a wand. Here I agree ... they won't. The secret behind good software defined networking is solid hardware defined networking ;) Sorry buddy. You're way wrong here. Either you've never built an infrastructure that anyone cares about losing for an hour or two while you figure out what went wrong or you're so tanked up on "cloud" kool-aid that you've forgotten how we got to this point and have failed to understand how large systems scale up from small ones. I dunno. Stepping back, I remember James Hamilton from Amazon at re:invent clearly was aiming directly at network vendors as the last bastion of costly proprietary mainframe thinking that will be commodified by software-defined cloud services on commodity hardware. It will take time. But they're pretty jazzed about it. "Show me a large multinational bank that's storing their transaction data on S3 or Ceph and I'll show you a bank that's not "systemically important". Ceph, I agree. Amazon OTOH has won the object storage game. S3 is the de facto API for all object storage now, whether it's from EMC, NetApp, etc. And mission critical banks are definitely using it, at humungous scale. I have no idea why you'd think S3 is appropriate for transactional data, it's an object store.
- Nrsolis 11y agoRight now most of my focus is in the financial industry. I would say "in the US" but the truth is that my clients are multi-national behemoths. I'm regularly on plane flights to EMEA and APAC. Believe me when I tell you that they are still running systems that were around in the 70's. They have a significant investment in code that can only properly run in a mainframe environment and isn't going to get thrown away anytime soon. This is not to say that they don't have any interest in "cloud" technologies...quite the contrary....they are deploying just about ALL of them: ESX, OpenStack, Cloudstack, etc. But what often emerges as a barrier to deployment is the operational details: things like upgrades to infrastructure, minimization of downtime, security, and integration with the rest of the network and computing infrastructure. They don't have the luxury of starting from scratch and they certainly can't just forget about how to make things work with their larger infrastructure. Does OpenStack even have a way to upgrade from Juno to Kilo with ZERO downtime? Questions like that are a huge part of the testing and design that go into their thinking. These guys spend $1B EVERY YEAR on computing. And here is ANOTHER barrier: they can't readily do business with startups. It just doesn't work for them. The possibility that a critical part of their infrastructure is dependent on the fortunes of a group of maybe 50 people being successful. And it's not enough to say "well they have the source code" and can support it themselves. That doesn't work for them when the auditors come out and need to identify WHO is responsible for taking care of support and the lifecycle of the code. They write the code for their applications; you can't expect them to code big parts of their OS too. SO....please take my comments in the spirit they are intended. You can't be successful unless you are able to sell you solutions to the broader market that includes lots of customers that aren't GOOG or AMZN. Don't forget staffing either. If your infrastructure requires a CS PhD to support/upgrade, you're going to have a hard time selling it out there. Handling of outages tends to be business-specific so NoOPS style models don't work everywhere. It's fine for FB, but not NYSE. As for storage, EMC is the standard. Transactional databases are in far more places than you'd expect. If your app has scaling issues with accessing a non-virtualized database or datastore, then that's going to be a problem for you. If your OS can't handle redundant datapaths or confuses the Ops people about which piece of physical hardware is causing the issues you're seeing at the virtual layer, then you're going to have even more problems. So I'm not drinking the Kool-Aid just yet. I love virtual infrastructure but there are still too many open questions that need answers before it's going to be a complete solution for mission critical stuff.
- jpgvm 11y agoI'm not the one down-voting you. I see your points and they are valid I just don't necessarily agree with them. 1. Yes they do, unfortunately said "marrying" of them to their own management platforms makes them un-necessarily difficult to work with. This is fine when you have an army of network engineers or you have bought into their entire proprietary system (still usually needs army of network engineers) but it's a pain otherwise. There are plenty of other viable Broadcom Trident based gear etc out there that is used at huge scale (see Quanta) that if you use aftermarket firmware actually works really well. Bing was built using vanilla Quanta gear, I am not sure about Azure but I would assume similar. 2. I think you are misconstruing my words. I wrote my own platform for my purposes. It scales well enough for what I do and that is great. I am not GOOG and am not pretending to be, at their scale all problems become immensely more complicated. They do however have immensely more resources to throw at said problem. The revealing point however is they actually went that path though. I am sure if buying a bunch of vendor gear and going with MPLS would have fixed their problem then that is what they would have done. 3. Mesos and Kubernetes are just 2 implementations of a much more wide spread cluster job scheduling idea that has been around for a very very long time. The specific examples mean very little, the point is more that they change the paradigm in which you operate your datacenter. You no longer think about specific hardware and networks being assigned to specific departments or projects, instead everything is scheduled on one pool of resources and SLAs etc are all built into the control plane. Once you have this for compute it follows the same will be done for networking using the same datamodel and infrastructure. My point here is this is why GOOG, AMZN and MSFT either build their own gear or go the more open route, you can't do this unless you are able to hook into your network plane at a low enough level. 4. I though this point was pretty straight forward. If you have a massive WAN with tons of datacenters and loads of compute workload that needs cross DC mobility you probably also have people smart enough (or the money to get them) to come up with a good architecture that solves the problem in a reasonable way. The more important point which you seem to have missed is that cross DC IP mobility is probably what people have been trying to achieve and that is likely the problem. If you stop trying to able to shift IPs anywhere you want and embrace the fact that workloads are dynamic and IPs will come and go then suddenly things are a whole lot easier, more scalable and more resilient. If anything you agree with my point by confirming billions of dollars of transactions are involved. I am not tanked up on any cloud "kool-aid", I have just refused to subscribe to enterprise architectures that have never made sense and never will. If you actually follow evolution of systems I think my argument is even more compelling, not less. They don't need to. Their large HDS, EMC or other overpriced SAN is a completely reasonable replacement for S3 or Ceph, I just used those as examples. Indeed they are, that is why it's so important that the network plane is open. So that tooling can be built or integrated that provides all of those things. While it's not then you are constantly hitting up against the age old vendor argument of "That feature is not currently implemented" or "We don't expose those counters" or "We realise this violates the spec but we can't make a breaking change to fix it". As for regulators, I am yet to meet one that could even follow this conversation. Maybe your experience has been better but I wouldn't say "regulator friendly" is a good thing on the balance of probabilities. As for your other comments about how much money is spent on these sorts of platforms vs existing enterprise and carrier datacenters I completely agree, it is tiny. However what I disagree with is that is a meaningful metric. The trends are pretty clear, computing is becoming commoditised to the point that it's not going to make sense to do it in-house any more. When that time comes it's not those old and crufty enterprise models that are going to be left standing.