3 ms·
In my last team, our team suffered from a CPU bug and a SSD bug that caused painful debugging nights. We used multithreading heavily, and it could have been lib
by nemonemo 7y ago
In my last team, our team suffered from a CPU bug and a SSD bug that caused painful debugging nights. We used multithreading heavily, and it could have been library/kernel issues or some bugs in our code, which cover massive amount of potential target code that needed to sort through initially. I don't think small companies can afford such gigantic debugging effort, so either they should rely on their luck (or flaky tests) or on someone else who can afford it (a cloud/infra service provider or a startup that depends on you.)
- lima 7y agoHow is that related to bare metal vs. clouds? That could have happened on either.
- nemonemo 7y agoBig cloud providers or some big VC backed companies have the resource to figure out the issues by wrestling with the gigantic and unyielding vendors, or by working around it with their engineering resource. Not many teams have the time for even wanting to understand the issues. This is not an inherent problem of bare metal, but one aspect to consider if you were to choose the path. But you are right, cloud providers have their own quirks.
- nocman 7y ago> Big cloud providers or some big VC backed companies have the resource to figure out the issues by wrestling with the gigantic and unyielding vendors, or by working around it with their engineering resource. This is all operating on the assumption that those entities will care enough about your problem enough to do something meaningful to fix it. As others have stated here, at least in the context of "Big cloud providers", that frequently is not the case. Often you just get the runaround (continuous delays, requests for "more information", other attempts to stall in first level support). Add a third party "managing" some piece of your company's infrastructure via a cloud provider into the mix and it often gets even worse. It is true that there are real costs to having your own hardware/software onsite and people who know how to manage it. However, the promised reductions in cost/hassle of moving things offsite are frequently offset or even exceeded by the costs/hassles you get by not having control of things yourself.
- Nullabillity 7y agoCloud providers also tend to run much more exotic configs for pretty much everything (since tenant isolation is a top priority), so such issues are more likely to happen in the first place.
- syn0byte 7y agoExcept they won't. Not for an esoteric or exotic bug that doesn't affect 99.9% of their customers. Unless your businesses contract costs more than that team of technical and engineering resources(Eg your Netflix et al), it's not worth their time to bother looking into it. All that assumes they even admit the issue is their hardware not your code, which is a mighty big assumption itself.
- turk73 7y agoBecause when you're caught between those who manage the hardware and those who manage the K8s cluster, you get ping-ponged between them in the blame shifting. It's very annoying.