14 ms·
General guidance when working as a cloud engineer
- lockedinspace 4y agoA helpful list of things to have in mind when working with anything tech related.
- myfirstproject 4y ago> Git should be your only source of truth. Discard any local files or changes, what's not pushed into the repository, does not exist. Completely agree with that.
- pondidum 4y agoWhat about secrets? I like to do short lived credentials using Vault (e.g. vault can create say db access credentials dynamically), but for things like API keys where I can't do that..? Is the Vault KV store the source of truth?
- rexarex 4y agoWe utilize version control for config/secret management as well…encrypted of course. Edit: now that I think of it, for generated short lived passwords we also use SSM but for anything set by a human it’s in version control…
- lr4444lr 4y agoConfig that should be pushed into the env: it's not code or assets.
- adra 4y agoAt least in cloud providers, they have secret vaults accessible to their customers. The individual secrets are stored in source code but they're encrypted. We've used SOPS as a valuable way to manage these secrets. You can certainly stand up your own secretserver or equiv but may not have all the same.integratuon bells and whistles.
- intelVISA 4y agoGit feat. Nix = no more worries. Ever! Probably.
- lofatdairy 4y agoi think notion uses nix in production. I can't remember what their ci/cd pipeline and version control system is outside of that, or if it was even mentioned in that one comment i saw about it
- f4c39012 4y ago> should Completely agree > only Fine, but can substitute "git" as appropriate > discard any local files or changes Ok for when deployment is completely and always automated, but for that _one special case_ maybe keep a copy of the old that you can revert to until you're _really_ sure of no unwanted effects. In the meantime, find out how & why that local change got made and what can be done to automate it next time
- martynvandijke 4y agoNice guide, just curious are there more of these guides ?
- mustafabisic1 4y agoSome solid career advice in there as well. I feel like this could used as one of those "How to 10x career" articles - and be better than all of them.
- WolfOliver 4y ago"Microservices should only perform a single task." -> I guess this advice is the reason there are so widely misunderstood, see: https://linkedrecords.com/challenging-the-single-responsibility-principle-9800f39c186f https://linkedrecords.com/challenging-the-single-responsibil...
- adamisom 4y agoWow and I thought functions should only perform a single task. I need to keep up with the times! Apparently you need an entire deployable app and API to do anything these days. I guess it makes sense. How else could we justify so many software engineers!?
- elric 4y agoSo many? Last I checked there was a huge shortage, and with the exception of a couple of notable bloatware companies, most seem to be understaffed?
- dagss 4y agoUnderstaffed because with the microservices fashion you can get 100 engineers to be satisfied and think they are doing meaningful work while producing the output of 10 engineers. The industry believes the fashion is the only way to work by now, and shortage results. (Of course, there are several other ways to get 100 engineers to have the output of 10; but I think microservices makes the engineers a lot less frustrated when it happens compared to more traditional alternatives)
- water-your-self 4y ago>but I think microservices makes the engineers a lot less frustrated when it happens compared to more traditional alternatives. Can you explain this reasoning?
- dagss 4y ago
- pondidum 4y ago> Do not make production changes on Fridays I ~hate~ dislike this advice. If you can't deploy on a Friday, you need to fix your deployment strategy. By removing Friday from when you can deploy, you're wasting 1/5 of your available days. Note: deploy != Release[1]. Use flags, canaries etc. [1]: https://andydote.co.uk/2022/11/02/deploy-doesnt-mean-release/ https://andydote.co.uk/2022/11/02/deploy-doesnt-mean-release... Edit: hate is far too stronger word for this
- lopatin 4y agoInterestingly my company only deploys on Friday because it has to wait for (most) markets to close for the weekend.
- dopylitty 4y agoThis one made me laugh. I've been places that only allow deployments on Fridays because it gives the whole weekend to fix things if they break. It's a good interview question as a candidate. If you ask the interviewer when they deploy and they say only Friday (or worse only once a month) then perhaps look elsewhere for your own sanity because it's a sign of serious malfunction either organizationally, technically, or both.
- fragmede 4y agoDepending on your role, that is. If your desired position is straight dev with minimal to no ops work as possible, then yeah, red flag. However, if you're an SRE/DevOps-type person, setting up a continuous deployment system so they can deploy more often than that is a perfect landing task to dig your teeth into. Different strokes for different folks.
- kator 4y agoDon't forget "pets vs cattle", thinking of servers as ephemeral and working towards quickly being able to scale up/down based on demand. So often I see people "lift and shift" from a dedicated server model into the cloud and never convert their pets into cattle. This reduces flexibility later, not to mention makes it harder to respond to patching needs, scaling, and moving to optimize latency or costs.
- candiddevmike 4y agoCitation needed? There are tradeoffs to both, one is not always better than the other.
- paulryanrogers 4y agoWhat's the advantage of pets? Simplicity?
- gtirloni 4y agoA 128-core 4TB "pet" is much easier to manage than the equivalent Kubernetes cluster with same capacity. Not saying it's the way to go for every situation.
- paulryanrogers 4y agoDo you mean bare metal?
- jdub 4y agoIt should still not be administered as a pet. In fact, it's even more important when you have a single instance of some importance to make it entirely rebuildable and replaceable.
- vanviegen 4y agoIf you manage to run your entire service on a single (heavy weight) server, uptimes will usually be excellent. Reliability for a single server is high. It gets lower for each server you add, unless you're adding redundancy. But that adds complexity, and that hurts reliability too. My point: a single dedicated server may be a more reliable, simpler and much cheaper solution than the cloud provides.
- elric 4y ago> Certify yourself with official courses. Can anyone recommend some certifications that are worthwhile? I realize that this is a very broad ask, but the advise is also rather broad.
- eikenberry 4y agoJust about any Certificate is worthwhile depending on your reasons. Best case I've seen them used for is to help you break into new technology areas, EG. you want to work as an SRE for AWS services, having a few AWS Certificates under your belt might be just enough to get you that interview (plus you'd kill at AWS trivia nights).
- intelVISA 4y agoLove AWS trivia. Most effective way to harm your employer's wallet? EKS?? EKS? EC2? EBS..? ELB..? Ah no way it was /data egress/ of course.
- SquibblesRedux 4y agoI have always wondered -- what makes egress expensive? Are they just trying to keep everything in-house, or is there some real cost to egress?
- oneepic 4y agoI'd absolutely go to an AWS trivia night/lunch hour at work. Maybe GCP? Azure?
- nijave 4y ago>Before jumping straight into a new technology, read and understand their docs The number of issues I've seen that turn out to be documented features... (or, more accurately, things just being configured incorrectly)
- birdymcbird 4y ago> A good monitoring system, well-organized repository, fault-tolerance workloads and automation mechanisms are the basis of any architecture. Monitoring/alarming, and knowing what to monitor. Also, properly instrument your services or whatever it is you have. Take time to reflect on what are the signals that tell you operational health. An error metric alone is useless if you don’t know the denominator. Also be careful to avoid adding noisy metrics that cause panic for no reason. I’m not sure what fault tolerance means in this context. Very handwavy statement. I think if you have dependencies, have a plan and understanding of which ones tipping over will bring down your service or how you can build resiliency. For example, some feature on your page requires talking to a recommendations service. If the service goes down, can you call back to a generic list of hard coded recommendations or some static asset? As for automation: yeah, have test workflows built into your CI/CD harness. And avoid manual steps there requiring human intervention. Use canaries to test certain functions are up and running as expected, etc
- lockedinspace 4y agoMaybe I was a bit vague in the fault tolerance statement. What I mean is to have a high availability in your services, e.g: using AWS ASG for your servers, having more than one replica for your Kubernetes pods. If one of the servers/pods fail, the process behind detects them as "unhealthy" (having a nice monitoring/alarming as you mentioned) and replaces them with a new server with the same software characteristics so, for the end-user, SO, your client. Nothing has changed, the load just moved to a single instance for about 5-10 mins until a new server was deployed.
- birdymcbird 4y agoCool makes sense
- abledon 4y ago> If you need to build an architecture which involves microservices, I am sure that your cloud provider has a solution that fits better than Kubernetes. E.g: ECS for AWS. Thank you! So many people running unnecessary things on Kubernetes
- rswail 4y agoOn the other hand, K8S provides you with orchestration abstraction across AWS, GCP, Azure, VMWare, bare metal. There are distinct advantages to that in terms of both development (running a local K8S cluster is relatively easy) and deployment. ECS has no distinct advantages over K8S (or EKS in AWS land). Particularly now that there are CRDs for K8S that allow you to deploy AWS functionality (eg ALBs, TGs) from K8S.
- raxits 4y agoOne more Have a good logging & rollback strategy well communicated across stakeholders
- raydiatian 4y ago> If you need to build an architecture which involves microservices, I am sure that your cloud provider has a solution that fits better than Kubernetes. E.g: ECS for AWS. Kubernetes is a fantastic toolkit, but only shines when all that it has to offer, gets used. As far as FAAS goes, I think more people need to go check out Cloud Run as a Knative implementation. Having used it for sometime now it feels like a near-perfect FAAS solution. The only gripe I have is that versioning is a bit dopey. But hey, if I can have autoscaling services with absolute impunity over how my HTTP interface is shaped (looking at you AWS lambda) and without needing to worry about Kubernetes headaches, I’m perfectly happy to embed version names in service domains.
- gizzlon 4y agoAgreed,but would call Cloud Run a Papas not a FaaS
- raydiatian 4y agoPapas=PaaS, or? If Papas I am unfamiliar with the term. If PaaS, isn’t Gcloud itself the PaaS? For instance cloud run, the product inside Gcloud, is ephemeral and stateless, which wouldn’t be at all good for trying to make a DB.
- gizzlon 4y agoHehe,yes,sorry. F@$# autocorrect I believe PaaS predates FaaS, there used to be only IaaS, PaaS and SaaS. In my book App Engine and Cloud Run are PaaS, while Cloud Functions is FaaS. OTOH, there might not be any super clean definition
- raydiatian 4y agoSo many "aaS"es. It's definitely a spectrum. Thanks for pointing out App Engine, I will probably be using this in the future now that I read the documentation!
- throwawaaarrgh 4y agoTruth is an interesting concept. It's often subjective and has many forms. Within the context of the cloud, almost all cloud services are only mutable, so "truth" is whatever the current state of the cloud actually is. Whatever is in Git is merely idealism. Whatever you are maintaining, read the docs completely first. And I mean cover to cover. Not just the one chapter you need to get a PoC up and running. You will wish you had later, and it will come in handy many times over your career. Consider it an investment in your future. Read books on microservices before you implement them. Whatever two-line quip you read on a blog will not be as good as reading several whole books from experts. Docker multi-stage builds won't work in some circumstances. Build optimization eventually gets complex, the more you rely on builds to be "advanced".
- pabs3 4y agoThe only truth is the memory and disk contents of the devices that make up your cloud. Everything else is an abstraction of that, which discards data and potentially is out of sync with reality.
- crdrost 4y agoThanks for the alternative microservices quip, it was better than the original. Indeed, I find that “microservices should only perform a single task” is a really dangerous way to phrase it because we have no idea what the article means by “task.” The classic microservices separation is to separate an ordering service from a shipping service, is each of those one “task”? Or at the most extreme, is saving an order distinct from returning the list of your outstanding orders? Even when people graduate to a language of DDD and refine, often they settle on “one microservice per bounded context” where “bounded context” means “separated however I want it separated at the time,” and has no consistent principle behind... This despite the fact that I think Eric was quite explicit in his explanation of the idea, he meant a mapping of the software idea to the fussy complex world of businesspeople and business language: perhaps a better way to phrase this is that it's one microservice per archetype of user, “we have people from the warehouses who all speak the same shipping jargon, we should have a microservice specifically for them which speaks their language,” and I think most developers target their microservices smaller than that, in which case it is definitely not “one microservice per bounded context”. Don't get me started on how “strong coupling” is shorthand for “coupled in ways I don't like” etc. ... Sometimes I feel like I'm on an episode of “whose line is it anyway?”, where everything is made up and the points don't matter.
- nielsole 4y agoAnother random selection: * When choosing internal names and identifiers (e.g. DNS) do not include org hierarchy of the team. Chances are the next reorg is coming faster than the lifetime of the identifier and renaming is often hard. * The industry leading tools will contain bugs. From Linux kernel to deploy tooling, there are bugs everywhere. Part of your job is to identify and work around them until upstream patches make it to you if ever. * Maintaining a patched fork is usually more expensive than setting up a workaround * Your hyperscaler cloud provider has plenty of scalability limitations. Some of which are not documented. If you want to do something out of the ordinary make sure to check with your account rep before wasting engineering time. * Bought SaaS will break production in the middle of the night. Your own team will have the best context and motivation to fix/workaround them. When choosing a vendor, include the visibility into their internal monitoring as a factor for disaster recovery (exported metrics and logs of their control plane for example)
- vladvasiliu 4y ago> * Your hyperscaler cloud provider has plenty of scalability limitations. Some of which are not documented. If you want to do something out of the ordinary make sure to check with your account rep before wasting engineering time. If only they'd tell you. We had this exact issue on AWS. Seemingly random packet drops. Metrics on both clients and servers were ok, latency specifically was very low when it worked. Call up support "yeah, you're running into our connection limit". "Oh. What's that limit?" "yeah, I can't tell you that". His solution was that, since this was somehow related to connection tracking in the security group, I could set this to allow all/all, and set up filtering at the NACL level. Turns out I could do it for this particular issue. This was before there was a possibility to monitor this [0]. Called up our customer manager. "Let me check". A few days later, "yeah, that's not something we divulge". --- [0] For those who don't know, it's now possible to keep an eye on refused connections (at least on Linux). https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/monitoring-network-performance-ena.html#network-performance-metrics https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/monitori... -> conntrack_allowance_exceeded
- toast0 4y ago
- qaq 4y agoDon't just read docs try things -- make a POC. The amount of time we hit something that "should work" according to the docs but doesn't is very high.
- virgilp 4y ago> Microservices should only perform a single task. If you are not able to achieve that isolation, maybe you should switch back to a monolithic architecture. Do not get fooled by the current trends, microservices are not meant for everything. I feel like this is spectacularly bad advice. "Do not get fooled by shades of grey, things are meant to be either black or white!"
- zikduruqe 4y agoEVERYTHING costs money. Tag every resource. Come up with ways to show cost avoidance and cost savings. This is will be appreciated more by management than any code you can bang out.
- rr808 4y agoI love monitoring but after a few decades working I still haven't found a good way to monitor everything. Still a mix of email, pagerduty, prometheus, cloudwatch, websites, kibana consoles. Surely there is a good way to do this? I figure some of the new BI dashboards would be good but haven't seen much usage.
- TrackerFF 4y ago"Learn to say: I do not know about this/that. You cannot know everything that gets presented to you. The bad habit comes when the same technological asset appears for a second time and you still do not know how it works or what it does." Absolutely. I've seen so many junior engineers / devs go on about it like this: Someone higher up: Could you please look at this problem? I need it fixed ASAP. Jr. Engineer, presented with a problem he's never seen before: No problem, I will look into it! Someone higher up (the next day): Did you fix the problem? Jr. Engineer: Sorry, I haven't still gotten around to look at it / I'm still working on it / etc. Someone higher up: We really need it fixed today, please prioritize it and give me a call when it is fixed. Jr. Engineer works on the problem all night, feeling stressed out, not wanting to let down his seniors.
- bobismyuncle 4y agoSome of these are lessons you only really learn once you make the mistake yourself