13 ms·
You need to be able to run your system
- hansor 6y agoAt my organization I'm unable to get even virtual machine instance(not local) for testing and CI, yet they are perfectly content with testing software on production server... Actually they are able to spin-up some virtual machine for me - but without root access... So if I need some software or library I need to ask them and wait for days... Plot twist: can't run any virtualization software on my development laptop and we are cut off from "the internet"(so no ssh or remote cloud). Advice?
- sitkack 6y agoGet them to install docker on the VM and add you to the docker group. You can now use docker like VMs to run tests.
- saberdancer 6y agoWhat is keeping you at this company?
- grumple 6y agoThe interview process in tech is often long and laborious, and there’s no guarantee the next place will be better.
- hansor 6y agoVery good money for my skillset(IMO I'm below average programmer), "prestige"(company is in TOP 5 in its field), and things I do are not ordinary boring stuff like web apps - but rather specialized stuff.
- aequitas 6y agoCanary deployments. Don't replace a production version, but install the new version next to it. Switch between versions via loadbalancer or use queues to decouple different implementations from storage. Adjust your development process to favor small incremental backwards compatible changes and feature flags (make deploying your code separate from enabling/disabling a feature). Lots and lots of monitoring and alerting.
- sitkack 6y agoLife finds a way. Within the boundaries of the space is a solution, we just don't know where the boundaries are most of the time. A bunch of years ago we were under a similar predicament, we needed to tune some software that was data dependent, yet we were cut off from prod data, and the test data as hot garbage. It took weeks to get new software into production. We had one code push left before launch but we still needed to tweak with production traffic. So we did what everyone does, ship remote firmware updates. In the last code push we included `eval` and the ability to POST code to an endpoint. We were just doing some query rewriting, but the code could have run anything. Once the last build made it through dev and test, we had a serverless deployment platform that allowed us to dynamically update the code in realtime. yolo
- easton 6y agoBring the sysadmin a box of donuts (or scotch, if it’s that kind of place) and ask again. Barring that, if you have admin rights and your laptop is running Windows you can enable Hyper-V[0] and get a great hypervisor that comes in the box. (If you’re already using WSL2, you’re already doing virtualization). If you’re on macOS, try to find something that uses Hypervisor.framework (multipass[1] is good if you can get along with Ubuntu) so you don’t need to install any kexts. But this is really something you should bring up with your manager, it sounds like it’d be difficult to get any good work done without being able to run tests in a not-production environment. 0: https://docs.microsoft.com/en-us/virtualization/hyper-v-on-windows/quick-start/enable-hyper-v https://docs.microsoft.com/en-us/virtualization/hyper-v-on-w... 1: https://multipass.run/ https://multipass.run/
- hansor 6y agoBeer and donuts doesn't work as its "policy". As for Hyper-V: virtualization is disabled on BIOS/UEFI level, and I'm unable to change it.
- geofft 6y agoTell your manager you can't get work done up to industry standards. The high-level problem here is one of resource allocation and policy: you will have a similar technical problem in the future and you need to be confident that it's solvable, which requires management to be willing to resource solving it. You need either some other team's management to decide that this pace of development is something they will prioritize solving or your management to "go rogue" and get you set up on public cloud. But if you want a solution to this specific problem, the first suggestion I'd have is to install a non-root-requiring package manager like Linuxbrew or Nix (which is better if you can get a directory /nix writable by you, but workable even if not). Then you can install software without contacting the sysadmins and waiting. Alternatively, do you have unprivileged user namespaces, i.e., does "unshare -Urm" work and get you an apparent-root shell? If so, you can install rootless Docker in your homedir, or use https://github.com/containers/bubblewrap https://github.com/containers/bubblewrap . (A few years back I wrote https://github.com/twosigma/debootwrap https://github.com/twosigma/debootwrap which ties together bubblewrap and debootstrap to get you an unprivileged Debian chroot.)
- antisol 6y agoQuit. Developers need full access to their machines or they are stifled. This organisation is not interested in you doing quality work. Your options are therefore to either resign yourself to not doing quality work, or leave that organisation, ideally making sure that management knows exactly why you're going. (Of course, the chance of management actually caring is about one in thirty seven trillion - if they cared about quality of work or employee satisfaction, your machine wouldn't be locked down. So not bothering to tell them is fine too. Tell them if you're feeling charitable)
- jrochkind1 6y agoIs it possible to use, say, S3, and meet these criteria? (Then, something more complex or special purpose, like say Amazon Elastic Transcoder or something).
- theamk 6y agoSure, if you use a separate bucket for your teat instances, and reset it between tests. This can probably work with more complex services too.
- jrochkind1 6y agoto avoid conflict between two entities running tests at once (say via CI), I guess it would need to use Amazon API to construct the bucket and tear it down around each test? But I'm not sure that would comply with the requirements, they seem to forbid running against external services at all, no? '"calling out to a few hard-to-run external services" doesn't count.'
- theamk 6y agoYep, to use AWS S3, you'll need a lock to ensure a single job at a time, or create buckets programmatically for each test. I think every team needs to decide how much to follow the original advice, there is certainly a balance between simulating external service locally (with possible associated behavior changes) or using an external service. If you do want to run everything locally, there are servers which offer S3 API, like "minio" and "fake_s3". Probably would be very nice for localhost operation. But if you need a complex thing like Transcoder this is not going to work.
- csharptwdec19 6y agoI had to actually do something like this. Due to the hilarious nature of the business, getting an S3 account to use for 'local' (i.e. dev machine) development was a huge chore. We wound up creating an Interface to abstract our usage of S3, and wrote a 'local' provider that just used the local filesystem. Of course, many use cases that would not work.... But the big -advantage- of it was, we at least knew that if we ever wanted to switch away from S3, it wouldn't be hard. The Interface had a spec that could be used to ensure that any plugin would work with our application's usage. It wasn't -that- much more work.
- _kblcuk_ 6y ago> In my experience, they're always wrong. These systems can be run locally during development with a relatively small investment of effort. Typically, these systems are just ultimately not as complicated as people think they are; once the system's dependencies are actually known and understood rather than being cargo-culted or assumed, running the system, and all its dependencies, is straightforward. Whenever I hear these statements, it always sounds like "I need to have an identical copy of this skyscraper in order for me to be able to replace one tap on floor 42." Also good luck running a system that operates on few hundred of terabytes of for instance YouTube data locally. Also running the whole system locally is usually a pretty good way of creating a "distributed monolith" -- yea, there might be microservices, but also a dozen of assumptions here and there that different parts of system are being deployed simultaneously (usually they are not), or that certain distant parts of a whole system share some behavior that can be changed simultaneously. So no, you don't need to run the whole system locally. On the contrary, you need to be able to run smallest part of it (hello, microservice) locally, and that part should be responsible for one thing. APIs and frontends can easily share JSON schema to make sure they send and receive valid data, and each service can have tests against that schema, the ultimate source of truth for them. Boom, suddenly I can develop "big system" piece by piece in isolation on my 7-year old macbook with no problems against tests / storybook / debugger.
- pydry 6y ago>Whenever I hear these statements, it always sounds like "I need to have an identical copy of this skyscraper in order for me to be able to replace one tap on floor 42." The building industry also has the problem the OP describes: https://i.stack.imgur.com/yHGn1.gif https://i.stack.imgur.com/yHGn1.gif The problem as I see it is that people who go all in on unit tests tend to be dogmatic about it and suffer the above type of issue whereas the people who want, broadly speaking, to run things as realistically as possible are pretty aware of the real constraints. Moreover modeling larger is also frequently cheaper because the real thing often comes for free while creating elaborate frequently changing unit test mocks has very high opex.
- _kblcuk_ 6y ago
- jusssi 6y ago> "Run most services in a mostly-isolated development environment, calling out to a few hard-to-run external services" doesn't count. That's pretty harsh, as it seems to rule out any system that integrates 3rd party features, such as payments.
- tlb 6y agoPayment processors all (?) have a test API you can run against. Stripe, for instance, gives you a separate account with an API key that behaves the same but no money moves. You can test your whole process including exceptions, reconcilation, and chargebacks. Running a service that handles payment without testing the entire flow is insanity.
- sitkack 6y agoI don't think GP was saying they didn't test, they are calling out to the test apis, still 3rd party. One should go a step further and stub out or have a local payment processor so that everything runs locally and with no delay.
- machello13 6y agoThat breaks the "no mocks" rule.
- jstrong 6y agoexternal services don't always have a test API, and it's not always as cut and dry as processing user payments. sometimes it's less important than that, and testing it would require replicating an external server to be able to test it. in those cases the tradeoff of time spent vs. value isn't as clear.
- sitkack 6y agoStubbing out payments is really nice, you can have end to end integration test flows that execute in milliseconds. Use the affordances that the concept of interfaces provide.
- 6y ago
- pmontra 6y agoTwo recent anecdotes from two different projects: Project 1, Elixir. > Run an individual service against mocks" doesn't count. We got an error in production a couple of weeks ago. The code was pretty old so it was a surprise. Found the problem, a rare case we never thought about, and the culprit: almost all the tests using that function where run against a mock of that function because it requires quite an elaborate database setup to work. The function itself was tested only in a couple of unit tests which exercised only the naive cases. Big mistake! Project 2, Python. We spin EC2 VMs to run some long running CPU bound task and kill it. Of course we can do that in a development environment. We're already using SQS in development. And yet spinning remote VMs takes longer than locally. So we build a couple of scripts around VBoxManage (VirtualBox) to run those VMs locally. Booting them from an SSD is really fast and it works when working offline. We have two classes with the same methods, one wraps VirtualBox, the other wraps EC2. The configuration file of the application picks the right one given the runtime environment.
- tutfbhuf 6y ago> In my experience, they're always wrong. In my experience, if someone claims that someone else is always wrong, then he is wrong.
- lcfcjs 6y agoWhat trash is this
- midrus 6y agoSuch a terrible advice. I've been several times involved in companies where this approach was followed, and where you could only run your own service with mocks, and we had incredibly better results with the latter approach by a far, far, far amount. Just imagine an Amazon developer running amazon in its own laptop (or cluser, or server, or whatever... multiply it by the number of devs... it is insane).
- johnobrien1010 6y agoI one time met a guy who was trying to debug an issue he was having in production on a sprawling Django code base which he couldn't run in any other environment. It was impossible until he got the system running in the local environment. I also came on once to a set of projects where there wasn't a good separation of the environments; test and prod shared the same database, b/c it was "too hard" to replicate the prod data back to test. Consequently, no real DBA work could be done b/c it couldn't be tested, only small hot-patchy kind of work. Separation of concerns is a slightly different problem but the root cause of the issue is similar, you need to be able replicate the complexity of your system in production as much as possible in a lower environment, otherwise you won't be able to make as significant changes as you'd like to your system. I don't think this is terrible advice, while I recognize there are limits to what is possible to run locally/in lower environments, the central notion of 1) separation of concerns and 2)matching the complexity of the lower environment to production as much as possible is I think good advice.
- nonameiguess 6y agoSo many +1s > Developers of large or legacy systems that cannot already be run in their entirety during development often believe that it is impractical to run the entire system during development. This is just a flat out ridiculous statement this author makes. My wife works on an automation system intended to respond to incoming events in the form of positive tips in surveillance system collections that in turn cues further tipping rules to retask those collection systems. One of the external dependencies is the entire spy satellite control system of the United States. How are you supposed to run this in your development environment? Some things you have to mock.
- kerblang 6y agoI'll agree in general with the sentiments of the author, but for a lot of people it isn't practical, because things are already gigantic, overcomplicated, etc, and it's out of your hands. So instead: Be prepared to run _any part_ of the system locally, like databases, message queue thingies, web applications & services, etc. This gives you lots of flexibility and control you wouldn't have otherwise. The up-front work can be painful, but it pays off over time. Know how to proxy missing parts: For example, have a local nginx server proxying requests to a shared server that can handle some things. Set up integration tests that actually bang on running applications (over, say, HTTP) instead of testing little blocks of code. Even if these only run on-demand (full runs may take hours, require overnight scheduling, bleagh) they'll be super-handy. Ultra-distributed systems are a nightmare, but you do what you gotta do. In the end you'll end up being your own "devops" engineer and know your architecture better. Edit: A lot of people think of "mocks" as unit-testing types of things, but "mocking" an entire web server is often the way to get out of an integration-testing pinch, especially when some production-only remote thing interferes with the critical path you're on.
- arcturus17 6y agoWell this is timely. Just this morning, I’ve been extending a Firebase app I developed for a very large enterprise customer, and I’ve been constantly reminded of some of the things the author speaks about as I tried to recreate the FaaS environment locally. Not only did I waste valuable time setting everything up locally when the same would’ve been trivial in an old-school web framework; I now have an explosion of environment permutations and smells in my code, and I don’t completely trust that local and prod will behave the same so I will have to do extra testing. Despite Firebase coming a long way - some of these local emulation facilities plainly didn’t exist or barely worked a couple of years ago - there is no way I would choose it for a new project. The only module that has genuinely saved me time is auth; none of the rest have made me happier or more efficient than a web framework. Even for prototyping (the reason why I picked Firebase initially) I would still pick Django or Express or whatever else over this. (And I haven’t even delved into the big elephant in the room, which is lock-in.) Regarding costs and scalability there are a gazillion (fairly easy) things I could’ve done to achieve a similar cost range bringing my own stack. FaaS and the like may have their particular use cases but for anything that has even the slightest potential to turn into a full-blown app I’d go for a good old monolith every time.
- btown 6y agoFor realtime requirements, I'm keeping a close eye on Supabase, which is building an open-source realtime Firebase-style API on top of the Postgres WAL. Theoretically, it's just Postgres with a bunch of services on top, so if you bring your own database migration and fixtures system, you could run a copy locally. Not sure that tooling is fully there yet (and it could use some of the model-layer bells and whistles I remember from the days when Meteor was the way you'd go for realtime), but the dev experience is very promising: https://github.com/supabase/realtime/blob/master/examples/next-js/pages/index.js#L36 https://github.com/supabase/realtime/blob/master/examples/ne... https://github.com/supabase/supabase https://github.com/supabase/supabase
- arcturus17 6y agoYes, I also think that realtime is something where Firebase or Supabase may help. The app I'm talking about has no such requirements, and likely won't have them ever. Even then, if I were to build an app with a predictable realtime requirement I would (a) find a full-stack framework with this as a core feature (b) use a framework without this core capability, but look very hard for the simplest library that would allow me to pull off the requirements, and only if all else fails do (c) which is to isolate the realtime component into something like Firebase and build everything else in the core stack. (I haven't tried Supabase, but from what I understand it might fall more into (b) than (c), since you can bring your own infra)
- paxys 6y agoWhat does "system in its entirety" even mean? Should an Amazon developer be able to run every Amazon property simultaneously on their laptop? As long as you have strict contracts between parts of your system (whether that is in a form of microservices, IPC, classes or whatever else), a change made in isolation should be perfectly fine.
- asim 6y agoYou must invest heavily in development tooling. You must invest heavily in an architecture model that caters to local as a first class citizen. You must invest heavily in developer mindshare to actively and continually reinforce that local is a must-have, not a nice-to-have. And to do all that takes an excruciating amount of effort beyond just building your products. Ultimately its a tradeoff, what are you actually trying to achieve. At some point you'll cut a corner, or a developer on your team will, or the next person will. Eventually the system doesn't work locally because of edge cases and the only way you claw this back is by mandating it a policy to all that it has to work locally. I built a thing with a local first view and I still battle this everyday because whether we succeed or not will not be dictated by whether it worked locally, but by actually shipping a product that people want and are willing to pay for. Tradeoffs. Sometimes you have to just let go of the purist view point.
- glenjamin 6y agoThis post to me reads like the author has given up on defining contracts between systems. While I think it's useful to be able to exercise a system for real, I generally prefer to do it in production rather than attempt (and fall short) of reproducing production in an isolated environment. Meanwhile, I'll invest my time in defining and understanding the contracts between services, and develop against those contracts so I don't have to run the entire stack of connected pieces every time I want to make a small change.
- bob1029 6y agoTesting in production is the best thing we have ever done with our customers. The key is to be able to quickly undo some scary new thing and get back to a known-good point. Also, you need to make sure the testing you are doing is easily distinguished from normal business activities. For us, the humble feature flag is all it took to get unblocked on moving mountains of complex scary things to production. Being able to reassure the customer that we can instantly revert a piece of experimental functionality has been a game changer. We used to spend months agonizing over how the test environment is such a poor facsimile of production and complaining about vendor XYZ not setting up various things that we would need to prove our code works as expected. We used to make insane bullshit promises about how if things worked well in staging the move to production should be flawless. Those were some dark times for us.
- Spivak 6y agoIsn't this just blue-green deployments but pushed into your app's codebase instead of the deployment posture? I have no doubt that this works for you but why not just just flip between two versions of your app?
- milesvp 6y agoYeah, never underestimate the value of solid rollback steps. Never deploy anything without some plan to get the system into the previous state, ideally with a button push. And for those rare cases when you can only roll forward, like with a major schema change, make sure that is easy to do too. There’s nothing like the stress of a broken build to make you stupid.
- sojournerc 6y agoMajor schema changes can also be done in a backwards compatible/recoverable way - I like to follow this procedure. Specifically the "The Five Phases of a Live Schema Change" section: https://queue.acm.org/detail.cfm?id=3300018 https://queue.acm.org/detail.cfm?id=3300018
- bob1029 6y ago
- nitsky 6y agoThis webpage has no CSS and no JavaScript, just a `<head>` with `<title>`, an `<h1>` for the heading, and a number of `<p>`s for the text. It's styled according to my browser's settings for font, size, etc. How refreshing!
- twic 6y agoIt is quite elegant. If the browser's default style was a little better, it could be even more elegant - not a full window width, for example. Could browsers make such improvements to their default style? Or are assumptions about the default baked into sites in such a way that changes would break them? Could browsers detect those assumptions and provide an appropriate default accordingly?
- ojr 6y agoit looks unprofessional, it is hard for me to take technical advice on the internet from someone running an unprofessional looking site. There is another blog post written by the author that docker containers are harmful that I strongly disagree with but not totally unexpected given the design choices of the site.
- oftenwrong 6y agoTheir site doesn't provide any opinion styling. Blame your client.
- choeger 6y agoWhy do you even bother responding to someone who uses such unprofessional software?
- hactually 6y agoI wondered if the GP post was from a frontend guy who had a different definition of "professional". Clicking through their profile to their business/project website was a loading screen for 5+ seconds. I never got a true "professional" impression as I closed the window before it finished.
- rkangel 6y agoSo this is a a sample of the reason why I like Elixir (or Erlang). Your 'traditional' microservice architecture, running on Google/Amazon/Azure infrastructure, using lots of their pieces is very hard to test in an end to end way and you've got no chance of running it up locally. Kubernetes provides an abstraction layer that improves this but it's not perfect. Even a Django framework probably has a load of workers doing stuff off process. In Elixir, what you've got is a lot of Elixir. Sometimes you need other external services and there's some work there, but you can often get a surprising amount of your system in one codebase, running on one more instances of the VM. It makes testing SO much easier.
- awinter-py 6y agoI mean 'users' are part of the entire system
- ftio 6y agoIn hardware, it's not enough to design the object itself. Objects need to be made. At large scale, they need to be manufactured. When you manufacture novel objects, you must also design all of the object-specific tooling that goes into producing your thing, like molds and fixtures, which are just as important in producing your object as the design of the object. It could not be made at scale without designing those fixtures and molds. For complex objects, creating effective molds or fixtures is a high-skill job whose difficulty is equivalent to that of the object designer. After working on developer infrastructure at a big company for a few years now, what I've realized is that most companies don't spend very much time at all designing the 'molds' and 'fixtures' their software teams need in order to manufacture a worthwhile object that they know will work. They pay lip service to the importance of testing and solid dev tooling, but they're reluctant to invest in it. (I don't blame individuals here — they're behaving rationally given market pressure — but it's a sad state.)
- vinceguidry 6y agoI'll carry that torch when I feel it's appropriate. But the reality is, this practically requires making the decision to use free software and keep the entire open source community at a long arm's length. Why? Because business doesn't give two shits about your productivity. They care about one thing, the productivity of the software. And so they will buy other goods and services that will at some point be closed source. And then you're the one that has to integrate them. Are you deploying to GCP? Already you can't run your full stack locally. My last week was spent hacking Google's Anthos CLI tool, nomos into a kpt function. Kpt is open source, nomos is not, it's not even source-available. It's just proprietary. And the platform it interacts with, you can't run locally. Go ahead and check out the troubleshooting page for one Anthos tool: https://cloud.google.com/kuberun/docs/troubleshooting https://cloud.google.com/kuberun/docs/troubleshooting No sections on spinning up a local env. You can't self-host or run locally, Google won't let you. This isn't isolated to just Google. People need to get it into their heads that open source is the realm of business and not the realm of hackers. You're a cog in their machinery. If you can't run free software for your whole stack, then your troubleshooting steps will necessarily involve understanding at least two environments, one of which will be a very black box. Devs, if you want a better world, start contributing to copyleft software.
- skybrian 6y agoIt seems like running a separate instance in the cloud counts as running the system, though? The key point is that you, the developer, get your own instance.
- feoren 6y ago> Devs, if you want a better world, start contributing to copyleft software. Your version of "a better world" is one in which authors, musicians, movie producers, artists, and all other creative types are free to make money from their works, but not software developers. When they try to do it, it's immoral. No thanks.
- vinceguidry 6y agoWhat? I didn't mean to imply that. I'm an associate member of the FSF. I pay my dues with the money that comes from my corporate job which is open source all the way. I use Emacs, and keep my personal hardware as close to free as I can. One day when I'm into the ecosystem enough, I'll start contributing code to GNU projects. The more we can build up free software as an alternative to open source, the more the business world will be forced to use it. They need your dollars more than anything.
- tylermauthe 6y agoThis is quaint and wholesome. I long for the simpler days when I could agree with this, blissful in my naivety of large scale organizations. Nowadays, I accept this reality is largely impossible and you must always draw some boundary. This doesn't mean all your developers should use a shared MySQL because nobody knows how to bootstrap the database- but it means you have to consciously decide where you sit on the continuum. Always expecting to run the whole system on your laptop (or even in a cloud) is also an unreasonable expectation, unless you're Netflix and your revenue per employee is into the millions. The reasons for this are many and complicated, but at a high level the work required to make it happen will cost too much and won't be a priority. I'll quote here from the excellent "Testing Microservices the Sane Way" by Cindy Sridharan: > asking to boot a cloud on a dev machine is equivalent to becoming multi-substrate, supporting more than one cloud provider, but one of them is the worst you’ve ever seen (a single laptop) https://copyconstruct.medium.com/testing-microservices-the-sane-way-9bb31d158c16 https://copyconstruct.medium.com/testing-microservices-the-s... "Full stack in a box- a cautionary tale"
- amelius 6y agoTell that to people who design rockets.
- grumple 6y agoThey seem to be given an allotment of several tries that result in blowing up their systems.
- Diederich 6y agoThis is an incredibly bold assertion: > Developers of large or legacy systems that cannot already be run in their entirety during development often believe that it is impractical to run the entire system during development. They'll talk about the many dependencies of their system, how it requires careful configuration of a large number of hosts, or how it's too complex to get reliable behavior. > In my experience, they're always wrong. These systems can be run locally during development with a relatively small investment of effort. Typically, these systems are just ultimately not as complicated as people think they are; once the system's dependencies are actually known and understood rather than being cargo-culted or assumed, running the system, and all its dependencies, is straightforward. In my 28 years of professional experience working in 9 different organizations, only one of them did fully run the core system in development, and another two of them could have done that with enough time and effort. For the other six, there was no reasonable way to run the full system in development. > Typically, these systems are just ultimately not as complicated as people think they are In fact, most long-running systems are quite a bit more complicated than people think they are. Granted, a variable but usually large chunk of that complexity can be removed/simplified, but at their core, these 'legacy' systems, the ones that make real companies work, do their job because they have, over many years, figured out how to successfully come up against all kinds of real world complications. To be clear: I get what (I think) this person is fundamentally trying to say: the more expansive/end to end a development environment is, the better. That it's better to have as wide a view into the system you're working on as possible. Good stuff!
- whb07 6y agoAren’t you making the point of the author? Most places don’t have a way to run in development because they haven’t sat down and gotten the application to a running state. Every year it gets ignored, things get harder and more tedious but never impossible
- imoverclocked 6y agoI see both sides of this argument. Often legacy systems are tied to their deployment methodology/dogma and the world moves on from that model. Without a lot of resources poured into maintaining that methodology and upgrading it, the expected environment can age out of current practices far enough that running it locally does take a lot of effort. I think docker is an obvious enabler of this because people can just spin up an image from 6 years ago in a container and keep running against old libraries/dependencies. Even worse is depending on that container makes it difficult to upgrade dependencies to modern versions of software because the container is so old. Maybe a corollary to the article is, "without constant maintenance, software becomes un-runnable."
- compscistd 6y agoOur server went down for a couple of days and if we could "run our system on localhost", I'm positive we would have been back online very quickly as opposed to the multiple days it took to track down stored procedures not in version control. Front-end was left twiddling their thumbs during the outage because the server wouldn't run on local and our frontend wouldn't run without a server (we neglected updating our frontend model mocks for years). Did we learn our lesson from the outage? A big _nope_. I suspect it's because being able to run a somewhat complicated system on local requires thinking in brand new ways with benefits that aren't very obvious from the outset. After that experience, I sympathize a lot with the author's points and hope to work in an environment (ha!) where spinning up a docker container is all it takes to have a _full_ dev environment.
- hinkley 6y agoI often set up a "how to create a dev environment" wiki and then we exercise it many many times. IBM got a bad shipment of laptop hard drives that exhibited a MTBF of about 2 years, and our equipment dept bought a stack of laptops from that batch. Over a summer we had 6 machines go belly up. Mine was number #5. People still looked at me like I announced that I had stage 3 cancer. Oh you poor poor man. I found this reaction disappointing. By then the process was about as documented as any I've had. It just took me a day to get it up and running (because the base image left a very slow step until after 2nd boot, which I still maintain is dumbness squared). A coworker from that cohort had an experience that I still use as an example. He tried to put his work laptop in the back seat of his car. He missed and hit the door frame. Killed the laptop. Similarly, taking your laptop down the stairwell could be a one way trip to the trash bin. If the information is important for us GET IT OFF OF YOUR COMPUTER. As soon as you know. Put it in storage, or at the very least in some coworker's head/computer. If you do this, consistently, then losing your machine is a shitty inconvenience, but nothing more dire than that.
- gamacodre 6y agoStrongly agree. In the last three years I've rebuilt my macOS laptop from scratch twice (dead hard drive, and a dev tool run amok), and once swapped it for a different size because I really wanted the larger battery. My co-workers were appalled, but I had less than a day of downtime each time. With continuous backup to an external drive (always on while I'm working), a password manager I can also store certs in, and keeping everything important in a cloud-backed git repo, there isn't a single thing on my work computer that's hard to replace. It's been fantastic.
- dr-detroit 6y agoThe author should write better unit tests. This is why its called work and not do-whatever-feels-good.
- idlewords 6y agoThe easiest way to follow this advice is to bang on your production system directly.
- kelvin0 6y agoImagine you're working at MS and have to build/run windows from VS.
- kelvin0 6y ago* sigh *
- tobyjsullivan 6y agoWe used a variant of this[0] when I was at Canva (and I'm sure they still do). Having a hard requirement that you are able to fake out any external dependency in dev environments provides a massive boost to dev speed. The investment is tiny too - often a single interface and a day of dev time to build out a functional fake service for, say, S3 or SQS. [0] https://product.canva.com/hermeticity/ https://product.canva.com/hermeticity/
- xrd 6y agoI would bet 78% of managers trying to get their developers to go faster will coach them to try to get that PR done in one day instead of 3 RATHER than investing the three days in taking heed of this advice. The main reason developers are slow is because of problems described in this post. But, investing in that change will be really, really scary.
- dane-pgp 6y agoThere really needs to be an expectation during sprint planning meetings that developers get a chance to adjust a "multiplier" for all story points, as a way to keep in everyone's minds the long-term costs of short-term decisions to cut corners by foregoing proper testable designs and good engineering practices (including security). One approach I've suggested in the past is that every time a developer has concerns about a design and is over-ruled by a manager, the developer should take a jelly bean from a jar that starts full of jelly beans. The team can keep track of the weight of the (contents of the) jar, and use the inverse of the "fullness" fraction as a multiplier for how long a ticket will take. For example, if a story has a complexity of 2 points, but the jar is two-thirds full, then the ticket should be expected to take as long as a 3 point ticket would. That might not be exactly right for every ticket, but this process might at least be more accurate than ignoring the hidden cost of accumulating technical debt. Of course this means that once the jar is empty, all estimates become infinity, or undefined, but if it ever gets to that point then at least the developers will have had some jelly beans to cheer them up. Similarly, a manager can have the satisfaction of "seeing" developer productivity increase when, at the end of a sprint focused on paying down technical debt, the developers add more jelly beans to the jar, to restore the fraction towards unity.
- xrd 6y agoSounds a lot like the budget described in the Google SRE book. Agreed.
- peterbell_nyc 6y agoI fundamentally disagree with the article - sorry! I used to believe it - for many years. But as systems continue to add essential (as opposed to accidental) complexity, the only way to run production is to run production. Why do you want to be able to run the app elsewhere? To test new features? To reduce regressions? Reasonable goals, but at the end of the day the only thing that is identical to production is production. Unless you have a real time copy of all of the data and you continually run copies of all real time production requests, it's not production. It might be the same code and a very similar infrastructure, but without the same data and load, you are going to get regressions in production. Maybe it's flaky historic data or unexpected patterns of load, but whether we like it or not, we're all already testing in production. I'm a huge fan of unit tests, CI, and all of the other common best practices to reduce the number of bugs that are identified in production, but you also need to have the kind of tooling and processes required to minimize recovery time and to be comfortable with testing in production. Small, easily testable, quickly shipped units of work and some flavor of feature flagging so you can dark ship code, recover quickly from outages, and do things like canary roll outs to ensure the new queries don't break with the production data at scale!
- jl6 6y agoThe tech giants are each their own unique thing and not generally a pattern that anyone else can or should follow. Most organizations, for example, do not have hyperscale volumes and face no insurmountable barriers in setting up parallel environments for testing.
- Sebb767 6y agoI disagree. Yes, fully replicating the production environment is not possible for big apps or apps with customer data, no discussion there. But when you want to debug components, test error conditions etc, a copy is extremely important. You can do "[DEBUG] added some logging" commits left, right and center, but it's not going to replace a debugger. Additionally, you might want to create custom or corrupt datasets. Yes, you can theoretically add a flag, but then this flag needs to be checked everywhere and might come with its own bugs. Using the "customers_debug" table instead of "customers" table (for example) works as well, but then you replicated staging in your production environment - added complexity which is definitely _not_ needed. Lastly, this misses the other points the article makes - a local copy allows you to freely play with the configuration, shut down related services etc. You can not do that in production. But I'll give you that - resiliency in production is still very necessary and you usually won't be able to replicate everything locally. The ability to run local isn't everything - but that was not the point the article was making.
- joana035 6y agoYes, if you can not run the real thing locally you should not be touching a production system at all. I worked in so many places where nobody knew how the system worked and the most invaluable thing I do is to bring the production system into localhost. It pays off real quick and people get amazed how fast I can solve long standing issues in the system by using tcpdump and co.
- amdelamar 6y agoMonolith, sure. Microservices, no.
- euske 6y agoWe've seen one counterexample recently: NASA Perseverance Rover. They emphasized that this was their first time of running everything in the real environment although I'm 99.999984% certain that they've tested it in the closest setup as possible. I'm also certain that there must have been a few glitches here and there, but the overall system worked. There's a process of mitigating and containing the faults and failures, and NASA is known for that. I don't disagree with the principle of OP, but just repeating it in a dogmatic way doesn't make things better. I want to see practical examples and ways to fill the gap (between the dev and prod).
- asd4 6y agoMost hardware engineering is done ahead of time without a full production style environment. This is because the cost of iterating is much too high. You can't build a bridge every time you want to try a new cable or bolt. It forces designers to make models and assumptions about their systems and, inherently, puts downward pressure on complexity. It also forces them to truly understand the principles behind what they are building. The fact that Perseverance and other Mars rovers have been successful is amazing and took an incredible amount of work to accomplish. These are complicated systems that were vetted using models and simulations without ever having been run "in production". This comes at a high cost. Critical software is never tested in production or run "in system" before it is deployed. Airplanes, banks, medical systems all require extensive validation through testing on models and simulations. You can't test your changes for the first time on a live aircraft or living tissue. Costs reflect that. The truth is, a lot of software is not critical. You can get away with hacking / trial and error type development and never fully understand the system you are helping build. Frankly there is a lot of money to be made providing brand new services that are unreliable or quirky or ephemeral because those services never existed before. My point is that how you test and validate your software product depends on your application. Sometimes the costs don't make sense to "run everything" and sometimes its physically impossible. I agree that you should always advocate for the highest fidelity testing your business case can afford, but be prepared to settle for less than everything and rely on your engineering skills to buy down risks in the gaps.
- theknocker 6y agoI have to say, it's surprising to me that self-respecting software engineers can be comfortable with ad-hoc systems that are difficult to reproduce, running in production or anywhere. It should be a huge red flag about dependencies and component boundaries.
- srich36 6y agoI’ve always been curious, if you’re working on a really complex system with lots of disparate services (or even if you are using a managed database like Spanner for example), what does your development environment look like? Do you spin up containers for all the services? Run a compatible RDBMS instead of the managed database? All my experience has been with systems that can be set up locally - how do you go about developing/testing/debugging without that?
- rurban 6y agoAnd then it gets interesting when your are not able to run your system by yourself. Which happens in plenty of circumstances. Eg You are writing software for a planned device which is operating on Mars. First, the device doesn't exist yet, second, it's hard to simulate the Mars environment locally. I did such projects for very expensive devices. Think of nuclear power plants or formula 1 race cars. Thanksfully you cannot complain loudly "I need to be able to run this thing locally!". Your job is also to figure out how to simulate the environment, and to do a proper job without being able to test everything beforehand.