6 ms·
I've definitely lived with the zombie flags problem. Teams ship experiments that double the size of a piece of code, but never go back to refactor out the unuse
by zero_shift 3y ago
I've definitely lived with the zombie flags problem. Teams ship experiments that double the size of a piece of code, but never go back to refactor out the unused code branches. In shared codebases this becomes a nightmare of thousands of lines of zombie code and unit tests.
This is a social problem as much as a technical one: even if you have LaunchDarkly, DataDog etc making very clear that a flag isn't used, getting a team to prioritise cleanup is difficult. Especially if their PM leaned on engineers to make the experiment "quick n dirty" and therefore hard to clean up.
At The Guardian we had a pretty direct way to fix this: experiments were associated with expiry dates, and if your team's experiments expired the build system simply wouldn't process your jobs without outside intervention. Seems harsh, but I've found with many orgs the only way to fix negative externalities in a shared codebase is a tool that says "you broke your promises, now we break your builds".
- esafak 3y agoA softer solution is to name and shame with periodic "leaderboard" emails to the org showing how many experiments each team has failed to clean.
- JohnFen 3y agoBut isn't spam like that exactly the sort that will ultimately get completely ignored? I probably get a dozen or so such barely-relevant internal emails a day where I work, and have learned how to recognize them by the sender and subject line, and ignore them. A "leaderboard" email would 100% be one that I ignore as a time-waster.
- esafak 3y agoNo because the whole org sees it and the org head can tell your team's manager to get the house in order. We did this at my last company, albeit with migrations rather than feature flags. Same idea. I believe I read about it in a book (perhaps Software Engineering at Google) in the context of test coverage; using a leaderboard for gamification.
- JohnFen 3y agoAhh, so the real target for such an email is management rather than the rank and file? That makes more sense. But surely, it would be better to send email just to those people and put the data up on the company intranet for those devs who are curious.
- esafak 3y agoYou can but transparency is a good thing, and if you see it you can fix it before your manager chides you.
- masklinn 3y agoIt assumes management gives a rat's ass about it. In the original comment: > Especially if their PM leaned on engineers to make the experiment "quick n dirty" and therefore hard to clean up. The PM then just needs to say that they don't have the time to cleanup because they have shit to do, and there's folks out there who pretty much live for this, and are usually well seen by management because they're management's executioner: no matter how shit the idea and execution, they'll get it rammed in.
- esafak 3y agoThe PM is not the EM. Accountability is driven by the org head through the EM. If none of them care then this system will not work of course. This system is good because it gently targets multiple stakeholders.
- JohnFen 3y agoThe point that I was making is that email might not be the best way of distributing this generally, because in many companies (and most of the larger ones), there is a high rate of companywide emails that amount to just spam. This increases the chances that these particular emails will get mentally categorized the same way and ultimately ignored. I know that I delete unread about 80% of the internal company emails I get because they're not actually useful, and have better uses for my time. An email like this, assuming that my own projects are not often mentioned in them, would rapidly get ignored, I imagine. If the goal is to have the affected teams alerted to the situation, it seems like it would be better distributed directly to the teams individually (a targeted email just to the teams involved) rather than spammed to all of the devs in the organization.
- benpapillon 3y agoTotally agree that it is a social problem as much as a technical problem. This is one reason why I had the thought here of FM tools starting to own some "feature management" jobs that aren't typically placed under the devops umbrella, and may be more of interest to product or marketing stakeholders. Perhaps that would do something to help with the issue of getting buy-in to do the maintenance.
- marcosdumay 3y ago> getting a team to prioritise cleanup is difficult This is the limitation that breaks every development practice people come up with. That idea of formalizing the cleanup and requiring it for deployment is very interesting. It may be possible to extend it to other contexts.
- 6D794163636F756 3y agoIt runs counter to the pressures a developer faces. I've at times been told that tech debt is fine because products only last 3-4 years before a replacement gets made. When you're on that timeline who cares if you've cleaned up after yourself?
- marcosdumay 3y ago> products only last 3-4 years before a replacement gets made Is that your experience? I imagine it can be a kind of self fulfilling prophecy, but even then I can't imagine people replacing everything each 4 years. And if that's not your experience, then the point is moot.
- 6D794163636F756 3y agoThat's how long it takes before the person who signed off on it has moved to another job and made the product someone else's problem. At that point the new person will usually kick off a new project because the old one is bad and releasing a new product is better for their chances of promotion
- IggleSniggle 3y agoI worked at a software shop with very high retention rates (like, 30 years, 10 years on average), and the inverse can also be an issue, "I own this product and it's my problem not yours." Having seen both situations, I personally believe that it can also be that a new project comes to exist simply because the old one is too complicated to understand; some things you need to work through in order to get. Someone here on HN said recently, "people forget that the primary job of the software engineer is as a learning agent for the org" or similar, and the more I see, the more I believe it. I used to think it was all about efficient automation, but I'm not so sure anymore.
- sitzkrieg 3y agojust went through a massive layoff and reams of flags in LD no one understands, extra fun! best part LD is so expensive we have to share hot seats
- jpdaigle 3y agoThe tricky requirement that ends up existing and torpedoing attempts to clean up feature flags is a requirement for long-term holdback. e.g. "Test was successful so it's rolling out to all users, minus a 0.5% holdback population for the next 2 years" This then forces the team to maintain the two paths for the long-term, ensuring the team might get re-orged / re-prioritize their projects sometime a year later making the cleanup really hard to eventually enforce.
- SketchySeaBeast 3y ago> "Test was successful so it's rolling out to all users, minus a 0.5% holdback population for the next 2 years" Man, I couldn't imagine being a user in such a situation. "Oh, I guess I'm just not getting the better functionality?" Even worse if I were a paying customer.
- Brian_K_White 3y agoYou are probably the lucky elite who got to keep the functionality you wanted.
- adamesque 3y agoIt’s actually usually the paying customers asking via support to be added to the holdback, improved experience or no. This is more true for larger flags that substantially change the experience and may not implement niche or edge-case functionality. Obviously you want to avoid these kinds of tests if possible but it’s not always possible.
- esafak 3y agoUsers should not be allowed to select their treatments; it defeats randomization, which is what allows causal inference.
- mandelbrotwurst 3y agoSure, they'll be more predictive that way, and simultaneously it's valuable to not piss off your customers.
- brightball 3y agoI've seen that easy enough to address with a frequent review (quarterly, per PI, monthly, etc). If you're operating in some methodology that has a consistent cadence, it should be manageable but you do have to be deliberate about it. Doesn't take long.
- no_wizard 3y agoThis is the only way, more or less, to enforce any code standards, whether its refactoring, quality, testing, docs etc. If it doesn't break the build (or do anything else that stops it from moving forward) there will always be external pressure to get things out "ASAP" despite in most circumstances "ASAP" isn't required. If you can't full on stop whats happening, it becomes exponentially harder to enforce anything.
- deathanatos 3y agoI agree with you, entirely, but I've had developers fight tooth and nail when the build breaks. "We need to get out there ASAP!" is definitely the cry, and usually "compromises" are made, such as "what if we made the overall CI run not fail if this test fails?" — which is as good as killing the test, IMO. The impetus is necessary, or the problem is ignored. The devs are really just proxies for the stress a PM is inappropriately pushing, though. But they are paid to not understand this problem, so getting them on board is impossible.
- tomhallett 3y agoAgreed, but at least with the override mechanism being as visible as the test suite, you have assured that more people will know about it (bringing it "to the surface") and you might even have the override in source control (documented). These aspects are increasing the chances someone will say "Let's just follow the process."
- deathanatos 3y agoIn our case, the override mechanism caused subsequent confusion. People were confused: "why does my CI run fail on [the security test]?" when the security test was set to specifically not cause the larger run to fail. The security test would still red-X itself (to visibly indicate that it was, in fact, failing), but not block the entire run. But people nonetheless went: "the run failed" -> "that test failed" -> "why is this test failing?" (There was a second failure in the run that was the actual reason the build, as a whole, failed, but that was blindly missed.) That triggered a large discussion the "compromise" of which was, to not "confuse" people, to have the security test green check itself on failure. And so now it is truly invisible. (I've actually turned it back to the "red-X but don't block the larger build" mode since then … but it still causes confusion. I do not know how to further help people who cannot understand the output from a build that has two failures, one of which is failing on master which is on the whole green, and one of which is only failing on your branch.)
- sb8244 3y agoThe big issue I ran into with zombie flags is that new features were always prioritized over cleanup. Engineering could "fight for time" to get things done, but there were always other priorities that needed to be addressed. No tool you have will solve that, whomever owns the product team time allocation needs to be onboard with the idea of cleaning up old code.
- alexjurkiewicz 3y agoWhat prevents teams extending the expiry date repeatedly?
- arein3 3y agoThe lead and other membera of the team that understand that temporary fixes should not be forever.
- hakfoo 3y agoOur feature flags tend to be "staged deploy" feature flags, and it hits an internal Slack channel when they hit 100% available. This usually triggers someone to queue up a ticket for "rip out the old code".
- deleted 3y ago[deleted]