6 ms·
One of our internal experimentation systems handles this by capping rollouts to 95%. So the author is forced to clean up their code if they want a full rollout.
by ublaze 6y ago
One of our internal experimentation systems handles this by capping rollouts to 95%. So the author is forced to clean up their code if they want a full rollout.
- sjtgraham 6y agoBrilliant. This reminds me of the apocryphal tale where NASA spent millions developing a pen that could work in microgravity… and Russia just used a pencil.
- dean177 6y agoPlease stop repeating this, it’s total nonsense.
- smabie 6y agoThat's why he said apocryphal..
- dean177 6y agoWhich means they are of doubtful origin. There is no doubt that this particular story is a complete fabrication
- jolux 6y agoYup. Completely false. The Russians ended up using that pen too or their own version of it; the Americans started off using pencils but it turns out having electrically conductive graphite shavings floating around is bad for electronic equipment.
- obilgic 6y agothis eliminates the ability to do a revert though if things don't work as expected at %100
- capableweb 6y agoI'm not doubting it's impossible, but I can't come up with any scenarios where something would work perfectly at 95% but then break at 100%. What could be some examples of things that can break like that?
- Shish2k 6y agoI saw something like that one time - the new version had one rarely-used broken API endpoint; clients who hit that would silently retry until eventually hitting an old instance which worked and then they’d be on their way. It was rarely used enough that the total number of retries didn’t trigger any alarms, and failed fast so that failing 19 times before working on the 20th attempt didn’t cause any timeouts. We now have monitoring of success-rate-per-endpoint, so even if 99.9% of requests are successful, if a single function crashes 100% of the time we should still notice :)
- fsociety 6y agoRace conditions can definitely do that, speaking from experience. I’ve fixed some that have been in the source code for a year before they were discovered.
- deleted 6y ago[deleted]
- JoshuaDavid 6y agoIf you have a marketplace of some sort, and users of your new code (at 95% rollout) cannot see postings made from some subset of other users, but the 5% on the old version still can, you might not notice that volume of sales of things posted before the rollout has dropped by 95% where you would notice if it dropped by 100%. Most of the time this would happen, the difference between "95% of people can't see a posting" and "100% of people can't see a posting" would be pretty small, but I guess if the marketplace is very liquid the difference might actually be substantial. But unless you're uber or someplace like that, where there's a big difference between "95% of drivers can't see a ride request (so it takes slightly longer to find a driver)" and "100% of drivers can't see a ride request (so no driver is ever found)", I can't imagine too many cases where the difference is likely to make a practical difference.
- heavenlyblue 6y agoHow does the system decide which code needs to be deleted?
- mkr-plse 6y agoThe flag management system has information on activity pertaining to various flags. After prolonged inactivity, a flag is considered stale and a diff is generated. This is a heuristic (in the absence of expiry date for a flag) and the final decision on cleanup is made by the flag owner.
- ThouYS 6y agoI'm sorry, but I don't entirely get this. Does it mean that after I've finished something I want to merge, I have to remove 5% of that? Or of the entire code base? Or something completely different?
- rkangel 6y agoI think they're saying that any new feature controlled by a flag is only rolled out to a maximum of 95% of the user base. To make sure it gets rolled out to 100% you have to remove the flag controlling it (and therefore the old version of the code).
- bluesign 6y agoThis is solving half of the problem though, failing experiments still clutter the code.