3 ms·
Experiment, Measure, Repeat
- nnutter 9y agoThey measured long enough in this specific example that they avoided the problem but something I see people repeated forget is to establish a baseline before making a change. If you don't know how variable something is before you make your change you might naively think you made something better or worse when it's within normal variability.
- ecaml 9y agoGood point, and measuring things as they are before experimenting can also serve as a valid motivation to start an experiment.
- oftenwrong 9y ago"avoid maintaining useless things" If your company is small, heed this warning. If you're spending a lot of time on maintenance of non-essential things, and struggling with supporting legacy things while creating new things, you are missing a huge opportunity. Right now, "trimming the fat" is as easy as it will ever be. As your company grows, it will only become more difficult. Be ruthless now with maintaining focus on things that "move the needle", and with killing things that don't. Analogy: It's like you're gearing up for a long hike. There is a natural tendency to take things along "just in case", even if you know you probably won't need it. At the trailhead, you could easily leave some of that non-essential stuff in car. You try on your pack, and it doesn't feel that heavy. You think "what's the harm". Of course, you are still fresh and full of energy at the trailhead. A few days into your hike, you start to really feel the weight of that extra stuff. The pack digs into your shoulders and hips. Your entire lower body is sore. You regret bringing the extra stuff, but now that you are in the wilderness, you cannot just dump it. You have to carry it back with you, and you wish you had exercised restraint when you had the chance.
- maxxxxx 9y agoI am working more and more on this. We have a ton of legacy code we drag along because nobody understands it so it's too scary to touch. I have started to push the idea that it's simply not acceptable to have code we don't understand. We have now refactored several parts that were painfully complex and convoluted by analyzing what they really do and then rewriting or changing. Most of the time it is not that difficult once you commit to the task.
- mwexler 9y agoIs it just me, or is this example not an experiment? There was just pre and post change measures, with no comparison group. The measure of success was "use more threads" which was the same as the treatment, instead of the actual goal, making an improved perception that a Slack channel was faster and easier to read (and potentially improving productivity: faster ship, more tests passing 1st time, etc.). A better method might have been something like: Pick two channels with equal traffic and relevance to business, require one channel to emphasize threads and respect quiet, let the other go on as they are, compare groups at the end on the actual metric of concern (of users in each group, quick check/survey of perceived utility, ease, value of the channel). Could even have done same measure in the beginning as well to show change over time comparison. Still not random selection, but better. "But that's a lot more work", some might say. But without this extra work, there is no actual experiment. The test just says that threads are good and when we ask folks to use them, they do. But look at the data; Slack message count decreased from Q4 17 to Q1 18; any change in actual utility could be due to seasonality, shipping vs. bug bashing, other changes that resulted in fewer messages so threading wasn't needed; maybe threading caused fewer messages, maybe threading caused people to choose not to comment when they would have otherwise... but we can't tell from this design. I'm not saying there's anything wrong with iterating. But call it that: "Iterate and change, measure, repeat if change correlated with goodness". A formal Experiment is designed to show that the change you made _caused_ a change in something else, something important to you or your business. Without a formal experiment, you just have correlation, hope, tribal knowledge, instinct, experience... all great things, none of which support causation. And not everything needs this level of rigor, and that's totally fine; maybe that's the case in top post example. But if a change is expensive in terms of workflow, effort, or actual expense, perhaps it's worth doing a more structured test before committing. If you've never had an experimental design experience, try reading anything at http://exp-platform.com/ http://exp-platform.com/ (Kohavi's work at Microsoft) or search for DOE design of experiments at your favorite search engine; also articles on "A/B Testing" often give suggestions on how to best structure a controlled experiment. And recognize that most older work focuses on traditional ANOVA and t-tests, but there are all sorts of other modern ways to assess impact. (edit: corrected typos)
- ecaml 9y ago
- cjf4 9y agoThis sounds an awful lot like the lean/six sigma (lss) tool set that every company of a certain size and age has experience with. But that's not to invalidate the ideas here, as lss can unfortunately be prone to a cultish zealotry that mutates the original principles.