3 ms·
> If you knew all the story and layout beforehand and you know you will be able to just plug in the numbers and make the paper stand, then you are not doing rea
by stevenbedrick 6y ago
> If you knew all the story and layout beforehand and you know you will be able to just plug in the numbers and make the paper stand, then you are not doing real science, you are playing the academic game, a cargo cult of science, a game stealing and consuming the prestige of science.
I actually agree with your immediate statement here, but that is not at all what I understood the OP to be saying. I read their "You should know what major type of finding..." argument as being one of starting with a concrete and well-formed research question, and carrying out carefully designed experiments, and thought that it was excellent advice.
In my little corner of computer science, I very frequently see people (at all stages in their scientific careers) start working on some new bit of research by a) getting a bunch of data, which they then b) feed into some nifty model du jour, and then c) spend a ton of time overcoming all manner of technical trials and tribulations, then finally d) get a number out the other side. They then e) find themselves totally stuck when it comes to actually interpreting their result, because before they ran their "experiment" they hadn't actually bothered to formulate a concrete hypothesis, and so it's not clear what they are supposed to _do_ with their shiny new number, or where to go next.
That's what people often don't get about science, whether it's wet or dry. The mechanical process of actually performing the experiment itself is usually the easy part, relatively speaking. The hard part is thinking carefully about the thing you're trying to study, formulating a theory, coming up with testable hypotheses, and designing experiments to perform those tests.
Part of that last phase involves planning ahead very carefully to what you're going to measure, what your control and intervention criteria will be, what specific statistical analysis you'll perform on the resulting data, and what your various interpretations will be. The more concrete and explicit you can make this, the better: "We're going to measure 'X' under conditions 'A' and 'B', because we think that 'X' will be a valid/useful measure of $PHENOMENON_WE_CARE_ABOUT, for reasons ____, ____, and ____. If X_A ends up being bigger than X_B, our interpretation will be ______, and if X_B is bigger than X_A, we will instead conclude ______'; if they are the same, that will suggest ______." [1]
Obviously you don't yet _know_ which of those conclusions you'll be drawing (if you did, it wouldn't be an experiment), but it is absolutely essential that you've gamed out the various possibilities to at least this level of detail _before_ you do the experiment. This is doubly true for exploratory analyses where you don't really have an intuition about what the outcome will be, as it helps keep the analysis from turning into an endless fishing expedition ("Well, what if I normalize this variable _this_ way? OK, what about _that_ way? ...").
In my experience, one of the best techniques for doing this is, yes, to actually write out blank versions of the tables that you think you'll need to tell the story of your experiment ahead of time, and to make dummy sketches of the various figures you'll need to help interpret the data. Not only will this help you clarify your thinking about what you are hoping to learn from doing the experiment, it has the added benefit of making sure that whatever code you write actually logs/outputs all of the needed data elements! More than once I've had to re-do an experiment because there was an important piece of data that I hadn't realized I would need until it was time to do the analysis. With just a bit more prior preparation, that poor performance would have been prevented.
To return to the OP's argument, they weren't saying that you should pre-specify your conclusions (which would be a terrible idea, for the reasons that you clearly spell out in your post). They were saying that you should have a plan about what specific experiments you're going to run and _how_ you're going to describe the motivation and results of those experiments.
And, if I may editorialize for a moment here, having a more structured approach to doing and writing about research can go a long way to helping to reduce the angst that comes with doing a PhD. I do very much think that many CS PhD programs are dropping the ball in terms of teaching experimental design and evaluation- but that is a rant for another time, as my TED talk today is already running long enough. :-D
--------------------------------
1: Very, very, very often, the process of formulating things this way takes several iterations, because usually once one is forced to write it out this explicitly, all sorts of little questions pop up- "Wait, is that actually what it will mean if X_A > X_B? What if means ____ instead? Hmmm... maybe I should be measuring X', instead? Oh, I'll need different data, in that case, because..."