3 ms·
> I would think that the reduce function would need some sort of "accumulator" input, and that you'd only get one thing as an output, as opposed to more files o
by mayank 8y ago
> I would think that the reduce function would need some sort of "accumulator" input, and that you'd only get one thing as an output, as opposed to more files of data.
It may be easier to think of the reduce step more like a SQL GROUP BY rather than a function of a list. The map phase emits a bunch of (key, value) pairs, and all values with the same key are processed by the same reducer function (but each key gets a new reducer, modulo implementation details).
So in your paradigm, there are many reduce functions, each starting with a null accumulated value, resulting in many outputs rather than a single one.
- jedimastert 8y agoInteresting. That actually makes a lot of sense. I've done a touch more reading, and I actually only just now realized that the "word count" example doesn't count all of the words in a document, but rather all occurrences of each word. That clears a lot of questions up for me.