45 ms·
Does anyone know of an application of M/R to non-text data like image data or time series data? I'm trying to think about how to process a hge set of 3D atmOsp
by metaobject 15y ago
Does anyone know of an application of M/R to non-text data like image data or time series data? I'm trying to think about how to process a hge set of 3D atmOspheric data where we are looking for geographic areas that have certain favorable time series statistics. We have the data stored in time series order for each pixel (where a pixel is a 4KM x 4KM area on Earth) and we compute stats for random pixels and try to find optimal combinations of N pixels/locations (where N is a runtime setting).
- eshvk 15y agoGenerally speaking (considering I don't know the detailed specifics) one problem you might run into is due to the density of image data which means that transferring data, sorting it become relatively expensive operations. Again, there might be chunks in the data pipeline where the MR framework might help.
- grncdr 15y agoDon't know of an application specifically like that, but I'd love to hear more, what kind of dimensionality does the per-pixel data have? (is it just RGB * N time series values?) There's nothing very text specific about Hadoop, just a lot of text-oriented examples. (e-mail is on my profile)
- ahalan 15y agoMapReduce is applicable wherever you can partition the data and process each part independently of others. I used Hadoop/Hbase for EEG time series analysis, looking for certain oscillation patterns (basically classic time-series classification) and it was an embarrassingly parallel problem: Map: 1. Partition the data into fixed segments (either temporal, say 1hr chunks or location based, say 10x10 blocks of pixels). Alternatively you can use a 'sliding window' and extract features as you go. In some cases you can use symbolic representation/piecewise approximation to reduce dimensionality, as in iSax: http://www.cs.ucr.edu/~eamonn/iSAX/iSAX.html http://www.cs.ucr.edu/~eamonn/iSAX/iSAX.html , "sketches" as described here: http://www.amazon.com/High-Performance-Discovery-Time-Techniques/dp/0387008578 http://www.amazon.com/High-Performance-Discovery-Time-Techni... or some other time-series segmentation techniques: http://scholar.google.com/scholar?q=time+series+segmentation http://scholar.google.com/scholar?q=time+series+segmentation 2. Extract features for each segment (either linear statistics/moments or non-linear signatures: http://www.nbb.cornell.edu/neurobio/land/PROJECTS/Complexity/index.html http://www.nbb.cornell.edu/neurobio/land/PROJECTS/Complexity... ). The most difficult part here has nothing to do with MapReduce but decide which features carry the most information. I found ID3 criterion helpful: http://en.wikipedia.org/wiki/ID3_algorithm http://en.wikipedia.org/wiki/ID3_algorithm, also see http://www.quora.com/Time-Series/What-are-some-time-series-classification-methods http://www.quora.com/Time-Series/What-are-some-time-series-c... and http://scholar.google.com/scholar?hl=en&as_sdt=0,33&q=time+series+dimensionality+reduction http://scholar.google.com/scholar?hl=en&as_sdt=0,33&... Reduce: 3. Aggregate the results into a hash-table where the keys are segment' signatures/features/fingerprints, and the values are arrays of pointers to corresponding segments (Based on the size this table can either sit on a single machine, of be distributed on multiple hdfs nodes) Essentially you do time-series clustering at the Reduce stage with each 'basket' in a hash-table containing a group of similar segments. It can be used as an index for similarity or range searches (for fast in-memory retrieval you can use HBase which sits on top of HDFS). You can also have multiple indices for different feature sets. ----- The hard part is problem decomposition, i.e. dividing work into independent units, replacing one big nested loop/sigma on the entire dataset with smaller loops that can run in parallel on parts of the dataset, when you've done that, MapReduce is just a natural way to execute the job and aggregate the results.