3 ms·
> Giving meaning to ... Ingesting such amounts of data is a challenge indeed. But problems will become much more complicated if it is necessary to perform comp
by asavinov 8y ago
> Giving meaning to ...
Ingesting such amounts of data is a challenge indeed. But problems will become much more complicated if it is necessary to perform complex analysis during data ingestion. Such analysis (not simply event pre-processing) can arise because of the following reasons:
* It is physically not possible to store this amount of events. For example, assume you collect them from devices and sensors
* It is necessary to make faster decisions, e.g., in mission critical applications
* It can be more efficient to do some analytics before storing data (as opposed to first storing data persistently and then loading it again for analysis)
Such analysis can be done by conventional tools like Spark Streaming (micro batch processing) or Kafka Streams (works only with Kafka). One novel approach is implemented in Bistro Streams [0] (I am an author). It is intended for general-purpose data processing including both batch and stream analytics but it radically differs from MapReduce, SQL and other set-oriented data processing frameworks. It represents data via functions and processes data via column operations rather than having only set operations.
[0] Bistro: https://github.com/asavinov/bistro https://github.com/asavinov/bistro