3 ms·
Disclaimer: I work at Cloudera as a Tools Developer What do you mean by unstructured? Do you mean the data has yet to be parsed into a format which could be lo
by monstrado 13y ago
Disclaimer: I work at Cloudera as a Tools Developer
What do you mean by unstructured? Do you mean the data has yet to be parsed into a format which could be logically grouped into columns? Or do you mean that it's deeply nested?
Since log data doesn't really change, it might be overkill to use something like HBase (or any database for that matter). On the tools team at Cloudera, we've found that writing the data into HDFS and using Impala to analyze it works pretty well.
We typically analyze chunks of log data and then ingest it into HDFS (due to the use case), but if you're looking to ingest data in "real-time", you'll want to use something like Apache Flume.
With the data separated into partitions, we're able to run queries that analyze GBs of data in under a second (15 nodes). This is log data (LOG4J) that has been extrapolated into columns, and then loaded into a columnar storage format (RCFile, soon to be Parquet).
Let me know if you have any questions, glad to help.
- m0nastic 13y agoThanks, I've been testing out a bunch of hare-brained schemes and it didn't even occur to me to just use HDFS directly (and seeing you and karterk both suggest it helps). Basically, at a high level, the system I'm working on aggregates and processes security information (It's a SIEM, if that product category means anything to you). At the point the logs get ingested, the server determines if they're "actionable" (which is determined by rules I load into Redis), in which case it parses them and stores them in a Postgres event table; or "not individually actionable, but may cause an action in conjunction with some other log" that I want to just store somewhere for batch processing. I don't really need to tokenize those logs, as at the point I care about them I'm just going to be searching through them. So, they're "unstructured" in the sense that there's about 15 different collection points, each with it's own format (many just an ugly facsimile of syslog with some JSON in the middle). So, I think your suggestion will work out very well. Thanks again.
- monstrado 13y agoNo problem, glad I could help. Your use case sounds pretty interesting, HDFS should fit the bill for sure. You should take a look at Parquet (http://parquet.io/ http://parquet.io/) for storing your data. It's an open source columnar format that was designed for Hadoop, it even supports nesting (https://github.com/Parquet/parquet-mr/wiki/The-striping-and-assembly-algorithms-from-the-Dremel-paper https://github.com/Parquet/parquet-mr/wiki/The-striping-and-... <-- really interesting). Also, it already works with a lot of the Hadoop ecosystem components (MR, Hive, Pig, Cascading, Impala, ..), so your data doesn't have to move once it's in HDFS. Good luck!