4 ms·
Since this seems to be trending: any questions about Linkurious or graph databases are welcome. Linkurious CTO here :)
by david_p 11y ago
Since this seems to be trending: any questions about Linkurious or graph databases are welcome. Linkurious CTO here :)
- lennartkoopmann 11y ago... and Graylog CTO here. Ask us anything! :)
- alanpost 11y agoA question for both of you. I don't understand how, if you have it together enough to centralize your logging, you need to be told about interdependencies between your services. What am I missing? Do I underestimate how easy it is to set up centralized logging? Or how complex a deployment becomes before you even wonder if you need centralized logging? I have a centralized logging system, but I can't imagine being so confused about it that I need my logs to tell me that two components interact with each other. What don't I understand?
- lennartkoopmann 11y agoIt is fairly easy to set up centralized logging. Interdependencies between services can be pretty complex. Think about it from the operations and not the development side: Which firewalls are between the public internet and service X? Why are two Windows workstations talking to each other? Which services are talking to an internal API? The moment more than a few people are involved with your systems, it can get so complex that visualizing dependencies can be extremely helpful and bring a lot of insights.
- alanpost 11y agoYour appeal to think about it operationally helped me understand. I have to answer questions like the ones you pose whenever something odd pops up in the logs. Thank you.
- lmeyerov 11y agoAs a concrete example: we work with enterprises with people numbering anywhere from 10K to 500K to government-scale, and each person may have a desktop/laptop/phone, and all the servers/printers/switches those connect to, and at the logical layer, all the applications and services for making it useful. We'll see multiple central logging systems, hierarchies of administrators, and the results of mergers, acquisitions, and one-off or zombie projects. These organizations are getting sophisticated enough to log 10M, 1B, etc. alerts a day (ex: using graylog or splunk), so we need to focus on the next step of being able to point to one alert and asking what's happening around it. It's a really fascinating data problem, so we've been loving building tools for seeing into it!
- henrikjohansen 11y agoIndeed it is. Ingesting 100k events per second into one or more centralised log management platforms will not be efficient it your're relying on a row-based analyses approach.
- henrikjohansen 11y agoImage having to do that for 30k machines, 8-9k access points across 100+ different locations accessing hundreds of different systems - it does not work efficiently with visualising the dependencies automagically.
- rspeer 11y agoI abandoned Neo4j -- and graph databases in general -- years ago because there was no reasonable way to load in several million edges that were not already in a graph database. Does Neo4j have a better importing story now? I see a blog post from 2014 [1] that makes importing merely a million edges between 500 nodes sound like it's still a terribly difficult operation, giving me the impression that graph databases aren't quite ready for <s>big</s> medium data yet. [1] http://jexp.de/blog/2014/06/load-csv-into-neo4j-quickly-and-successfully/ http://jexp.de/blog/2014/06/load-csv-into-neo4j-quickly-and-... If, for example, I wanted to load an N-Triples file that's approximately the size and shape of DBPedia, can I reasonably do so? What tools should I use to get the job done quickly without descending into a Java nightmare?
- lmeyerov 11y agoWe settled on "medium"-batching into Titan, and sharing your experience, I'm hoping the Datastax acquisition means ingest will improve. For the bulk of our work, we do what you'd expect -- load terabytes into HDFS, (Py)Spark for straight SQL and some join helper functions, and occasionally, add some GraphX scala libraries. I'm curious -- what did you end up doing, & for what?
- rspeer 11y agoI maintain ConceptNet [1]. It's a large-ish, messy semantic graph linking together crowd-sourced knowledge (from Open Mind Common Sense, Wiktionary, some Games with a Purpose, etc.) with expert-created resources (WordNet, OpenCyc, etc.) [1] http://conceptnet5.media.mit.edu http://conceptnet5.media.mit.edu It's turned out to be a good input for machine learning about semantics, which has changed the goals of its representation a bit -- not only do I need to be able to load in data easily, I also need to be able to iterate over all of it. But some graph operations would be nice to have, too. Many technical people I describe the project to immediately ask me what graph database I'm using, both before and after the ill-fated semester of grad school where I actually tried to use graph databases. The answer to what I use now is: a bit of SQLite and some flat files. No need for HDFS, it still fits easily on a hard disk.
- schwarzmx 11y ago