3 ms·
Graphistry is really cool but the point here is not to compete on the visualization side, we needed a tool that scales for a few of our use-cases (mostly as a h
by mbuda 4y ago
Graphistry is really cool but the point here is not to compete on the visualization side, we needed a tool that scales for a few of our use-cases (mostly as a highly configurable graph visualizes) + a library that can easily be extended :D
- stuntkite 4y agoHeh. You and me both. When I worked there something I really wanted to get to was GIS integration and abstraction with graph. They are doing some really, really cool stuff and I think their future offerings will be things to take note of, but it's a service too and it's not open source. If you find a solution to what you're talking about, I'm interested. Let me know. Also, I talk to Leo. If you have specific requests that you want to list, I'll make sure he reads this thread.
- taubek 4y agoWhat kind of integration with GIS where you considering? What did you want to accomplish?
- stuntkite 4y agoWell I haven't let go of the idea and am in the process of releasing a GIS data processing platform, but that's a story for another day. My thought is that all data, especially large datasets from the top down only has a few ways it can be displayed and sliced. Relationships, multi temporal (slicing and or playing by time but then also playing back multiple time sections so they can be compared. compute vs. real time, time to exec on server vs time experienced by user, and also finding various time slices to compare. So multi temporal being that your timeline playback isn't just a film strip), and spatial. The ground truth for all data is that it happens somewhere in the world so a graph should be able to be put into a GIS space for display, bonus points for 3D, but then also you should be able to build composite fake GIS spaces and visualize the relationships between them. Like for instance, your products network traffic all going to colocation spaces and to users and back again. That has normal GIS locations and that can provide constraints for graph display but also, especially in a distributed system there may be artificial geography that can be defined, lets call it "The Astral Plane" that has relationships that are important and grouping just like state and country boundaries that can be defined in shape files and put into a spatial database in a projection you just make up or is defined by the graph data and physical locations and you should be able to slide between all those things. That's where I wanna go and still intend to get to. EDIT: Obviously when I say everything has GIS coordinates I'm not quite talking about outer space, but I think outer space is even covered by the rest of this idea. So is tiny space. This realization came to me when I was reading about people using PostGIS as a database for chemistry simulation. I have no link for that, I haven't thought about it in years, but now that I'm talking about it, I'll try to find it and post a link if I find it as I would like to readdress.
- taubek 4y agoInteresting take on GIS application. I used to work only with ArcGIS some 25 years ago. It was pretty new concept for me at that time. The whole spatial concept of linking databases with geo coordinates.
- lmeyerov 4y agoThanks for kind words all :) Scale - backend: As we're rapids.ai-native (helped start early days of both Apache Arrow + Nvidia RAPIDS.ai), we work with customers doing billion-level nodes/edges in interactive time on GPU servers. Mostly for fast ingest, ETL, + graph neural nets / manifold learning, and we're slowly pushing that into the visual stack. Scale - frontend: We normally recommend reducing down to about ~2M edges or less. For sensible visual experiences, add in auto algorithms that cut to more like 500K. We've planned a way we think we can do another 100X, just not (yet) an engineering priority. Fun fact: your browser's JS VM is limited to ~1GB of RAM, so we're already at that limit in practice. RE:Scaling graph visuals as an engineering practice, it looks like memgraph is starting where neo4j reached a few years ago, and makes sense. That approach doesn't really work well for the use cases memgraph advertises for, because as soon as a bunch of user/customer/IT/etc events happen and get visualized, the browser crashes. Optimization approaches like wasm and workers are clever -- a v0 prototype of graphistry did that! -- but we found is too unreliable to be the path for good performance across users of most operational teams ("Works on my machine" syndrome). Do it, but that shouldn't be the main source of 100X performance, just a 2X boost. We end up connecting GPUs in the browser to GPUs in the datacenter not just for 100X'ing this kind of stuff, but for a predictable performance way that limits how often your user's browser crash on real datasets. Also maybe not obvious, this article focuses on interactive rendering, but a lot of the challenge after they figure out how to solve it is interactive analytics too. Most layout algorithms have non-linear complexity, so O(500K) is actually a challenge. A lot of our GPU offloading work nowadays isn't just rendering but layout, ETL, ML/AI clustering, etc. This ends up overwhelming the browser (why we do distributed GPU), and OLTP graph DB's aren't good at that either -- Neo4j basically had to write a V2 DB-in-a-DB to make their Graph Data Science module perform. And no worries: There's no competition because Graphistry isn't a graph database :) Most of our users will do something like databricks dashboard / jupyter notebook / powerbi / etc query <> graphistry visual. I bet pairing a great streaming db like memgraph with Graphistry would combine respective engineering strengths quite well!