3 ms·
Full disclosure: I'm a software performance nerd who is building a deeper understanding of hardware using whatever tools are I work with a lot of assembly trac
by dwrodri 3y ago
Full disclosure: I'm a software performance nerd who is building a deeper understanding of hardware using whatever tools are
I work with a lot of assembly traces emitted from CPU sims in Verilator. These traces can be several gigabytes in size, with tens of millions of entries. I've found the typical interactive matplotlib backends to not work the best with tens of millions of points on my M1 Macbook Pro. This is no slight against Matplotlib, or cairo, or agg, or the default macOS backend—they prioritize cross-platform "It Just Works" support over extreme performance.
But it does irk me that there seems to be such a gap between the amount of data we can visualize in the gaming domain in comparison to the scientific domain[1]. The difference between 50 million and 5 billion is 100X. I'd be ecstatic if I could get within 10X difference. I understand that this gap exists for lots of reasons, of which I'd be happy to hear in detail in the replies to this comment.
1:https://youtu.be/eviSykqSUUw https://youtu.be/eviSykqSUUw
- thadt 3y agoGame engines spend a whole lot of work to not draw things. Even if you have an 8K monitor, that's still only around 33 million pixels. If the data set you want to visualize is 5 billion elements, and if you decided that you were going to use every single pixel on the screen, you'd still have to collapse those 5 billion elements in a visual space of less than %1 the size. In a video game it's easy to figure out what not to show - just hide everything that would normally be too small or far away to see. But in a scientific domain, what to show and what to hide might be a much more interesting question. On the other hand, if you're just looking for a plot of a few billion elements with the ability to zoom in and out, then a straightforward decimation of the data and drawing it to the screen can work. In the past I've written custom tools to do just that, but at this point in life I would probably throw it at something like SciChart and let it take care of that.
- rcme 3y agoGraphics is not as trivial of a problem as you're making it seem. Even if you're just rendering a single "thing", that thing can have an unbounded number of vertices. And it's definitely not a trivial problem deciding which vertices are important and which are not.
- thadt 3y agoOh, I just meant that it was conceptually easy - things should look 'right' - whereas with scientific data figuring out how to map that data to a 2D screen often involves a lot of carefully thought-out discretion about what to show and what to hide. The actual implementations of 3D rendering are gargantuan feats of engineering, science, and art. Having spent time writing code to turn terrain height maps into quantized meshes for 3D rendering, I agree entirely that deciding which vertices are important is not a trivial problem. The OP linked a video about Nanite, which IMHO is a marvel of engineering.
- rcme 3y agoMy point is that both you and the person you were replying to have a kind of fundamental misconception about why data is slow and graphics are fast. The misconception is that data isn’t slow and graphics aren’t fast, at least relative to one another. For graphics, the main performance metric is (generally) triangle count. You can draw lots of low-poly objects on screen more efficiently than you can draw one very high-poly object. The same holds true for data: you can render low amounts data more efficiently than you can render large amounts of data. Nanite doesn’t magically render high-poly meshes in real time. Nanite needs to pre-process the mesh and produces lower poly meshes that maintain the geometric properties of the original mesh. This is the main innovation of nanite, because traditionally it’s been very hard to reduce triangle count while keeping the overall geometry roughly similar to the original. And in this way, data processing has traditionally been much more efficient than 3D graphics. There have long been various statistical aggregations you can do on data to keep the same rough statistical properties using less data. But you have to be willing to preprocess the data, and that is slow. I haven’t used nanite, but I imagine the import process is also slow relative to the rendering speed after processing.
- johnnyanmac 3y ago>The misconception is that data isn’t slow and graphics aren’t fast, at least relative to one another. depends on the data/graphics. modern commercial GPUs are beefy (even when talking about integrated ones) and I suspect the Matlib kinds of tools aren't even tapping into a fraction of a percent of its power. But at the same time 5 billion draw calls raw will bring even a decent gaming GPU to its knees, at least for responsive, real time applications. The trick is to first understand your data (e.g. that 5 billion triangles are useless on a monitor that has 1-4 billion pixels). As you said, even Nanite isn't truly trying to process a trillion triangles raw. The steps from that understanding to a good enough approximation are indeed some dark magic.
- semi-extrinsic 3y agoHave you tried holoviews with datashader? https://holoviews.org/user_guide/Large_Data.html https://holoviews.org/user_guide/Large_Data.html
- bfrog 3y agoYou can't draw that data on the screen in any meaningful manner without zooming. At various zoom level what you actually want is a low pass filter over a slice of the data, such that the low pass filter matches the display. E.g. I can display 1080 pixels across, but there's 1 billion points. Ok... meaningful data is limited. Optionally you could do what other device displays do and show a sort of fuzzy intensity gradient at each display x axis over y points to represent the number of actual samples at that point. It's kind of crazy honestly. There's many options to display billions of points across 1080 pixel wide displays but none of them really show you the reality. Only some subset of information of the reality there.
- rcme 3y agoIt’s worth noting that, as far as I know, nanite involves preprocessing the mesh to convert it into nanite’s special format. You could do the same thing with data easily. The issue is taking a large batch data you’ve never seen before and displaying it efficient in real time, which nanite doesn’t do either.
- johnnyanmac 3y agothere are indeed several reasons, but in this case I suspect the problem isn't even with the matplotlib (well, most of it). You mention traces that are GB's in size and that's the first big problem. For reference, the newest Zelda game is a total of 18.2GB of storage, and of course Zelda isn't trying to audit the entire game on every load. Games spend a lot of time making sure their assets are lean, and at the end it is deployed in some binary format to further reduce its impact on a game. trace files that focus on human readability lose this compressibility. So regardless of how fast the graphical plotting capabilities are, I imagine such an app is CPU bound from simply trying to parse that trace data. That'd be the first place I'd look to optimize (you know, without properly profiling your app. Probably something I'd say in an interview setting).
- KolenCh 3y agoThere exists scientific visualization tool that goes far beyond what you can visualize in gaming domain, e.g. some applications are deployed on supercomputers so that you can change parameters interactively with the visualization. Gaming is highly specialized applications, so you should also be comparing to highly specialized applications in scientific domain.