5 ms·
I've never really understood the point of these visualizer things. The idea that a model is always well represented by a directed acyclic graph seems extremely
by brrrrrm 2y ago
I've never really understood the point of these visualizer things. The idea that a model is always well represented by a directed acyclic graph seems extremely dated.
I really would love a PyTorch/JAX profiler that shows, in annotated Python, where your code is allocating memory, using compute or doing device copies.
- lamename 2y agoIt may be that the way you like to think about things and the way others like to are different. I find that quickly grasping a new architecture is easiest with a graph-based diagram first. Then code for details. All with the goal of internalizing the information processing steps. Not memory allocation per se. In my mind, how the network implementation allocates memory is a different question. But I think both of our desires just reflect our jobs, our interests, and simply how our brains conceptualize things differently.
- brrrrrm 2y agoI think it's a trap of visual elegance. When you start thinking of models this way you miss the way a lot of models are actually written. E.g. how do you represent an online fine-tuning process? I want to randomly switch between a reference impl and an approximation method, but when using the approx method I want to back-propagate so that it gets better over time. full disclosure: I've written plenty of these little visualizers and also fallen for the trap of "everything should be a declarative graph."
- gedy 2y agoNot mocking you, but I'd make a guess you don't like to draw block diagrams when discussing designs or code architecture with others either? Some folks aren't "visual" thinkers, and took me a long time working to realize that some folks are like that.
- almostgotcaught 2y agoLol I interned inside pytorch a few years ago (you and I even met/talked about tangential things :)) and worked on tracking such allocations (although I didn't hook it up to profiler). Spoiler alert: you can't track such provenance because everything gets muddled in the dispatcher. EDIT: not completely accurate to say you can't do it. I prototyped a little allocator that would stamp every allocation (the pointer itself, in the unused bits, a trick I learned from zach) with the thread id and a timestamp (just an incrementing counter) and then percolate that up to the surface. Obv that didn't land lol.
- brrrrrm 2y agoI wonder if you could track provenance by operating at the highest layer of the dispatcher and capture any calls to GPU operations (ala Cuda Graph)? > in the unused bits I feel like this is a PT rite of passage :P
- chillee 2y agoWe actually do track such provenance now (https://pytorch.org/blog/understanding-gpu-memory-1/ https://pytorch.org/blog/understanding-gpu-memory-1/) - works pretty well I think :)
- almostgotcaught 2y agowell you should tell bram then :p but also while generally "in all things i defer to horace" (ok not really) so maybe i'm not looking closely enough (and missed it) but the bottom of that stack shows (roughly) the autograd dispatch key and not the python call site (or some such). and maybe it's a pedantic difference (depends on what bram wants) but i wanted provenance back to the TS op so that i could then do static memory allocation things with that representation (now i've probably fully de-anonymized myself...) and for that use-case, even what you have now, isn't enough (you can't get a total sum for how much each TS op or whatever allocates and when the corresponding free happens).
- mathematicaster 2y agoI don't see any reference to acyclic as a requirement.
- brrrrrm 2y agoI think it uses TF's graph construct which has that built in? it's like a weird mix of dataflow and control flow graphs.
- hatthew 2y agoA DAG visualization of a model is a good abstraction to learn the general structure of the model to help contextualize the code you're reading.