3 ms·
It may be that the way you like to think about things and the way others like to are different. I find that quickly grasping a new architecture is easiest with
by lamename 2y ago
It may be that the way you like to think about things and the way others like to are different.
I find that quickly grasping a new architecture is easiest with a graph-based diagram first. Then code for details. All with the goal of internalizing the information processing steps. Not memory allocation per se.
In my mind, how the network implementation allocates memory is a different question.
But I think both of our desires just reflect our jobs, our interests, and simply how our brains conceptualize things differently.
- brrrrrm 2y agoI think it's a trap of visual elegance. When you start thinking of models this way you miss the way a lot of models are actually written. E.g. how do you represent an online fine-tuning process? I want to randomly switch between a reference impl and an approximation method, but when using the approx method I want to back-propagate so that it gets better over time. full disclosure: I've written plenty of these little visualizers and also fallen for the trap of "everything should be a declarative graph."