23 ms·
What I wish someone had told me about tensor computation libraries
- cygaril 6y agoSeems to have missed the existence of jax.jit, which basically constructs an XLA program (call it a graph if you like) from your Python function which can then be optimized.
- easde 6y agoTorchScript JIT (torch.jit.script) is similar for PyTorch.
- komuher 6y agoNot even cloese, jax.jit allow you to compute almost anything using lax.for_loops, lax.cond and other lax and jax contsturts pytorch jit does not allow that its just extra optimization for static pytorch functions.
- hyperbovine 6y agoNo autodiff for most of these though.
- alevskaya 6y agoJAX autograd will work on most any jitted fn - the control-flow limitations are no autograd for code with for/while loops since there's a statically unknowable trip count through the loop body. Much looping code can be handled differentiably using a "scan" though.
- JHonaker 6y agoIn the section title, JAX: > But JAX even lets you just-in-time compile your own Python functions into XLA-optimized kernels...
- nestorD 6y agoThe authors gives that quote (from the JAX documentation) but does not seem to interiorize it as his conclusion says: > This is the niche that Theano (or rather, Theano-PyMC/Aesara) fills that other contemporary tensor computation libraries do not: the promise is that if you take the time to specify your computation up front and all at once, Theano can optimize the living daylight out of your computation - whether by graph manipulation, efficient compilation or something else entirely - and that this is something you would only need to do once. It is exactly what JAX does. There is a computational graph in JAX (its encoded in XLA and specified with their numpy like syntax), it is build once, optimized and then runs on the GPU.
- yongjik 6y ago> with dynamically generated graphs, the computational graph is never actually defined anywhere: the computation is traced out on the fly and behind the scene. You can no longer do anything interesting with the computational graph: for example, if the computation is slow, you can’t reason about what parts of the graph are slow. Hmm, my experience is the opposite. When I used Tensorflow, there was no way I could figure out why something is slow, or require huge memory. All I have is a gigantic black box. Meanwhile, in PyTorch, all I have to do is run it with CUDA_LAUNCH_BLOCKING=1, and it will give me an accurate picture of exactly how much milliseconds each line is taking! (Just print the current time before/after the line.) With nvprof it will even tell you which CUDA kernels are executing. * Disclaimer: Haven't dabbled in ML for ~a year, so my view might be outdated now.
- whimsicalism 6y agoEh. I love pytorch, but it can definitely be difficult to reason about at times. For instance, due to async dispatch on GPU, you could get assertion errors where a line fails, but the real error was actually several lines above. That was difficult to reason about.
- atorodius 6y agoWouldnt this be fixed by CUDA_LAUNCH_BLOCKING=1? Or putting a bunch of torch.cuda.synchronizes in the suspected lines.
- whimsicalism 6y agolol whoops yeah that would definitely solve the problem. I wasn't familiar with `CUDA_LAUNCH_BLOCKING` but `torch.cuda.synchronizes` does work.
- 37ef_ced3 6y agoNN-512 (https://NN-512.com https://NN-512.com) Generate fully vectorized, stand-alone, human-readable C99 code for neural net inference, and understand exactly what's happening. For example, watch the code run with Linux's perf top and see the relative costs of each layer of the computation. Total transparency, no dependencies outside the C POSIX library
- DSingularity 6y agoYummy. Thanks. Gonna bookmark that one.
- joshuamorton 6y agoIn what sense is this "better"? The generated code is like __m512i wfs16 = _mm512_castsi256_si512(_mm512_cvtps_ph(wf25, _MM_FROUND_TO_NEAREST_INT|_MM_FROUND_NO_EXC)); fs16 = _mm512_inserti64x4(wfs16, _mm512_cvtps_ph(wf26, _MM_FROUND_TO_NEAREST_INT|_MM_FROUND_NO_EXC), 1); _mm512_mask_storeu_epi32(wfPtr1+230400+38400*i5+768*c2+128*k1+64*m2+16*f3, 3855, wfs16); _mm512_mask_storeu_epi32(wfPtr1+345584+38400*i5+768*c2+128*k1+64*m2+16*f3, 61680, wfs16); (which is a set of 4 lines that appear in the middle of an ~800 line function). That's not "human readable". Sure you can use asan or gdb, but if gdb profiles slowly, what can you do? You're still at the mercy of the code generator to be able to optimize things.
- 37ef_ced3 6y agoGoogle those _mm512_... intrinsics (they are part of GCC) to see what they mean. The code you pasted is converting single-precision floats to half-precision floats, and storing the half-precision floats to memory, 32 at a time. That's filter packing, which happens during initialization (and never during inference) I agree, if you don't know anything about how convolution is implemented (filter packing, data packing, matrix multiplication, sum unpacking), you could be lost. But it's very shallow compared to a JIT or CUDA library scheme, and a knowledgeable ML performance engineer would have no difficulty The inference function (at the end of the C file) is a series of blocks, each block corresponding to a convolution or other complex operation. It's straightforward to see which, by looking at where the weights come from (a field in a struct that has the same name as the layer in your graph) If you use perf top (for example) you can see which convolution was most expensive, and why. Does the shape of the tensor produce many small partial blocks around the edge, so the packing is inefficient (a lot of tile overhang), for example? You can see that by glancing at the code and seeing that there are many optimized blocks around the edges. As a rule, if NN-512 generates small code for a tensor (few edge cases) you have chosen an efficient tensor shape, with respect to the tile Or you might find that batch normalization is being done at inference time (as in DenseNet), instead of being integrated into the convolution weights (as in ResNet), because there's fanout from the source and a ReLU in between. You can see that easily in the generated code (the batch norm fmadd instructions will appear in the packing or unpacking code) Is the matrix multiplication slow because there are too few channels per group (as in ResNeXt)? Easy to see in perf, make your groups bigger. Are you using an inefficient filter shape, so we have to fall back to a slower general purpose convolution? You can easily see whether Winograd or Fourier was used And so on
- dragandj 6y agoLet me chip in with some self-promotion. This book explains and executes every single line of code interactively, from low level operations to high-level networks that do everything automatically. The code is built on the state of the art performance operations of oneDNN (Intel, CPU) and cuDNN (CUDA, GPU). Very concise readable and understandable by humans. https://aiprobook.com/deep-learning-for-programmers/ https://aiprobook.com/deep-learning-for-programmers/ Here's the open source library built throughout the book: https://github.com/uncomplicate/deep-diamond https://github.com/uncomplicate/deep-diamond Some chapters from the beginning of the book are available on my blog, as a tutorial series: https://dragan.rocks https://dragan.rocks
- kyllo 6y ago"Deep Learning in Clojure with Fewer Parentheses than Keras and Python" Love it! :D What better way to define a neural network in code than an S-expression?
- piokoch 6y agoI am not sure if I am that enthusiastic. The problem with Lisp is not a number of parenthesis but where they are and what is their role. In c-like languages parenthesis help parser compiler but they also help humans to read the code. In case of Lisp they are just for the sake of the parser. Let's look on the code: Python: model = Sequential() model.add(Conv2D(32, kernel_size=(3, 3), activation='relu', input_shape=(28, 28, 1))) Clojure: (defonce net-bp (network (desc [128 1 28 28] :float :nchw) Which one is more readable? looking on the Clojure code I see 128 1 28 28 thrown on me, without digging in the documentation I have no idea what's happening.
- susam 6y ago> Which one is more readable? Both are equally readable to me. Now, granted 128 1 28 28 can be difficult to understand without documentation but that is not due to Lisp's fully parenthesized prefix notation. The Clojure code would also look equally readable if it had used keyword arguments. Are you sure you are not confounding familiarity with readability? With Lisp, after a while, the parentheses become invisible to the programmer.
- Const-me 6y agoI wonder does any of them have proper Windows support, i.e. DirectCompute? CUDA is NVidia only and vendor lock in is bad for end users. Both CUDA, OpenCL and VK require large runtimes which are not included in the OS, software vendors like me need to redistribute and support it, I tend to avoid deploying libraries when I can.
- cmarschner 6y agoTensorflow 1.0 has its roots in how Theano was built. Same thing, a statically built graph that is run through a compilation step, with a numpy-like API. So what makes Theano such an ingenious concept while TF is regarded as “programming through a keyhole”?
- dr_zoidberg 6y agoHere's my take about TF (in general, not particularly 1.x or 2.x): Like many things from Google, I always had the impression that the library, while better than alternatives at the time, is too tailored to Google use cases. And if you fall outside of them, bad luck. Still, at work we find it easier to deploy and interoperate with other tools than Pytorch. Hell, we have a guy working in Pytorch who converts his work to ONNX so that we can then connect those to some tooling we already have from back when TF was our only backend. Could there be a better way? Perhaps. But we have to ship models and TF "just* works" (with a big asterisk, yeah).
- bravura 6y agoI recently used TF 1.0 (former Theano author, current PyTorch user) and found TF 1.0 to be hellaciously difficult to grok and seemed to include a lot of unnecessary abstractions. There was existing TF 1.0 code I was trying to extract gradients through (nsynth-wavenet). I spent over 8 hours on it unsuccessfully; I asked for help from a friend at Google who worked on TF and he couldn't figure it out either. I emailed the original author of the code and he acknowledged that he didn't know how to do it either, and he had an old notebook he could dig up that kinda would work with a lot of fixes.
- bravura 6y agoAlso see my comment here: https://news.ycombinator.com/item?id=25439073 https://news.ycombinator.com/item?id=25439073 I am definitely interested in a higher-level Pytorch API that uses TF as an execution engine.
- 6y ago
- albertzeyer 6y agoI was not aware that the PyMC developers have forked and continued Theano: https://github.com/pymc-devs/Theano-PyMC https://github.com/pymc-devs/Theano-PyMC It seems very active right now. Here some further information: https://pymc-devs.medium.com/the-future-of-pymc3-or-theano-is-dead-long-live-theano-d8005f8a0e9b https://pymc-devs.medium.com/the-future-of-pymc3-or-theano-i... I haven't really found references to its new name "Aesara". Apparently, the main new feature for Theano will be the JAX backend. I wonder though, my experience when working with Theano, and also deep with the internals (trying to get further graph optimizations on theano.scan): - Some parts of the code are not really clean. - The code is extremely complex and hard to follow. See this: https://github.com/pymc-devs/Theano-PyMC/blob/master/theano/scan/op.py https://github.com/pymc-devs/Theano-PyMC/blob/master/theano/... - This also made it very complicated to perform optimizations on the graph. See this: https://github.com/pymc-devs/Theano-PyMC/blob/master/theano/scan/opt.py https://github.com/pymc-devs/Theano-PyMC/blob/master/theano/... - In this specific case, it's also a problem of the API: theano.scan would return the whole sequence. But if you only need the last entry, i.e. y[-1], there is a very complicated optimization rule which checks for that. Basically many optimizations around theano.scan are very complicated because of that. - Here is one attempt for some optimization on theano.scan: https://github.com/Theano/Theano/pull/3640 https://github.com/Theano/Theano/pull/3640 - The graph building and esp the graph optimizations are very slow. This is because all the logic is done in pure Python. But if you have big graphs, even just building up the graph can take time, and the optimization passes will take much longer. This was one of the most annoying problems when working with Theano. The startup time to build the graph could easily take up some minutes. I also doubt that you can optimize this very much in pure Python -- I think you would need to reimplement that in C++ or so. When switching to TensorFlow, building the graph felt almost instant in comparison. I wonder if they have any plans on this in this fork. - On the other side, the optimizations on the graph are quite nice. You don't really have to care too much when writing code like log(softmax(z)) -- it will optimize it also to be numerically stable. - The optimizations also went so far to check if some op can work inplace on its input. Which made writing ops more complicated, because if you want to have nice performance, you would write two versions, one which works inplace on the tensor, and another one not. And then again 2 further versions if you want CUDA as well.
- 6y ago
- prideout 6y agoAre these libraries ever useful in non-deep learning applications? It sounds like Theano is a bit more general purpose, but why would I ever need it outside of a deep learning context? I wonder if it could be used for something crazy, e.g. setting up a graph that generates shadertoy-like images on the GPU.
- 6gvONxR4sf7o 6y agoThey are. Lots of numerical code benefits from GPU and lots of numerical code benefits from derivatives. Simulations, solvers, numerical optimization, good old fashioned statistics.
- timkpaine 6y agoIdk about using these libraries, but its almost impossible to find generic graph libraries that aren't designed around either ML or alternatively scheduling batches. One such example is my own, https://github.com/timkpaine/tributary https://github.com/timkpaine/tributary
- nerdponx 6y agoInteresting library & idea, almost like its own programming paradigm when you abstract away all the specificity for building software or running ETL jobs or whatever. But this is a completely different kind of graph. The graphs being discussed here are differentiable DAGs of mathematical computations.
- timkpaine 6y agoIs it that different? https://github.com/timkpaine/tributary/blob/main/docs/examples/autodiff/autodiff.md https://github.com/timkpaine/tributary/blob/main/docs/exampl...
- nerdponx 6y agoNow that is definitely interesting. And you have some notion of "differentiability" for all of your various sources, sinks, and transforms? That said, Tensorflow and Pytorch are both very much general purpose numerical computing libraries. You don't have to use them for neural networks.
- jstrong 6y agoI'm a theano diehard, and I'll never get over how google came along, introduced a shittier version of theano, garnered worldwide acclaim for it, and killed the better library in the process.
- alevskaya 6y agoHaving written and debugged both Theano and TF plenty in the past, I think this is a somewhat uncharitable take, esp. recalling the absolutely enormous Theano compile times. :) I think Theano was genius, but a system that relied on python-string-based C++ code-emitters was always going to have trouble with long-term sustainability.
- bravura 6y agoI am one of the authors of the Theano work. I am happy to hear that the Theano project is now being maintained again. I will agree with alevskaya that the compilation times were an issue in my particular research ten years ago. I was trying to build neural-networks for parsing that were created at run-time. Since each parse tree had a different computation graph, I was not able to use Theano since it required compiling every single type of parse tree computation graph it encountered during training. [edit if you want more details: There is really interesting old-school work called "Recursive distributed representations" and later "Labelling recursive auto-associative memory" that used auto-encoders to consume a variable length sequence, e.g. text string, in a sequential fashion. My work with Yoshua Bengio---incomplete---was based upon the idea of doing unsupervised binary parsing of sentences using a hierarchical RAAM-style approach: At any given point in time, greedily find the two adjacent tokens that could be most easily compressed into one token with low reconstruction error. However, once you apply this recursively and end up with auto-encoding binary parse trees, you end up with a variety of different computation graphs, each of which required separate compilation.]
- deleted 6y ago[deleted]
- deleted 6y ago[deleted]
- nautilus12 6y agoThe last arguments about why you would want a static graph and even it's drawbacks and complaints sound basically similar to why you would want to do functional programming
- bravura 6y agoI will say that I am very excited by the tftorch.py effort from @sillysaurusx: https://twitter.com/theshawwn/status/1311925180126511104 https://twitter.com/theshawwn/status/1311925180126511104 The idea being that pytorch can just be a high-level API executing lower-level tensorflow under the hood.
- deleted 6y ago[deleted]
- dangirsh 6y agoRelated: The Simple Essence of Automatic Differentiation - Conal Elliot - https://www.youtube.com/watch?v=ne99laPUxN4 https://www.youtube.com/watch?v=ne99laPUxN4 - https://arxiv.org/abs/1804.00746 https://arxiv.org/abs/1804.00746
- galaxyLogic 6y agoWhy are they called TENSOR computation libraries?
- MaxBarraclough 6y agoAs someone who knows nothing about this area: > I get confused with tensor computation libraries (or computational graph libraries, or symbolic algebra libraries, or whatever they’re marketing themselves as these days). Aren't tensors a sort of generalisation of matrices? How are they equivalent to graphs?
- ogogmad 6y agoThe word tensor in this context refers to a multidimensional array, not to a tensor in the mathematical sense. The computation graph is simply a representation of a sequence of arithmetic operations that you're performing on some data.
- MaxBarraclough 6y agoI see, thanks.
- PoignardAzur 6y agoCan someone ELI5 what are the differences between the different libraries are? The article uses a lot of jargon, an something that frustrates me about getting into machine learning is that teaching material will either abstract away what the internals do or assume that you already know how the internals work. Some specific questions: > They provide ways of specifying and building computational graphs Is the article talking about neural networks? As in, arrays of arrays of weights, where input values go through successive layers, and for each layer the same instruction is applied to some values with the respective weight? Or is it talking about a graph as in, a functional graph, where manually written functions call other manually written functions? (hence why a later paragraph talks about if-else statements and for loops) > Almost all tensor computation libraries support autodifferentiation in some capacity (either forward-mode, backward-mode, or both). What are those? From the wikipedia article, it sounds like autodifferentiation basically means running f(x+dx)-f(x), but if there are entire frameworks handling it, then there's probably something fancier going on. > According to the JAX quickstart, JAX bills itself as “NumPy on the CPU, GPU, and TPU, with great automatic differentiation for high-performance machine learning research”. Hence, its focus is heavily on autodifferentiation. The earlier description makes it sound like JAX does some cutting-edge compilation stuff to transform semi-arbitrary functions (with ifs and else and loops and stuff) into a function that returns it derivative. So how can that stuff run on the GPU? It sounds like there would be a lot of branching code. And how is that related to machine learning / neural networks?
- sidhu1f 6y agoFor the heavy lifting of the actual linear algebra computations, these tensor computation libraries typically use some variant of BLAS or eigen.