4 ms·
Thanks for the pointer. I can believe that a language that looks so different will find that different patterns and primitives are natural for it. My experienc
by imurray 3y ago
Thanks for the pointer. I can believe that a language that looks so different will find that different patterns and primitives are natural for it.
My experience from writing a lot of array-based code in NumPy/Matlab is that broadcasting absolutely has made it easier to write my code in those ecosystems. Axes of length 1 have often been in the right places already, or have been easy to insert. It's of course possible to create a big mess in any language; it seems likely that the NumPy code you saw could have been neater too.
In machine learning there can be many array dimensions floating around: batch-dims, sequence and/or channel-dims, weight matrices, and so on. It can be necessary to expand two or more dimensions, and/or line up dimensions quite carefully. Einops[1] has emerged from that community as a tool to succinctly express many operations that involve lots of array dimensions. You're likely to bump into more and more people who've used it, and again it seems there's some overlap with what Rank does. (And again, you'll see uses of Einops in the wild that are unnecessarily convoluted.)
[1] https://einops.rocks/ https://einops.rocks/ -- It works with all of the existing major array-based frameworks for Python (NumPy/PyTorch/Jax/etc), and the emerging array API standard for Python.
- mlochbaum 3y agoOf course the method with length-1 axes will look good if it's the only way to do broadcasting. I've hardly touched it (and never saw the NumPy code mentioned in the quote), but some people in the APL community have so I'm reporting on their thoughts. Although length-1 axes already being in the data sounds worrying: you have an array that doesn't depend on an axis, but there's an axis anyway to show where it would go if it did? Named axis systems like einops are much better regarded. I've seen some implementations of similar functionality in APL, and various discussions of how it could be built into an array language (I have my own design where that's a core idea, but not using APL syntax).
- imurray 3y ago> Although length-1 axes already being in the data sounds worrying: you have an array that doesn't depend on an axis, but there's an axis anyway to show where it would go if it did? Example: A / A.sum(axis=3, keepdims=True) Make all of the vectors along axis=3 sum up to one. There are other ways of doing it, but this way seems fairly clear to me. The shape of the denominator is the same as the numerator, except for a 1 in position 3. Unfortunately we have to specify `keepdims`, because the default of `False` removes the dimension being summed over, which doesn't work in general. `keepdims=True` is the behavior in Matlab/Octave, so the example becomes A ./ sum(A, 4) with 4=3+1 because Matlab is 1-based like Fortran.