6 ms·
This is really interesting although a little dense/cryptic as slides. Does anyone have a link either to a slightly more leisurely exposition of the theory behin
by mjw 14y ago
This is really interesting although a little dense/cryptic as slides. Does anyone have a link either to a slightly more leisurely exposition of the theory behind it, or perhaps a more user-focused summary of what kind of features and optimisations this new magic might enable?
As I understand it from a quick non-expert skim (please correct!): the idea is to use a fairly flexible and expressive metadata language under the hood to specify the physical memory layout of a multi-dimensional array -- more expressive than numpy's existing support for strides. Sounds like this will allow arrays to be reshaped and joined together without copying, support arrays larger than core memory, etc.
I wonder if the language is too powerful, if this could make memory management (e.g.: can I free this contiguous chunk of memory?) quite hard to reason about.
Also: seems like this makes it easier to push the cost of array layout manipulations in a somewhat lazy fashion from the producer onto the consumer of results. At which point does one decide that traversal is now excessively complicated and the thing should be copied into a flat layout. Especially since this might affect one's ability to use certain BLAS routines, take advantage of certain hardware optimisations etc. Is the idea that a compiler makes this decision for you, or you make the decision explicitly or implicitly?
Interesting stuff anyway, look forward to hearing more about it!
- pwang 14y ago> Does anyone have a link either to a slightly more leisurely exposition of the theory behind it, or perhaps a more user-focused summary of what kind of features and optimisations this new magic might enable? No, but those will be forthcoming as we build out more of it. We are targeting an end-of-November preview release, as the slides indicate, and it will give people a much more concrete idea of the things Blaze can do. > the idea is to use a fairly flexible and expressive metadata language under the hood to specify the physical memory layout of a multi-dimensional array Yep, that's exactly right. The challenge is to constrain this is that it is not too general, i.e. enters into PhD research land, but still useful for a large number of use cases. One of the core ideas is to more coherently represent the metadata about location, locality/compute affinity, and index spaces. Numpy has the beginnings of some of these ideas in its various flags and View object semantics, but it's all stuck in its very rigid, contiguous-memory origins. > At which point does one decide that traversal is now excessively complicated and the thing should be copied into a flat layout. This would be a great Master's research project. :-) The hope and the belief is that one can punt on resolving this in generality, and still solve a number of interesting, concrete problems in an efficient way.
- mjw 14y agoNice one. Sounds like you guys have thought hard about the right balance between expressivity and allowing these structures to become excessively hard to reason about or traverse. Incidentally I found the use of topological manifold terminology (atlas of coordinate charts, homeomorphisms, ...) in the slides quite intriguing, are there topological connections here or is that more just an analogy?
- freyrs3 14y agoThere is a connection, we can view mappings between linear memory spaces to higher arrays constructs as a series of coordinate transformations in the same way we deal with transformations between charts on manifolds. Except in our case our charts are necessarily disjoint, so the boundary conditions on charts are trivial.