5 ms·
This must have piqued a number of readers interedts, mine included. Are you reffering to something like Matlab? Julia?
by Jenz 6y ago
This must have piqued a number of readers interedts, mine included. Are you reffering to something like Matlab? Julia?
- enriquto 6y agoThere's many programming languages where multidimensional arrays of floats are a natural language construct. From low-level (e.g. Fortran) to high level languages (e.g. Octave/Matlab, Julia). Even in C you can work with multidimensional arrays of floats, but it is a bit cumbersome. This is not the case for Python, where the language cannot express in itself this data structure, and it needs external library support like numpy. EDIT: I'm always very surprised by the wide acceptance of python/numpy for numerical algorithms, as it seems a notion quite foreign and awkward for this language. Python has excellent native support for string processing and advanced data structures like dictionaries, but those are mostly useless for numerical computation. If you want to do math in Python, you need to write libraries in other languages and provide a Phython interface. This is what scipy does, and it is a very good job, but it does not feel "natural" from the point of view of the language. If Python was a good scientific programming language, the most natural way to implement scipy/numpy and all that would be using Python itself, but this is not at all the case.
- pletnes 6y agoI think numpy is mostly great. Although I do agree - numpy should be in the stdlib, period. So many users come from the scientific/data/simulation communities. It’s just that there are no (or too few) core devs from that tradition. Python lacking numpy in the base language is as awful as Fortran lacking dict-like data structures in its runtime.
- enriquto 6y ago> numpy should be in the stdlib, period. I disagree strongly with this. Numerical facilities should not be in the stdlib, they should be in the language itself, without need to "import" anything. I can create strings and dictionaries without any imports. I should be able to create a multidimensional array of floats, and perform natural operations with it (like matrix multiplication) without importing anything. As per the dictionaries in Fortran, there are libraries for that. But I think that not having complex data structures in Fortran is a feature, not a bug. If you actually need these data structures, many Fortran programmers will tell you to just use a different language. You'll never hear such a plainly honest answer from Python programmers; they'll tell you instead to use bizarre libraries with unnatural interfaces, and you'll be forced to multiply matrices using a function "numpy.dot" or an ugly operator "@".
- fishmaster 6y agoThis is exactly why I'm going to Julia. It either has these constructs in the language itself or is powerful enough to express them adequately. Numpy is a necessary crutch for Python.
- pc86 6y agoYou're going to Julia so that you can avoid typing "import numpy as np" at the beginning of a file? Or is there more to it?
- fishmaster 6y agoAlso so that I don't have to use things like "np.dot(A, B)" and similar annoyances.
- axi1 6y agoAs a sidenote, I don't write "import numpy as np" in jupyter notebooks (at least as long as I don't share them with others) for I have this line in my ipython startup file.
- dklend122 6y agoMathematical programming in Julia is just more natural and integrated with the language and type system. From dot broadcasting, to multiple dispatch, to parametric types (if you want ), multidimensional array comprehensions. No worrying about 4 different container types (list, np array, pytorch tensor, symbolic etc etc) Everything works with one array abstraction, including gpu and multi threaded
- fishmaster 6y agoThis is what I mean with it being more natural and elegant in Julia. Also, "@code_warntype" is incredibly useful for efficient code.
- rbanffy 6y ago> I disagree strongly with this. Numerical facilities should not be in the stdlib, they should be in the language itself, Overloading the language with features that have limited use for most of their users would tie their release cycles and end up a disservice to both users of the language in general and those who use it for heavy number crunching. As for @, my guess is that Python is running out of ASCII symbols for operators (and unwilling to go full APL with operators). I would imagine that, if NumPy arrays don't support @ (don't they?!), it'd be a desire to be compatible with pre 3.5, which is when the @ infix operator was introduced. It shouldn't be much trouble to implement __matmul__ and __rmatmul__ for them, assuming NumPy's own interfaces are sane.
- aldanor 6y ago”Can not express this data structure"? It can of course, there's not much to express. Go and write a (slow) tensor library yourself in pure Python. To be fair, there's even a standard library `array` module that is a direct wrapper around C arrays, and NumPy shares some functionality with it (namely in regards to buffers, memoryviews, data type encoding etc). Edit: there's also a built-in buffer type that supports multi-dimensional strides, suboffsets and all that (which, again both `array` and `numpy` make use of in establishing a standardized buffer api).
- enriquto 6y ago> Go and write a (slow) tensor library yourself in pure Python. Paraphrasing Alexander Stepanov, "complexity is part of of the specification" [0]. If your data is not contiguous (as the third figure in the article shows), it is not the same data structure. Thus you cannot implement this data structure in pure Python, period. The "array" module is just a wrapper to C data just like numpy is, it is not a pure Python construction. [0] Here's the full quote, from http://www.stlport.org/resources/StepanovUSA.html http://www.stlport.org/resources/StepanovUSA.html In 1976, still back in the USSR, I got a very serious case of food poisoning from eating raw fish. While in the hospital, in the state of delirium, I suddenly realized that the ability to add numbers in parallel depends on the fact that addition is associative. (So, putting it simply, STL is the result of a bacterial infection.) In other words, I realized that a parallel reduction algorithm is associated with a semigroup structure type. That is the fundamental point: algorithms are defined on algebraic structures. It took me another couple of years to realize that you have to extend the notion of structure by adding complexity requirements to regular axioms.
- aldanor 6y agoIf you're talking about non-contiguity due to multi-dimensionality (or strided slicing), PyBuffer already internally supports that (and thus memoryviews, stdlib array, etc) - [0], is that "pure Python" enough? [0] https://docs.python.org/3/c-api/buffer.html#c.Py_buffer.strides https://docs.python.org/3/c-api/buffer.html#c.Py_buffer.stri...
- lgessler 6y agoWhat if NumPy and SciPy were part of the Python stdlib--would this change your view at all? (Note that for the majority of Python's scientific users, this might as well be the case, since they're using Anaconda to get their Python distribution set up.) Like I mean sure, maybe it'd be nice to have some syntactic sugar at the language level for tensors (though IMO the slice syntax for numpy.ndarray is already plenty), and maybe it's a microannoyance to have to write `import numpy as np` over and over, but what else is lacking?
- jofer 6y agoI've used fortran, C/C++, matlab, and IDL for many years but primarily work in python these days. I really do think numpy gets the interface correct. Modern fortran is quite nice as well, but I'm a _big_ fan of the way numpy does broadcasting. I think this is one thing it gets _much_ more correct than, say, matlab does. Also, the fact that numpy arrays are semi-contiguous chunks of memory just like a C arrays makes them _very_ easy to reason about. It's easy to predict memory usage and understand what operations will make temporary copies (assuming you understand how numpy works, anyway). The way numpy uses python's built in operators makes it feel very native. I really don't find it awkward at all. I used numeric and numarray in the pre-numpy days, and those did feel more "bolted on". While numpy is really similar to numeric, a lot of little things were fixed during the transition to make numpy very much a native part of python. Everyone keeps insisting that numpy/scipy/etc are all really "just written in other languages with a python interface", but that's _far_ from being true. Core parts are (e.g. ndarray itself + built-in ufuncs), and a lot more is wrappers around widely-used libraries (think BLAS), but you might be surprised at how much of numpy really is python. Ditto for scipy (though for purely historical reasons, scipy is a dumping ground for X-implemented-in-C stuff). Ditto for things like skimage/sklearn/etc. The bulk of it really is implemented in python. Sure, operations like + aren't, but it might be worth looking through the numpy codebase before stating that it's not written in python. Next, yes, if you naievely loop over a python array it will be slow, and this is a real limitation of the language + numpy model. Yes, you need to use a different mental model when implementing things. Some problems (e.g. finite-difference-like problems) are best implemented in a way that has poor performance with python+numpy. You do need to drop down to fortran/C/etc (cython is _great_ for those cases) when you have problems that can't be modelled in a particular way. However, that's not the limitation most folks seem to think it is. It's a well-known limitation that there's a ton of support for switching approaches when you hit it (e.g. cython, numba, f2py, etc). Yes, fortran and julia are both nicer in those cases (though my experience with julia is mostly "toy"/learning projects). However, the problems where python+numpy breaks down performance-wise are not anywhere near as common as folks seem to think they are. In my experience, the other types of non-scientific problems are _much_ more common, so it's nice to have a very general purpose language to draw on. As for matrix multiplication, is it really that bad to do things like x.T.dot(y) instead of x' * y? I really don't understand the argument that the `dot` method of an ndarray is ugly. Also, I actually hate that matlab/et al treat any 2d array as though it's a matrix and fundamentally don't have the concept of a 1d array (row vectors and column vectors are 2d, not 1d). I really do want a 1d array and not a row vector or a column vector most of the time. That's simply impossible in matlab. Also, if you want that behavior in python, it's absolutely possible! Use `np.matrix` instead of `np.array`. Arrays don't behave that way because element-wise operations are more common in practice. I get the impression you're looking at things from a Julia perspective. Julia is great, but similar to matlab and fortran it falls apart when you start to move outside its core domain. I wouldn't want to write a web service in Julia (though I'm sure it's possible). I do very frequently need to implement web services that do relatively heavy number crunching or deal with things that are fundamentally math-on-big-arrays-of-numbers. Python is _great_ for that type of application. It's also actually pretty good for many "hardcore" number crunching tasks. I've done things where python falls short, sure. However, python + numpy/et al hits a sweet spot that other languages currently don't in what it combines. Domain specific languages are better in some ways, but surprisingly little of a scientific codebase winds up being pure numerical operations. There's an awful lot of other stuff in there, and that's where the scientific python stack is incredibly effective.
- Chris_Newton 6y agoIf Python was a good scientific programming language, the most natural way to implement scipy/numpy and all that would be using Python itself, but this is not at all the case. To be fair, with modern hardware capabilities, I still wouldn’t be writing out linear algebra algorithms in the language itself even if I were programming in a relatively low-level language like C or C++. I’d be using an implementation of BLAS/LAPACK or some other suitable library, which in turn might well have been written in carefully tuned assembly language on the target platform in order to use any parallel or otherwise specialised operations provided by a CPU or GPU, partition data for optimal cache use, etc. The situation with numpy and other numerical libraries in Python seems analogous. We are still importing highly optimised code to do the numerical heavy lifting and then using a higher-level language to write the less performance-sensitive logic and glue everything together conveniently.
- dunefox 6y agoBut as soon as you actually need to implement an efficient library you must use C/C++/Fortran. Julia, for example, allows you to stay in Julia and get comparable performance.
- rbanffy 6y agoIf you need to implement an efficient vector math library on modern hardware you'll end up using intrinsics, padding structures for cache alignments, identifying which CPU, GPU, NPU, and vector extensions you are running on to better use specific features, and so on, and it won't look like idiomatic sequential C/C++/Fortran for almost every value of "idiomatic".