8 ms·
NumPy 1.20
- jphoward 6y agoI _love_ numpy, and I am getting excited about jax, too. However, I do have one request for it. Getting the argmax of a multi-dimensional array, in terms of the array's dimensions, is difficult for new users. np.argmax(np.array([[1,2,3],[1,9,3],[1,2,3]])) is 4, rather than (1,1). I understand why, but it seems strange to me that argmax cannot return a value the user can use to index their array. Having to then feed that `4` into unravel_index() with the array's shape as a parameter seems less elegant than say passing a parameter of "as_index=True" to the argmax.
- kakadzhun 6y agoConsider this: In [1]: np.argmax(np.array([[1,2,3],[1,9,3],[1,2,3]]).flat) Out[2]: 4
- montebicyclelo 6y agoAlternatively you could use flat: a = np.array([[1,2,3],[1,9,3],[1,2,3]]) idx = np.argmax(a) a.flat[idx] # 9
- 6gvONxR4sf7o 6y agoDoes that work the same way with strided arrays?
- montebicyclelo 6y agoAssuming you mean what I think you mean, it does work. e.g. a[::2, ::3].flat[idx], where idx is from 0 to width*height of the view (idx can also be a NumPy array, for getting multiple values)
- user2049 6y agoIt is great to see type annotations! A huge step in the right direction.
- enriquto 6y agowe are slowly approaching the ease of use and efficiency of fortran. In a few decades more we'll be there!
- gugagore 6y agoI assume you say it at least a bit in jest, and with a good deal of sincerity. Did you ever know an interactive environment for producing Fortran code? Anything like a REPL?
- cycomanic 6y agoIn development, but there is lfortran: https://lfortran.org/ https://lfortran.org/
- klyrs 6y agoModern fortran has classes, too! I've occasionally joked that fortran 25 and python 4 are to be the same language...
- teddyh 6y ago“I don't know what the language of the [future] will look like, but I know it will be called Fortran.” — Tony Hoare, 1982
- jabl 6y agoThe next step after type annotations will be having the type specified by the first character in the variable name.
- michaericalribo 6y agoNumPy is incredible to me: it serves not only as a critical bare-metal layer, but also an essential tool in its own right, and I can’t imagine using Python without it. I do numerical work—ML/data science/statistics—and I simply couldn’t accomplish any of what I do without the functionality provided by NumPy. np.array alone is worth all the king’s gold. It is interesting to me that in many respects it serves as “guts”. I definitely drop down to use numpy directly with regularity. But it’s also possible to do 80-90% of the job without ever explicitly using the module itself. It’s baked in everywhere, to the point that it feels like just another standard library. Exciting to see progress! Keep up the good work, numpy team :)
- gspr 6y agoI agree with everything you say, but I don't understand this part: > it serves not only as a critical bare-metal layer What's bare-metal about NumPy??
- michaericalribo 6y agoHeh, I’m probably misusing that term—forgive me, I’m just a script kiddy! :D I meant that it can serve the role of low-level, nitty gritty machinery, entirely separate from its use case as a module unto its own right. Eg, pandas (in some sense) is just a convenience layer on top of numpy—but to me that’s like saying any piece of software is “just a convenience layer on top of python.” It’s partly a means to an end! Not just an end unto itself. Numpy enables a whole new class of functionality, independent of its direct use as part of my software. Is there a better word for that?
- 6y ago
- chalst 6y agoQuick summary: "This NumPy release is the largest so made to date, some 684 PRs contributed by 184 people have been merged." > Annotations for NumPy functions. This work is ongoing and improvements can be expected pending feedback from users. > Wider use of SIMD to increase execution speed of ufuncs. Much work has been done in introducing universal functions that will ease use of modern features across different hardware platforms. This work is ongoing. >Preliminary work in changing the dtype and casting implementations in order to provide an easier path to extending dtypes. This work is ongoing but enough has been done to allow experimentation and feedback. > Extensive documentation improvements comprising some 185 PR merges. This work is ongoing and part of the larger project to improve NumPy’s online presence and usefulness to new users. > Further cleanups related to removing Python 2.7. This improves code readability and removes technical debt. > Preliminary support for the upcoming Cython 3.0. Type annotations seem the biggest deal to me. I'd say if you care a lot about SIMD and the performance issues, you should be thinking of moving to Julia: it's still a valuable technical achievement.
- notagoodidea 6y agoI would rephrase your statement in : "If you care about SIMD, performance issues and type annotations, you should look into Julia". Numpy is an incredible piece of software and provides performance for one of the most mainstream language. It has been one of the main building block in the python takeover in data science, ml, etc. But if I had the choice, I would have move to Julia during my precedent work/projects as soon as it reached v1.
- deleted 6y ago[deleted]
- chalst 6y agoThe type annotation story is indeed better with Julia, but having type annotations for NumPy is beneficial for many users for whom Julia isn't a win, where number crunching isn't the main thing going on and Python's better library situation is important and you want to avoid the complication of calling Python from Julia.
- RocketSyntax 6y agoHmm. I am checking dtypes via `np.floating` and `np.signedinteger`. What will this change to?
- calhoun137 6y agoA really great way to learn more about numpy is with Math Inspector[1]. It creates a block coding environment which works with the entire numpy/scipy stack, it also has an interactive 2d and 3d plotting library that updates the functionality in mathplotlib. Also, I made it =) Just spent the past 3 months working day and night and released a massive update a few days ago. Check out the youtube video on the bottom of the page for more in depth information. [1] https://mathinspector.com/ https://mathinspector.com/
- ksm1717 6y agoFantastic. Jupyter is a handicap disguised as dataman’s best friend. This is actually intuitive
- calhoun137 6y agoThank you!!! I worked so hard on this project, and really appreciate the kind words
- Scene_Cast2 6y agoCould you elaborate more on Jupyter? I use it almost daily, and despite its warts (such as lack of source control and a lack of explicit dependencies), it's pretty darn useful.
- ksm1717 6y agoNo doubt useful, I just find that it’s failings are whisked away as “you don’t need to worry about source control because it won’t work!” It prevents users from learning things that would make their lives easier because it can’t do those things, it gets them stuck at a local optimum. It’s great for what it’s great for but as soon as you leave its usefulness domain jupyter becomes a burden to integrate even in a simple way.
- 411111111111111 6y agoThere are ide plugins to use the jupyter kernels thought. Atoms hydrogen[1] or vscodes Python Support [2] comes to mind. Jetbrain IDEs have similar plugins. They enable pretty much both package managers and scm without any overhead to speak of. > 1: https://blog.nteract.io/hydrogen-interactive-computing-in-atom-89d291bcc4dd https://blog.nteract.io/hydrogen-interactive-computing-in-at... > 2: https://devblogs.microsoft.com/python/data-science-with-python-in-visual-studio-code/ https://devblogs.microsoft.com/python/data-science-with-pyth...
- ZuLuuuuuu 6y ago"sliding_window_view" method will be a great addition. Sliding windows are used quite frequently when analysing data and currently you can kind of achieve it using "as_strided" method but that method is very cumbersome to use. But from the examples given, "sliding_window_view" is much easier to use.
- isoprophlex 6y agoYes! This is a big win for me! Basically, across projects, I've been reusing a snippet that uses some as_strided magic for years now. The snippet looks seriously deranged, it will be great to refer to something built in... also for my colleagues who now have to understand my as_strided shits.
- sandGorgon 6y agoThe most interesting question here is the newest competitor to Numpy - Tensorflow - https://www.tensorflow.org/guide/tf_numpy https://www.tensorflow.org/guide/tf_numpy With the added advantage that TF is natively accelerated on M1. Will TF-Numpy maintain parity ?
- qoqosz 6y agoTF-Numpy is not a competitor to NumPy. It just introduces a small subset of NP API to TF codebase. In TF you still have immutable tensors, so e.g. sth like: tensor[mask] = new_value doesn't work.
- u678u 6y agoNumPy/Pandas is one of the reasons I'm stuck on Python. I'd prefer a strongly typed language like go/java/rust but there isn't the same library or community. Any recommendations? Its for applications not just DS/ML so Julia isn't really in scope.
- throwaway10102 6y agoLatest Python (> 3.7?) typing combined with mypy --strict is the best of both worlds. Highly suggest you try it out.
- eigenspace 6y agoJulia is not just a datascience and ML language. It's a very nice general purpose programming language.
- optimalsolver 6y agoPython is strongly typed. Did you mean statically typed? You can use optional type-hinting in Python now, and install MyPy as a type-checker.
- jjgreen 6y agoI've been getting some odd warnings (thousands of them) from CI runs against 1.20, typical: some-file.py:83: DeprecationWarning: `np.bool` is a deprecated alias for the builtin `bool`. To silence this warning, use `bool` by itself. Doing this will not modify any behavior and is safe. If you specifically wanted the numpy scalar type, use `np.bool_` here. Deprecated in NumPy 1.20; for more details and guidance: https://numpy.org/devdocs/release/1.20.0/notes.html#deprecations return self.z[j][i] But all of the lines mentioned in these warning do not reference np.bool
- Lev1a 6y agoI haven't seriously used Python for quite a while so I don't know if this is possible/usual, maybe there's some kind of code generation happening at runtime that would throw off line numbers or added/removed/substituted lines compared to the on-disk script files?
- sleavey 6y agoAre you using libraries that use `np.bool`? These could trigger warnings to stderr that you would see.
- jjgreen 6y agoThat was the issue: https://github.com/numpy/numpy/issues/18281 https://github.com/numpy/numpy/issues/18281
- throwaway9930 6y agoIf you know Python, then trying NumPy is pain. Because for some reason data scientists have managed to misuse the python syntax. E.g. `np.ogrid[ -200000000:200000000:100j,-500000:500000:100j]` it's completely unexplainable what this does. They have managed to overload index/slice operator and imaginary numbers to produce two arrays. (Example taken from here https://asecuritysite.com/comms/plot06 https://asecuritysite.com/comms/plot06 )
- mrcactu5 6y agoThere's a random shuffle function. Instead of writing one out using an algorithms textbook. Lots of data types get randomly shuffled during coding. Shuffles arrays and lists. https://numpy.org/doc/stable/reference/random/generated/numpy.random.shuffle.html https://numpy.org/doc/stable/reference/random/generated/nump...