4 ms·
Your complaint applies to array processing libraries in general. It applies to numpy. Yet those syntactic niceties are exactly what make them so convenient. Sur
by superbatfish 5y ago
Your complaint applies to array processing libraries in general. It applies to numpy. Yet those syntactic niceties are exactly what make them so convenient. Sure, it takes time to get comfortable writing array-oriented code. But it isn’t a good idea to optimize for complete beginners. Some familiarity with the general domain is a prerequisite to using array (and dataframe) libraries.
- matsemann 5y agoBut instead of having X ways of using the [] operator, it would be almost just as convenient if that were actual named functions. Then the editor could understand whats going on, one could jump to definitions, look up stuff, more easily actually learn that it's possible to do etc. I feel Notebooks aren't used as much because they're a great tool, but that they are used so much as a necessity since you need to look at live data to have any hopes of understanding what's going on. Why couldn't a compiler instead tell me what I'm trying to do will fail?
- goatlover 5y agoIt's just like learning any unfamiliar programming approach. Array processing has a long history, and it's supported by languages like R which are heavily used in data science. Once you understand the different uses of [] operator, it becomes second nature, and you realize how useful vectorization and broadcasting can be. Also, manipulating live data is an excellent way to learn about your data and figure out what you need to do. Python is a dynamic language, so interacting with a rich REPL environment makes a lot of sense.
- superbatfish 5y agoI agree with you in spirit. Having separate names is preferable to context-specific semantics for the []. In pandas, using [] alone is discouraged, in favor of using .loc or .iloc. But admittedly, that only partially improves the situation. Both accept multiple types (int vs int-list vs bool-list). Last year, I had an in-person discussion with a numpy core developer in which I proposed adding pandas-like syntax to numpy, e.g. allowing the author to use: a.boolmask[b == c] (but maybe not so verbose). A problem immediately arises: An array can have multiple axes, and each could be indexed with a different type within the same slicing call. One idea might be to allow the author to explicitly wrap the arguments with some dummy class to make clear what types they’re using: a[np.boolmask(b == c), d, np.intlist(e)] (Again, perhaps choosing shorter names in practice.) The idea would be that the wrapper merely “looks like” the thing it wraps, but contains an assertion to verify the type of the argument. Anyway, I agree that notebooks (or REPLs in general) are great for verifying your assumptions about how little snippets will behave. But I think they would be useful even in a fully statically-typed compiled language, too.