5 ms·
> I'm curious who or where the author heard that from (not specifically the people themselves but the domain they are in). In the telecom domain, I've dealt wi
by pg314 9y ago
> I'm curious who or where the author heard that from (not specifically the people themselves but the domain they are in).
In the telecom domain, I've dealt with data big enough that Python wasn't really feasible. Think 100 of millions of records in CSV format that need to be parsed and processed. Doing that in Python is going to be painful.
- classybull 9y agoPython is insanely fast at data processing and analysis because it has very fast libraries. As a matter of fact, don't know if you've heard, but data processing it kind of like.. Python's thing...
- mattkrause 9y agoYou're violently agreeing with each other. Python itself can be pretty slow. Doing image processing on data stored as list-of-lists-of-integers would be brutally slow. On the other hand, numpy is an import away, and it can be quite fast, especially if it's been built with an optimized BLAS/ATLAS, etc.
- AstralStorm 9y agoBy blazingly fast you mean 100x slower than C++ equivalent and only 20x slower is you're very careful to avoid accidental copies. For reference, MATLAB is about 30x slower with no special care. Pure Java on Hotspot was 5x slower except it dies on big data input due to very slow GC and goes to 50x slow. Source: handled big audio data from hdf5 database, gigabytes sized. C++ equivalent had no vectorization or magic BLAS or anything.
- joshuamorton 9y agoAs I'll often say to these comments, then you're doing things wrong. Numpy code can be written to never leave the numpy sandbox, and at that point it should be as fast or faster than naive c++ (because you'll be getting SSE and stuff for free). There's a reason almost all deep learning is done in python.
- pg314 9y agoNot all data is a good fit for Numpy: some data is non-numeric or not a homogenous array. > There's a reason almost all deep learning is done in python. The heavy-lifting in e.g. TensorFlow is done in C++. Bindings to Python make sense because it is one of the few sanctioned languages inside Google, and it is widely used outside of Google and easy to pick up.
- joshuamorton 9y ago>The heavy-lifting in e.g. TensorFlow is done in C++. Bindings to Python make sense because it is one of the few sanctioned languages inside Google, and it is widely used outside of Google and easy to pick up. That's exactly the same as with numpy. I'm not sure what your point is. C++ is also one of the few sanctioned languages inside google, as is Java. >Not all data is a good fit for Numpy: some data is non-numeric or not a homogenous array. I'm curious what kind of data you're working with that can't be represented and effectively transformed in a tensor (numpy array).
- pg314 9y ago> That's exactly the same as with numpy. I'm not sure what your point is. I was replying to "there's a reason why...". You didn't specify that reason, so from the rest of your comment I took it to mean that Python (with numpy) was fast and good enough to write deep learning stuff. That doesn't seem to be the case for TensorFlow. > I'm curious what kind of data you're working with that can't be represented and effectively transformed in a tensor (numpy array). I'm not intimately familiar with the internals of numpy, but my understanding is that the basic data structure is a (multi-dimensional) array of values (not pointers). That leads to a number of questions. If you have an array of records (dtype objects), and one of the fields is a string, am I correct that each element needs to allocate memory to hold the longest possible value that can occur for that field? What if that is not known beforehand? How do you deal with optional fields (e.g. int or null)? Do you need to add a separate boolean to indicate null? How do you deal with union types, e.g. each record can be one of x types, do you make a record that has a field for each of the fields of those x types? Do those fields take up space?
- pg314 9y ago> Python is insanely fast at data processing and analysis because it has very fast libraries. It doesn't have fast libraries for everything. E.g. for the use case I brought up, it doesn't (or at least none that I could find at the time). Not all data are homogenous arrays that fit in numpy. And are those libraries written in Python? Or in C because Python is too slow? If you take your reasoning, any language that can bind to C (which is pretty much any language) is as fast as C. That is not very helpful when comparing the speed of languages. A slower language will force you to drop into C sooner than a faster one. > As a matter of fact, don't know if you've heard, but data processing it kind of like.. Python's thing... Things like audio and video codecs, or crypto code aren't written in Python.
- agentgt 9y agoIts interesting you mention CSV as I had an issue with CSV with Java many years ago. I was trying to cut a couple of columns out of a CSV and tried to write it in Java and for some reason either the CSV library was slow or maybe I didn't have the right buffered input but I ended up writing a python script that outperformed my Java code (which reminds me I meant to revisit that). Anyway here is the python code (it might not be the exact code I used as I didn't update it but whatever): https://gist.github.com/agentgt/1383185 https://gist.github.com/agentgt/1383185