3 ms·
That's a fair point, most of the applications I've worked on have been relatively simple (or I've just gotten lucky). It makes sense that different issues crop
by rocmcd 5y ago
That's a fair point, most of the applications I've worked on have been relatively simple (or I've just gotten lucky). It makes sense that different issues crop up at different scales, given the complexity of different workflows and package dependencies. I also haven't worked in ML before, so that may be a class of complexity all on its own.
- PaulHoule 5y agoTwo issues turn up in in ML. One is that the dependencies are a beast. The risk that it won't find a solution between a number of libraries that are only compatible with certain versions is high. The other one is that ML projects themselves contain data, often large amounts of data. For instance a "word2vec" style model might be 1 gigabyte and it might be a part of another model. (Say you turn the words to vectors then put the vectors through a CNN or RNN.) The Python packaging system might be a logically correct place to store this data (a necessary part of the the model) but it's a big file that will cause hassles if you do everything right, big hassles if you do things wrong (like compress anaconda packages with slow bzip2, use whatever algorithm that Docker uses to superamplify I/O, ...) It is nice to pack all your "empty" models (ready to train) as python packages and maybe even your "trained models" (some data files to supplement the empty files) but it will take some iteration to make it all go smoothly.