4 ms·
I think this is it. I still don’t know what benefits poetry adds over the typical venv + pip + requirements flow, because that flow has worked pretty well for m
by rocmcd 5y ago
I think this is it. I still don’t know what benefits poetry adds over the typical venv + pip + requirements flow, because that flow has worked pretty well for me.
Is there something crucial I’m missing? I feel like package management has largely been solved in python, but I admittedly haven’t done a lot of research into what other workflows are out there.
- PaulHoule 5y agoIf you think ‘pip works’ then the programs you are building are too simple. I worked at a place where we did machine learning builds and neither pip or conda worked reliably. What was insidious was that it almost worked and people thought they could live around it’s limitations which turns every 2 hour job into a 15 hour job. (But they are data scientists…. A reliable build is like Bigfoot, cold fusion or Linear A to them.) Pip’s dependency solving strategy is not correct, especially if you install pip packages one at a time. A real dependency solver would download the metadata for all matching versions (can be done with two seeks for a wheel but gotta run setup.py for an egg…) do an smt solve and then install the wheels. Pip just starts installing packages, hopes for the best, sometimes backtracks, often gets stuck or does the wrong thing.
- rocmcd 5y agoThat's a fair point, most of the applications I've worked on have been relatively simple (or I've just gotten lucky). It makes sense that different issues crop up at different scales, given the complexity of different workflows and package dependencies. I also haven't worked in ML before, so that may be a class of complexity all on its own.
- PaulHoule 5y agoTwo issues turn up in in ML. One is that the dependencies are a beast. The risk that it won't find a solution between a number of libraries that are only compatible with certain versions is high. The other one is that ML projects themselves contain data, often large amounts of data. For instance a "word2vec" style model might be 1 gigabyte and it might be a part of another model. (Say you turn the words to vectors then put the vectors through a CNN or RNN.) The Python packaging system might be a logically correct place to store this data (a necessary part of the the model) but it's a big file that will cause hassles if you do everything right, big hassles if you do things wrong (like compress anaconda packages with slow bzip2, use whatever algorithm that Docker uses to superamplify I/O, ...) It is nice to pack all your "empty" models (ready to train) as python packages and maybe even your "trained models" (some data files to supplement the empty files) but it will take some iteration to make it all go smoothly.
- lmns 5y agoHow do you pin all of your dependencies to a specific version with just venv, pip and a requirements.txt? How do you upgrade them later?
- rocmcd 5y agoVersions are specified within the requirements.txt, if that's what you mean. Upgrading can be done by installing a later version of the package and rewriting the requirements file. I suspect my use cases haven't been as complicated as some of the others listed in this thread, which may be why I've never felt the need to look into poetry and others.
- lmns 5y agoI know what you mean, but as soon as you have transitive dependencies it doesn't work at all. At that point you can't reproduce the same state of your venv at a later date because some minor version of a dependent package could change and maybe break your build.