3 ms·
Per point 1, after interviewing a good few candidates in the space I think that most of them just don't have the backgound to effectively handle version control
by fundamental 5y ago
Per point 1, after interviewing a good few candidates in the space I think that most of them just don't have the backgound to effectively handle version control (it's more complicated when working with evolving datasets), automated reporting, or reproducible builds. They might have a CS background, but not the software engineering skills to execute any of those tasks effectively. So, whenever there's a time constraint from management they're going to skip those more formal software engineering practices to the organizations longer term detriment.
- PaulHoule 5y agoActually "data rich products" are complicated because they involve evolving code and data. There are tools that are good at one but not tools that are good at both. For instance Docker is a problem instead of a solution if part of your system is a several-gigabyte language model that changes frequently.
- fundamental 5y agoI tend to agree with tools either being good at code or data. I'm sure eventually there will be a solution, but for now it's a lot of bespoke tooling. Docker can be a headache at times, but being able to maintain a consistent environment is very handy. For the large binary resources (trained models, dataset, etc) I've found it easy enough to just use docker volumes to mount read only resources. Other people have resorted to leaving read only assets on the local network which might be fast enough for your needs. As long as you're not unintentionally copying data, mounting a TB of data read only takes no time in my experience.