5 ms·
Do bindings like this help much with the day to day work of machine learning research? Are the kinds of errors you encounter the kind that could be avoided with
by Gormisdomai 8y ago
Do bindings like this help much with the day to day work of machine learning research? Are the kinds of errors you encounter the kind that could be avoided with a smarter type checker?
- Cybiote 8y agoThere places where types will help when reading unfamiliar code. They'll also help when you're putting together more complex algorithms for novel architectures. They're less helpful if you're only stacking layers like lego-blocks and tuning hyper parameters of existing models. If you look at the issues, you'll see the developer contemplating GADTs to deal with tensor shape matching (but even phantom types will do). This is an issue that causes headaches for just about anyone whose dealt with neural nets.
- yawaramin 8y agoI would say so. There are 4,500 results on Stack Overflow for '[pandas] typeerror': https://stackoverflow.com/search?q=%5Bpandas%5D+typeerror https://stackoverflow.com/search?q=%5Bpandas%5D+typeerror
- Barrin92 8y agoI'm not sure that's a good metric. After all type checking does not really answer a question a developer has about type errors, it just informs you of a type error before you run your program. In fact I'd wager your average Haskell type error has generated more than a few frustrated stackoverflow questions.
- yawaramin 8y agoFair point. I guess it's fair to say that a typechecker would have avoided (almost) all runtime type errors and a fair portion of compile-time type errors provided the developer had type-guided tool assistance.
- Cybiote 8y agoType errors in functional languages (except Elm) are so unhelpful, you just get used to their general shape and your context to figure what the problem might be. Eventually these are rarely ever a problem but I agree that for beginners, type errors can be quite hostile. I'll say however, the benefit of an ML like Ocaml is less so the types than that functional languages are the most advanced in providing tools which allow for domain modeling in terms of composition and defining tailored algebras. This is something that fits both applied and experimental machine learning especially well.
- whateveracct 8y agoRight, but learning about the errors without having to run and hit every codepath in your program has a lot of value. So in that way, seeing how many people hit pandas runtime type errors is a useful thing to measure.
- Gormisdomai 8y agoI guess instead of "smarter" I should have said "smarter than what python does already".
- GIFtheory 8y agoSo, this doesn't apply to torch specifically, but to give an example from my own experience... I have a tensorflow model that takes 45 minutes just to build the graph. You can imagine how frustrating it is to wait most of an hour to run your code, only to see it crash due to some trivial error that could have been caught by a compiler. Although this is somewhat of an extreme example, I'd say most of my tensorflow models take at least several minutes to build a graph. Anything that reduces the possibility of runtime errors would therefore speed up my development cycle dramatically. I'll also note that the OCaml compiler is lightning-fast compared to say, gcc.
- Cacti 8y agoAre you sure you’re writing it in a reasonable way? I’ve made some awfully large nets (implementing current state of the art models) with TF and I’ve never ran into anything like that. I mean I might have 30 seconds to a minute before it gets moving but that’s the most I’ve seen, and that’s including the entire init process (reserving GPUs, preprocessing enough data to keep the cache moving, initializing variables for optimizers and such, and so on). Some of these models have an absolute ton of ops. What are you doing that requires a 10, 30, 45 min graph build?
- GIFtheory 8y agoI'm fairly sure the model is implemented in a reasonable way. It's an experimental deep generative model based on https://github.com/openai/glow https://github.com/openai/glow, though more complex because the warp and its inverse are evaluated at training time, and the outputs fed to other things. The warp has around 200 layers, IIRC. The model requires keeping track of the evolution of the log-determinant of the warp after each operation, along with the derivatives of those things... so the graph can get pretty huge.
- Cacti 8y agoInteresting, thanks.