11 ms·
Deep learning experiments in OCaml
- pc86 8y agoSo I was lost at the VGG19 example code, but probably because I have (a) no OCaml experience; and, (b) no ML/NN experience. Still seems interesting, though. If anyone has any suggestions on basic sources for getting a background on the concepts here I'd definitely give them a read.
- make3 8y agoPlease find Stanford's "Deep Learning for Computer Vision" https://m.youtube.com/playlist?list=PL3FW7Lu3i5JvHM8ljYj-zLfQRF3EO8sYv https://m.youtube.com/playlist?list=PL3FW7Lu3i5JvHM8ljYj-zLf...
- hackermailman 8y agoLook through youtube for university lectures, like these ones https://www.youtube.com/playlist?list=PL_Ig1a5kxu57NQ50jSuf0cjTe3zPLqv1O https://www.youtube.com/playlist?list=PL_Ig1a5kxu57NQ50jSuf0... Most intro classes just require familiarity with basic calculus (differentiation, chain rule), linear algebra and basic probability all of which you can just lookup directly on https://www.expii.com https://www.expii.com for a short tutorial. Toolkits are usually in Python or Lua, plus the numerous textbooks 'Deep learning with python' that are around and specific DL books such as http://www.deeplearningbook.org/ http://www.deeplearningbook.org/. Afterwards look around for Adversarial Learning, like detecting perturbations that force mis-classification and other attacks described in papers by Carlini and Wagner. Currently there isn't a perfect defense developed for all of these attacks, except robust optimization that provably defend some of them. Attacks are an interesting area in DL you can get into since we don't have access to large resources and can only do DL on a small scale (in my case anyway).
- mlthoughts2018 8y agoI had a very unpleasant interview regarding deep learning with Jane Street. I spoke to a member of their HR team to try to get significant assurances that the interview would actually be focused on deep learning and not puzzles or brain teasers, and that the job would really focus on deep learning for their actual business, and not just be a proxy for being generally smart and then work on whatever existing inhouse models. The HR employee reassured me significantly on both points. Then the interview was nothing but deck of card puzzles and random riddles where you have to articulate a careful model of some physical quantity like speed or frequency to solve the puzzle. I hate that junk, never found that it correlates with a way of thinking that matters in quant finance (which I previously did for a living) and suitably failed the interview. Worse, I would have been happy to decline that interview and tell them I know I’m not their guy if only the HR staff had correctly depicted the interview & job to me. Ok, enough grumbling. From this actual blog post, > “Type-safety helps you ensure that your training script is not going to fail after a couple hours because of some simple type error.” I really think this way of thinking about static typing is a very bad thing. This is not at all an actual benefit, because in any sane situation, you will use unit and integration tests that execute extremely quickly on small test data to exercise your end to end model training code. What I currently do for this on my team is to always require that model training programs are deployed inside of containers that capture not just the state of the code, but also make it configurable to mount the training data volume and pass in ENV that governs what the training job really is. So then Jenkins or whatever will build the container for any PRs that seek to implement or modify training, attach fixture data and fixture ENV settings, and give you quick feedback about the whole end to end training, even inclusive of GPU settings (we have to do a slight manual step to specify Jenkins running on a GPU server, but this is a vestige of some of our infra headaches). The point is that adding all sorts of extra code to embody type annotations, and limiting people from awesome dynamic typing features is a silly thing to do if you’re worried about type errors ruining a long-running job. That should be handled by fast integration tests. Now, there are perfectly valid other reasons to like static typing. I just always hear this one, especially in regards to Python, and it’s really the wrong way to look at it. The extra code and constraints of static typing are liabilities that should have to offer offsetting value to choose them. You already need integration and unit tests to reliably make changes and maintain the training code. If you can get the same benefit of overall job safety (or even 99% of the same benefit), from the tests, without paying the extra costs of static typing, then don’t! Turning it around to act like static typing is de facto always a benefit is a very one-sided way to look at it.
- bcyn 8y ago> Then the interview was nothing but deck of card puzzles and random riddles That's really disappointing. Was the position you applied for Software Developer, or a specific deep learning position?
- mlthoughts2018 8y agoSpecifically posted as a deep learning position.
- deepGem 8y agoThe type safety argument is total BS. First of all the training script will fail for the very first time if there is a type error. You'd be a moron to pass an argument of a different type 'a couple of hours' into the training. No sane programmer writes such code. What kind of nonsensical argument is this. What I have found static typing to be really useful for is in remembering what I have coded. It's quite hard to remember a dynamic type while you are writing code, given the number of variables you are dealing with. Seeing that type definition next to your variable name is a handy reference. I find it helpful to speed up coding a bit and being able to remember a lot more clearly what I have done.
- mlthoughts2018 8y agoStatic typing is also a nice way to communicate design intentions. But for this to work, the annotations have to be very expressive. I don’t know the first thing about OCaml, but I have worked professionally with Haskell and static typing is a joy when it adds clarity and makes the contracts of functions instantly readable. Contrast this with Scala, which I have also worked with professionally and the difference is stark. Scala type annotations are much harder to read, and the mechanism of implicits can make for extremely mysterious code that looks like it shouldn’t compile and only once you track down some distant implicit that’s somehow in scope, can you make sense of the way types are flowing through some function contract.
- deepGem 8y ago
- rememberlenny 8y agoFor reference, Jane Street is financial firm known for their widespread use of OCaml.
- remify 8y agoI'd add that the author Laurent Mazare is a fucking brilliant person.
- mi_lk 8y agoYou know him personally? Or he has some prior works that we can take a look?
- bachmeier 8y agoremify is Laurent Mazare.
- mi_lk 8y agoWhat a twist.
- remify 8y agoAhah no, I was just looking at his CV and the guy is jacked.
- 3rdAccount 8y agoHahaha! This is a comment I would normally make and be scoffed at. I'm always amazed at how smart some people are.
- 2_listerine_pls 8y agoor is it?
- gaius 8y agoType-safety helps you ensure that your training script is not going to fail after a couple hours because of some simple type error. This isn’t a failure mode that ever happens in DL... 2 hours into the job you will only be dealing with floats anyway no matter what language you are using. If you’re going to fail on anything typed it will be in the first 20 seconds probably, basically the instant you start your first epoch.
- habitue 8y agoThis is true, but it's also because Tensorflow is a typed language, but it uses the syntax of python. (At least the default way) in tensorflow, the graph is type checked on startup, and it'll fail if anything is wrong. Contrast this with pytorch, chainer, or tensorflow's dynamic computation graphs and they're much more likely to have a bug that happens later, since their graphs aren't verified up front. Unfortunately, typed languages won't help you much there. A big reason people use pytorch is because of its flexibility (i.e. they were bumping up against the constraints of a static graph system and wanted out)
- gaius 8y agoI admit I don’t have much experience with those, I’m a Keras and CNTK guy but the principles will be the same: marshal your data into a huge matrix of floats/one-hot and hand it off to training where it will spend 99.9% of it’s time. I am a fan of strong/static typing and was once very active in the OCaml community but that just struck me as a very odd thing for the OP to say... it’s just not something that people doing DL worry about. It could be valuable in the marshalling phase but that all happens before DL begins and (in my experience) in a separate program.
- phonebucket 8y agoThis is great. Functional languages have such an elegant representation of so many mathematical concepts. It's a bit of a shame that they don't have more widespread use in scientific computing.
- mlthoughts2018 8y agoI would also suggest looking into Keras and PyTorch too. I think they honestly achieve a greater degree of elegance and a greater degree of mapping the programming constructs into the mental model space of the domain expert, than any FP interface to neural nets that I’ve seen yet.
- phonebucket 8y agoI use PyTorch a lot; it's definitely my preferred framework at the moment. I just wish there was something as thoughtfully done and well-supported in a more functionally oriented language. Flux.jl on Julia is the frontrunner in this regard, IMO. The added benefit is that being written in Julia the whole way down makes it easy for practitioners to delve into the source code and extend it in a performant way without going into the C level nitty gritty.
- dnautics 8y agoFlux is great. Because it's Julia, I could write a custom datatype that has fewer bits and test to see if inference and training are possible, and then apply that datatype to ml models without writing custom kernels (except convnets, but I'm going to push code for that.)
- mlthoughts2018 8y agoThis is also very easy in Keras / Tensorflow using the FloatX parameter, or specifying e.g. float16 dtypes. However, I’d say desiring a framework that allows “easy” extensibility to choose float precisions lower than 16 bits and have it “just work” is actually a mistake. That type of flexibility is overkill. Instead, supporting a limited set of fixed types is better. To experiment with a new type requires some integration hurdle to make it recognized by the backend, and then requires published research or some similar type of evidence that there are use cases which materially benefit from that new additional fixed data type, to get a PR approved to add it. The reason is that permitting arbitrary complexity growth in the form of “easy” custom data type support has two big downsides, (a) the mechanism that makes it easy had to consume maintenance and development resources even if it’s a very obscure form of customization, and (b) more importantly, it proliferates and worsens the already insane problems of being able to export / import models from one language/framework to another. It’s a case study of KISS and YAGNI: this is super premature abstraction especially if it’s for experiments. And the hurdle of making a branch and adding your new dtype in the backend is not (and should not be seen as) a significant engineering hurdle. Rather it’s a very good check on complexity growth.
- xvilka 8y agoSorry for repeating myself, but since there is a machine learning and OCaml it worth mentioning Owl [1] - library for numeric and scientific computations, including ML. [1] https://github.com/owlbarn/owl https://github.com/owlbarn/owl
- yunfeng_lin 8y agoSo much bashing on static typing on deep learning:) Does any one from Google can explain the benefit since you guys are working on swift in tensorflow https://medium.com/tensorflow/introducing-swift-for-tensorflow-b75722c58df0 https://medium.com/tensorflow/introducing-swift-for-tensorfl...
- shoyer 8y agoStatic typing for catching errors is only a small part of the vision for Swift on TensorFlow. The real advantage of static typing is that it enables the compiler to reason to about your code, e.g., to automatically rewrite it for a hardware accelerator with guaranteed correct semantics: https://github.com/tensorflow/swift/blob/master/docs/DesignOverview.md https://github.com/tensorflow/swift/blob/master/docs/DesignO... This is obviously possible in Python as well (e.g., see Numba) but is clearly has additional challenges: https://github.com/tensorflow/swift/blob/master/docs/WhySwiftForTensorFlow.md https://github.com/tensorflow/swift/blob/master/docs/WhySwif... (I work at Google, but not on the TensorFlow team.)
- yunfeng_lin 8y agoThanks! that's a very interesting idea! Definitely worth exploring. Not sure it is my false sense. It seems that many python deep learning people are so proud of their chose, it is difficult to convince them.
- KenoFischer 8y agoStatic typing has very little to do with what the compiler can say about your code. You can have dynamic languages with very strong type systems and semantics as well as static languages with weak semantics. The only difference between static and dynamic languages is whether the compiler enforces completeness of the analysis or not.
- yunfeng_lin 8y agoI am not sure I agree with you. You do need compile time type to generate efficient hardware accelerated code. Python has strong type, but that is only available at run time, which is not useful to generate code. But now python also have optional type. this might be utilized in generating more efficient code though
- senorsmile 8y agoAm I the only one who gets confused by references to ML (ML derived typed FP vs Machine Learning)? The threads on this page are the represent a strange junction where I really have to think about what people mean, because they really could mean either!
- yaseer 8y agoThis happens to me too, having dabbled with ML for theorem proving. Thing is, ML is an obscure language for most people. The association with machine learning probably dominates in 95% of people.
- icc97 8y agoIt's becoming less obscure, F# (.net), Elm and Reason (JavaScript) are bringing ML to a wider audience. Plus Jane Street do a great job of promoting the use of OCaml.
- checkyoursudo 8y agoML with ML? I occasionally have to double check, yes.
- ummonk 8y agoI usually assume ML refers to machine learning and ML-like, ML-family, ML-derived, etc. refers to the FP languages.
- preparedzebra 8y agoI'm not convinced that functional programming will grow in terms of devs using it daily, but it has been very useful for myself in certain contexts (especially when I wrote math based libraries using permutations, heavy recursion, etc). The results of this seminar are awesome!
- mark_l_watson 8y agoVery nice. I have spent many evenings playing with the Haskell bindings for TensorFlow that don’t have the coverage these OCaml bindings have (e.g., character seq models). I have thought of learning some OCaml, maybe this will give me the kick in the butt to do it.