3 ms·
I agree. The only thing I believe the article gets right however is that the naming doesn't expose this default behavior, which might cause problems for people
by Macuyiko 7y ago
I agree. The only thing I believe the article gets right however is that the naming doesn't expose this default behavior, which might cause problems for people coming from other software.
Still, any good data scientist I know actually knows to read the documentation carefully and has learnt regularisation is the default.
- missosoup 7y agoThe call signature does expose it though. Am I a minority in terms of using jupyter notebook and an IDE that shows call signatures and docstrings?
- paulgb 7y agoOne thing that concerns me about the "you should read the documentation" argument is that even if the person writing the code did that for every function they used, anyone reading the code is likely to make (reasonable) assumptions. So bugs like accidentally regularizing, for example, could slip past code review.
- deleted 7y ago[deleted]
- ad404b8a372f2b9 7y agoIt's a fair point for standard software engineering libraries, that the defaults should be obvious, but I don't think that it's possible to hold scientific libraries to the same standards. They require expert knowledge, and the assumptions you make about the models that are being used should be very carefully checked. There is no obvious defaults in that case because we're not writing software, we're building scientific models. For example, look at the neural network classifier model in scikit-learn: https://scikit-learn.org/stable/modules/generated/sklearn.neural_network.MLPClassifier.html https://scikit-learn.org/stable/modules/generated/sklearn.ne... None of these defaults can be assumed. If you showed me a function call with no optional arguments I couldn't tell you, nor anybody else that hasn't read the documentation, what activation is being used, what optimizer, learning rate, etc... In fact when coming across these kind of models during code reviews the nature of these parameters is the first thing I ask about.
- rfeather 7y agoAgree. It's naive to assume your use case is the basic one when using these libraries or that you know the underlying implementation and all of the parameters, since in ML the implementations vary enough to affect the outcome and will have different controls. To the point of exploratory analysis in some of the parents ,I prefer statsmodels for that purpose. It's not quite up to where the similarly purposed tools are in other languages, but for most of my work where I care about interpretation, it hits the right spot between usability and providing the standard statistical outputs.
- zenexer 7y agoThen there shouldn’t be any defaults, right? All of these issues could easily be avoided by forcing users to provide explicit parameters. It only makes sense to offer defaults if they’re intuitive.