5 ms·
The scikit-learn cargo cults
- gleenn 5y agoThe author's beef seems to be "people use similar terminology across similar libraries/frameworks/platforms but they don't behave identically and represent subtly different things". Maybe I don't do enough data scienceing, but isn't this super common? Like, if I write a parser... I'd probably call the main function "parse", or if I'm writing a database connector, I'd probably call the function "connect" to do the connecting. I personally wouldn't expect those to work identically or mean the same exact abstraction. I personally love when things are named similarly so I can grok the meaning in a new codebase more quickly, even if things don't transfer identically.
- Jugurtha 5y ago>I personally love when things are named similarly so I can grok the meaning in a new codebase more quickly, even if things don't transfer identically. I think it is a sign of good design when things are named similarly, as it makes it easy to use. I'm not surprised it acts differently than something else, because that surprise implies an expectation that shouldn't be there, even more so when there is absolutely no standard both need to comply with, which is the case in the current ML/DS context. There are surprises when people have different implementations of a spec or whitepaper and people accept even those. Not saying it's a good thing. But having that expectation for standardless things just because the entities names are similar is too much to ask for. It would be great if they did for interchangeability, but it's okay if they don't. More work (bridges, adapters, and what not), but different people took a shot at something. The field is not mature enough but going in that direction with attempts at formats, interfaces, protocols either for data (protobuf, dataframes, events) or models (PMML, PFA, ONNX, etc).
- huac 5y agoSKL being first does not afford it a monopoly on ML object design. Nor should other libraries necessarily seek to emulate what came first (or support pickling...)
- rubatuga 5y agoThe author doesn’t really know what cargo cult is. It means doing things similar to other groups and expecting an unrealistically positive result. Not only do you have to prove that other ML libraries were imitating sklearn, but that copying it wasn’t useful. Like another commenter said, naming the functions: “fit” and “predict” are simply common names to easily convey meaning. It certainly has the positive effect of letting me know what the functions do. If that’s cargo culting, then so is any program that has a “main” or “init” function with different arguments. Also, to refute their last point, PyTorch is too low level to have a fit function, not because they aren’t trying to cargo cult.
- deleted 5y ago[deleted]
- deleted 5y ago[deleted]
- deleted 5y ago[deleted]
- deleted 5y ago[deleted]
- jwilber 5y ago“ Sagemaker “Estimators” do not have anything to do with fitting or predicting anything. The SDK is not supplying you with any machine learning code here.” The author is confusing the sagemaker service with the mxnet deep learning library (which sagemaker provides access to). Basically everything they wrote in that section is flat out incorrect.
- maldeh 5y agoYeah, given that the article started off establishing how an Estimator was basically an interface with simple rules about supporting "fit" and "predict" and how it could contain anything or do anything, I thought the argument laid out here would be about how these derivative implementations broke these rules. The rest of the article instead seems to have lost the plot though, somehow finding fault with various derivative or concrete implementations of this interface, for A) being inextensible implementations and not transitive interfaces themselves, as though "be anything do anything" no longer applied; or B) not being perfectly aligned with sklearn estimator details that the author didn't really identify as essential, like not following some sklearn-specific parameter naming rule or not being serializable via pickle (like seriously, pickle support is often not appropriate for production, why should this be a required pattern! It's not even a requirement of the interface unless you read between the lines like the author implies is essential to be at parity.) As other commenters outlined here, it assumes that sklearn's contract is absolute, as though other libraries couldn't reinterpret the core principles. The arguments against Tensorflow or Sagemaker's interfaces especially stretch quite a bit - what exactly is so offensive about these implementations given the very rules that the author establishes in this article? All "fit" is supposed to do is update internal state as the author asserts, but what precludes implementations of this interface from using cloud-based compute resources to achieve this end? And what about the fact that a docker container is deployed to the cloud by this command makes "fit" a lie? And honestly, what does the author have in mind for an estimator implementation that uses cloud resources like GCP TPUs or AWS EC2 that is also somehow more correct or pure than these implementations? More than anything, the author's dismissal of the value that GCP and AWS's implementations bring in eliminating infrastructure management via their Estimator implementations (equating it to "simply" writing Dockerfiles or running Docker containers on the cloud like there's no setup involved) implies that they're thoroughly disconnected from the realities of ML devops on the cloud. They're free to run their purist single-core sklearn estimators on their laptops as much as they'd like though (unless Dask somehow gets a pass from these arbitrary rules around how estimators can and cannot be used).
- gyrovagueGeist 5y agoHuh, didn’t think I’d see the writer of The Northern Caves on the top of HN. Back to this post: I’ve written some nearest neighbor code and definitely felt some pressure to make the API sklearn compatible. But I don’t think it’s as bad as the post claims in practice. Highly recommend checking out the posters other work. Its a lot of fun,
- nightpool 5y agoTo be clear, nostalgebraist has been doing ML professionally for years, so he's working with these APIs all day, every day. If he has a complaint about them, I expect it to be well-grounded in months-to-years of full-time experience, not just idle speculation that doesn't pan out in practice