5 ms·
This is a very dangerous article as it tries to make an argument that a person with minimum to no training can start to run these black box systems. One of the
by CarbonCycles 4y ago
This is a very dangerous article as it tries to make an argument that a person with minimum to no training can start to run these black box systems.
One of the biggest fallacies that I have run into within this space is the failure to understand the problem statement. For example, this article does not mention or address models performances on very unbalanced datasets. Most classic M/S/D-L models will perform mediocrely if care is not given to understanding the statistical distributions underlying the data (or even understand the dynamics of the system such as seasonality).
In addition, the author does not address how biases are introduced by using AWS' (or insert Google, Azure, etc) algorithms...not all algorithms are codified equally.
Finally, this article demonstrates how ppl are trying to trivialize the complexity of very complex algorithms, statistics, etc with a plug-n-play...guess it doesn't help that many companies (I'm looking at you Meta) treat their DS as high caliber business analysts/intelligence units.
Okay, time to step away from the keyboard...
- whoevercares 4y agoIn the article, the report from Data Wrangler actually tells about class imbalance and report a whole other bunch of ML specific issues like target leakage. There’s a blog post on it (granted there are other pre canned tool you could use to do this) https://aws.amazon.com/blogs/machine-learning/accelerate-data-preparation-with-data-quality-and-insights-in-amazon-sagemaker-data-wrangler/ https://aws.amazon.com/blogs/machine-learning/accelerate-dat.... Canvas also seem to have nice histograms in their UI
- CarbonCycles 4y agoThere is a difference between reporting a stat and understanding the stat. Another thing that many ppl (including ppl who work in this field) fail to understand is the theoretical (or statistical) underpinnings of the algorithms they are trying to deploy. Many assume IID or NID to make the problem tractable, but that's not how the world typically works even on the very very large data scale. More things on why many data science teams/groups/fail because too many ppl treat it as a BI/BA organization...I really should step away from the keyboard now. LOL.
- time_to_smile 4y ago> class imbalance Worrying about class imbalance is a classic example of how ignorant many data scientists are about how the statistical properties of the models they're using work. For example if you are to try to "solve" class imbalance issues in a logistic model by over/under sampling you're actually throwing out important information about the prior probability of an event. Class imbalance in the data, if it reflects the real distribution of observed events, is often valuable information that will give you a better model. Yet so many data scientists I've known see class imbalance as a major issue and when you ask "why?" it's clear they don't know even the basic principles underlying the models they're using. There are times we you are forced to deal with a class imbalance issue due to the way this can impact training data size, but these cases are relatively rare in modern data environments.
- CarbonCycles 4y agoAnd this is where DS with a deeper more rigorous understanding begin to differentiate themselves by being able to step back and reformulate/re-model the problem into something more like an anomaly detector (in keeping with this example). I can see the argument where ML Engineers and DS w/out a more advanced statistics/STEM background would fail since they would continue down their list of libraries w/in their prescribed toolbox. Granted many problems can be approximated to be good enough, and let's face it, the FAANG/MAANG/whatever companies aren't running things so super critical where a user getting one extra email or ad presentation will cause them serious injury or death. Btw, appreciate your comments.
- matmatmatmat 4y agoCould you expand a bit about what kind of information class imbalance can reveal?
- timy2shoes 4y agoIt's not about revealing information, it's about how you model the problem. Say you're using logistic regression. If you upsample the minority class (never downsample the majority class because that's statistically inadmissible), then your classifier might get better at modeling the decision boundary. However, that comes at the cost of losing the best part of logistic regression: probability prediction and calibration. Sometimes the trade-off is worth it, sometimes not. Depends on the type of problem. For more details I suggest reading Frank Harrell: https://www.fharrell.com/post/classification/ https://www.fharrell.com/post/classification/
- mistrial9 4y ago> This is a very dangerous article as it tries to make an argument that a person with minimum to no training can start to run these black box systems. it is obvious that a class of executives-in-training want to make Ford Factory Workers out of machine learning operators. Do you want to be a Ford Factory Worker? maybe not but the implication is that there are others who will do it
- kjkjadksj 4y agoA ford factory worker stamps the same part for 60 years. Running ML is not the same as making a cog of a priori known quality over and over again, but lets let the MBAs run their little experiment and watch the whole thing explode so no one seriously wastes more effort on this. You cant no code your way around generating testable hypotheses which is why data scientists get paid the big bucks. Anyone can write up a model in python following some 30 minute tutorial. That was never the hard part.
- time_to_smile 4y ago> Do you want to be a Ford Factory Worker? maybe not but the implication is that there are others who will do it Unfortunately the vast majority of data scientists I've worked with are already more than half way there. Very few data scientists know how to actually model a problem. They only know how to take data, turn it into a matrix and run it through a grid search algorithm on a bucket of SKlearn models and parameters. In fact I'm sure there are data scientists reading this saying to themselves "wait, isn't that modeling?". The frustrating part of this is, in my experience, there are a ton of important modeling problems to be solved at most companies, and yet the huge number of data scientists hired out there don't even see modeling problems that don't fit the form y = f(x). A good quick sanity check of this is to ask a data scientist how they would model churn. An overwhelming majority will say "I'll take N days of data in the past, label who churned there, build a model and then predict who will churn in the next N days", which is absolutely the incorrect way to model churn. If you ask these people about censored data they'll just shrug their shoulders, completely ignorant of the fact that for decades people have been solving this exact same problem in the medical research community using survival analysis. In reality there is only a small subset of real business problems that can be quickly modeled in SKlearn, and yet this type of "modeling" is the large majority of what data scientists are doing today.