3 ms·
This appears to be an thorough overview of machine learning. Even though many bases are covered, I wish there was more on how to create or select features for m
by kmax12 9y ago
This appears to be an thorough overview of machine learning. Even though many bases are covered, I wish there was more on how to create or select features for machine learning in a systematic way. Feature extraction gets 1 paragraph, but feature engineering and selection aren't mentioned much in the 200+ pages!
I don't have a Phd in machine learning, but I have spent many years using it as a tool to solve problems. While the details here can get you a long way, without understand feature engineering or feature selection, you will have a hard time building accurate models.
For any engineers looking for more on feature engineering after reading this, I maintain an open source library for automated feature engineering called Featuretools (https://github.com/featuretools/featuretools https://github.com/featuretools/featuretools). We also have demos on our website (https://www.featuretools.com/demos https://www.featuretools.com/demos) if you want to see it in action.
- bistro17 9y agoI was wondering how featuretools differs from https://github.com/AxeldeRomblay/MLBox https://github.com/AxeldeRomblay/MLBox https://github.com/crawles/automl_service https://github.com/crawles/automl_service and the proprietary and newly launched driverless ai (from h2o)
- kmax12 9y agoFeaturetools focuses on handling data with relational structure and timestamps. Here's an example to explain those two key points. Imagine you have a relational database from a retail store with tables for customers, transactions, products, and stores. Featuretools can make a feature matrix for any entity in the database using an algorithm called Deep Feature Synthesis. We wrote a blog post about it here: https://www.featurelabs.com/blog/deep-feature-synthesis/ https://www.featurelabs.com/blog/deep-feature-synthesis/. Basically, it tries to stack dataset-agnostic "feature primitives" to construct features similar to what human data scientists would create. This means that a data scientist can go from building models about their customers to models about their stores in one line of code. One aspect worth highlighting is that Featuretools can be extended with custom primitives to expand the set of features in can produce. As the repo of primitives grows, everyone in the community benefits because primitives aren't tied to a specific dataset or use case. Some of our demos highlight this functionality to increase scores on the Kaggle leaderboards. Featuretools is good at handling time. When performing feature engineering on completely raw data it is important not to mix up time. When your data is timestamped, you can tell Featuretools to create features at any point in time and it automatically slices the data for you (even across relationships between tables!). You want to avoid situations similar to training a machine learning model on stock market data from 2017, testing that it works on data from 2016, and then deploying it and expecting to make money in 2018. You can read more about how featuretools handle time here: https://docs.featuretools.com/automated_feature_engineering/handling_time.html https://docs.featuretools.com/automated_feature_engineering/...)
- swsieber 9y agoIs there anything for supervised document classification? Like raw plain text -> features?