4 ms·
Part of the problem is that if you gave 20 developers/data scientists/ml engineers the same the set of data and asked them the do data prep and feature engineer
by kmax12 7y ago
Part of the problem is that if you gave 20 developers/data scientists/ml engineers the same the set of data and asked them the do data prep and feature engineering, you'd probably have them come back with 20 different approaches.
To avoid pipeline jungles, teams need to agree to certain API's that their data processing code will follow e.g scikit-learn helped many people standardized around fit/predict/transform for their machine learning algorithms. In the future, I expect we'll see this expand to other parts of the process, such as feature engineering.
Towards that goal, I work on an open source library trying to do this for feature engineering called Featuretools. You can check it out here: https://github.com/FeatureLabs/featuretools/ https://github.com/FeatureLabs/featuretools/
- ende 7y agoAutoNormalize (part of FeatureTools, to those unfamiliar) is one of those most useful libraries I’ve used in awhile.