5 ms·
How to build your own feature store for ML
- LexSiga 6y agoAs this topic will inevitably become more trendy find some some additional interesting resources on the subject as well: - https://www.quora.com/What-are-the-implementation-challenges-of-a-machine-learning-feature-store https://www.quora.com/What-are-the-implementation-challenges... - http://featurestore.org/ http://featurestore.org/ (a list of -some of- the available feature stores)
- jamesblonde 6y agoI'm the author. Let me know if you have any questions.
- kuu 6y agowhat is exactly a "feature store"? do you store "features" by themselves? do you store whole data? Can you give me a bit of insight on this?
- SirOibaf 6y agoYou can see the Hopsworks feature store as a repository of curated features ready to be used in ML models. Or, a middle layer between data engineers and data scientists: - Data engineers write the data pipelines with the transformations and publish the features on the feature store. - Data scientists browse the feature store, pick which features they need and build the model In the Hopsworks Feature Store we group features together in feature groups. Feature groups can be then joined to create training datasets. (You can also select a subset of features from a feature group) Training datasets are stored in a ML Framework friendly format (e.g. TFRecords if you are using TensorFlow) and you can feed them directly to your model. If you are interested, we have a longer blog post explaining the core concepts of the Hopsworks feature store: https://www.logicalclocks.com/blog/feature-store-the-missing-data-layer-in-ml-pipelines https://www.logicalclocks.com/blog/feature-store-the-missing...
- overfitted 6y agoThanks! Also the first 10-15 minutes in this recording explains the Feature Store concept and how it can integrate with other ML tools (in this case Sagemaker) incl. slides, examples and demo if you are more of a watch & listen type. https://www.youtube.com/watch?v=3DaTA7o0FHY&list=PLgN6fhzkSui-NWA7_tPFhBACUaIzFjj9C&index=5 https://www.youtube.com/watch?v=3DaTA7o0FHY&list=PLgN6fhzkSu...
- dijksterhuis 6y agoIs this sort of lambda architecture for ML applications? Not referring to the "library" part of the flowchart obviously. Only major difference I can spot is the streaming layer is more about read speed latency instead of serving real time data.
- jamesblonde 6y agoThere is a similarity between feature engineering pipelines that feed the feature store and the lambda architecture, in that you have two sinks for your data - one for batch applications (and creating train/test data) and one for real-time serving of features. However, there is typically only one feature engineering pipeline (to make sure features are consistent between training and serving), whereas lambda has two (one batch, one streaming).So, you could come back and say it is more like the kappa architecture, but it could be either a batch or streaming application computing the features and saving them to the feature store.
- king_magic 6y agoPractitioner; not really convinced this is something i or my team needs. Maybe I just need a really dumbed down explanation of what a “feature store-as-a-Service” is. If we were talking about a super flexible/easy to use data catalog-as-a-Service that made it dead simple to store, version, manage & pull data from datasets, then I’d be super interested. But a “feature store” by itself? I just don’t get it - what am I missing?
- jamesblonde 6y agoI think for small teams with a small number of models, a feature store is probably overkill. Just like you wouldn't need a data warehouse if you only have one database. However, with lots of sources of features (oltp database(s), data lake, kafka), it becomes very hard for data scientists to find/use the data they need to train models. The feature store acts like a data warehouse for features for data scientists - and the more features are reused within the organization the more value you will derive from the feature store. You wouldn't ask a Tableau user to go to S3 or Athena or Big Query to get their data for reporting, and at organizations with feature stores, their data scientists can find most of the features they need in the feature store (Uber have >20k features last i heard in their feature store). Then, there are problems related to making those features available to online applications and not duplicating feature engineering pipelines and monitoring of models that are covered elsewhere on this page and in the literature (http://www.featurstore.org http://www.featurstore.org).
- michaeltlee 6y agoJust a heads up: There is a typo in the flow chart — “logicaclocks.com”.
- LexSiga 6y agoSome sad soul just lost their job. Also thanks; I fixed it.
- strgcmc 6y agoJust to add a counter-anecdote, as I see lots of (good/valid) questions about "why do I need this?", here's an anecdote about "yes we definitely benefited from this": - Years and years ago, we already had a data warehouse (DWH) - In the data warehouse, you would store data like, each and every order that all customers have made (i.e. up to and including full-fledged facts and dimensions about each) - Now, let's say axiomatically/hypothetically, a very useful and highly predictive feature for ML is "# of orders made in past 7 days" for each customer - Can this be computed from the data already in the DWH? Yes, absolutely, but it's a new computation and not an existing column/attribute in the dimensional model. - What if you need to recompute this feature daily, for millions of customers and orders? Well, we could always just add it to the dimensional model, compute it once, and let people just use it/share it... but why? Most internal users of the DWH probably don't care about something like "# of orders past 7 days" as something to be added to a customer dimension or per-customer grain (too specific or whatever), and moreover, the DS/ML folks want the same feature but for "every 1/3/7/30/90/180/365/730/etc." days breakdown, as well as a bunch of variations about orders and things other than orders (e.g. "average time between new orders, over past 7/30/90 days" or "average $ spent over past 7/30/90 days", as features that serves as a proxy for frequency of activity and level of engagement) - Hence, it makes sense to keep the "golden copy" of data in a canonical form in the classic/standard DWH as a baseline, and to separately/independently compute features out of that data and to store them in a different system (which can also be optimized for the different query/access patterns that DS/ML have, vs traditional BI). Over time, it also made sense in certain cases to go upstream of the DWH to source data from and process it more directly (for performance/efficiency reasons), though generally deriving features out of the transformed dimensional models was still very useful. - It took our teams ~1-2 years to really go through this evolution and reach a mature-ish state, but for the past 2-3 years, we've benefited tremendously from having an independent feature repository/store, that is separate from the classic DWH. Benefits came in all the obvious and some non-obvious ways, i.e. in faster iteration/cycle time, in better quality/repeatability, and in being able to automatically discover interesting relationships that no human could have anticipated - simply by having a very broad/large repository of features and running automatic feature selection over it.
- claytonjy 6y ago
- encyclopedia 6y agoSimilar to generating feature vectors for dataset augmentation here https://vectorspace.ai/covid19.html https://vectorspace.ai/covid19.html
- deleted 6y ago[deleted]
- StonyRhetoric 6y agoThis is a good idea - every ML operation should have something like this, to store, organize, version data, check for drift, do time-travel, backups/replication et cetera. But to borrow from Steve Jobs, I think this is a feature, not a product. If you've already done the hard work of setting up a data lake or data warehouse in a cloud provider, the cloud provider can give you backups and replication, and even some time-travel. Using something like Delta Lake or even just the standard Kimball DW audit columns will get point-in-time queries. Feature versioning is just query versioning in source control, and if you have schema, you can schema version with views if you need to. If you don't have a data lake, data warehouse ... well, you'll still need to gather and clean all your data before you put it into a feature store, and that's where 90% of the work is. I'd love to learn more, I'm sure I'm missing something, but it seems that they're re-solving the solved part - data storage and versioning. Checking for drift and data integrity is a nice bonus, but again, lots of libraries for that. I guess I could see it being beneficial for ML shops that don't have modern development practices, but if you don't have that, you have bigger problems anyways.
- jamesblonde 6y agoAll your points are valid points. However, operational models (models used by online applications, for example) typically need access to lots of historical features that are not available in the application. In that case, you need to go to a low-latency database/store to get your feature values (build your feature vectors). If you want to reuse those features in different models, you will need join support for building the feature vectors, so a key-value DB won't help there. Now, your features are duplicated between this online/serving layer and the data warehouse. How do you sync them up? The other thing you're missing is time-travel queries (temporal logic for SQL in data warehouse speak). Yes, Delta Lake gives you this, but you will need to wrap that data in APIs so that your data scientists will be able to use it. For data drift, a library alone won't cut it. You need to compare descriptive statistics/distributions of the data used to train the model and the live data coming in. Where do you get those statistics from - the feature store, in our case (with the help of versioning+metadata). Then, there is end-to-end governance of ML models - what training dataset was used to train this model, can i reproduce that training dataset if it hasn't been archived? You need metadata to manage all that. So, yes you can do it - but you have to build something (as the article describes) or buy it.
- tristanz 6y agoA great collection of real-world case studies and various implementations can be found here: http://featurestore.org/ http://featurestore.org/
- kadder 6y agoThis View point is an engineering driven view point to solve the featurization problem, the same problem can be approached in a much more simplified way which is more compatible with how practitioners work. Have a look at some of the abstractions sagemaker has in place for A modeling workflow, abstracting featurizarion at that level is more beneficial than approaching the problem this via a core engineering driven mechanism Based upon my experience on building a system like this, the percentage of Features reuses and searched across models is Generally lot smaller. A system which provides A publishing And a fast simple serving mechanism generally meets all the needs. Metadata management, history, audit, lineage, search are all good to have but not critical requirements to most practitioners Eg computing an aggregate lookup over a time range in spark, archive in S3, wrap it in a fast lookup implementation in a sickit learn transformer, and have the transformer pickle the lookup will give you offline , online parity out of box