5 ms·
Does anyone know if there are deep learning libraries that can model the relations between table based data? I see many that work on a table or a data frame but
by johnwatson11218 4y ago
Does anyone know if there are deep learning libraries that can model the relations between table based data? I see many that work on a table or a data frame but where I work our db is over 1000 tables and nobody can understand it.
I feel like the next frontier is a tool that you point to your oracle or sql server and it compresses the table space. Whether you consider it a kind of PCA dimensionality reduction or the logical extension of the db "normal" forms ... it is just compression.
- rockemsockem 4y agoI believe that random forests are still the primary way data scientists work with tabular data. Deep learning hasn't cracked tabular data like it has other areas.
- johnwatson11218 4y agoEverytime I look into this stuff it is just one table/dataframe. Nobody is modelling the relations between things. I think there is a huge opportunity for a product that can look at Customer, Items, Orders, Returns, Payment_Methods etc. and first of all show me when things tend to co-occur, so that it could generate synthetic customers that have the statistically correct number of registered payment methods where those payment details are also synthetically generated. Next would be the ability to decompose or factor my entire db into the subcomponents and make suggestions for combining tables. The use case would be a legacy enterprise system that has grown so complex and tangled that devs are afraid to do basic db refactoring. From where I'm sitting and working this is the next gold rush, apply DL methods to the bread and butter computing, log flow analysis, etc.
- iansane 4y agohi john, your first paragraph is my phd (2021). its a massively under researched area because relational data has more... dimensions to it and is understandably not as exciting to most. happy to discuss (email in profile)
- andrewnc 4y agoWe have some preliminary work in this direction https://github.com/gretelai/multi-table https://github.com/gretelai/multi-table I love the idea of "table space" though. It would be fun to traverse this space and output a new database at each step, like a VAE.
- cush 4y agoThere's a company called Tonic.ai that reminded me of your comment - they deeply analyze a db, including the data, and are able to find and strip out all of the PII, while keeping the various relationships in tact. Super interesting podcast on it here: https://softwareengineeringdaily.com/2021/09/29/faking-data-using-tonic-ai-with-ian-coe-and-adam-kamor/ https://softwareengineeringdaily.com/2021/09/29/faking-data-... They mention in the podcast that most customers end up finding relationships in their tables that they didn't know they had - that weren't explicitly in schema
- chiarabeth 4y agoThanks for the shoutout, cush! And yup, our platform Tonic enables developers to realistically de-identify their data while preserving relationships and consistency across tables within their DBs, to optimize dev and test with real fake data. You can sign up for a sandbox here: https://www.tonic.ai/ https://www.tonic.ai/ We've also recently released a new platform called Djinn that is specifically designed for data science workflows. It enables you to query from tables across your DB to build customized views of only the data you need and synthesize high-fidelity data based on models trained on those views. Relationships are fully preserved and no external scripting is required. You can create an account and take it for a spin here: https://djinn.tonic.ai/?signup https://djinn.tonic.ai/?signup Full disclosure, I'm Chiara Colombi, Product Marketing Manager at Tonic.ai. Cheers!