5 ms·
Generate Synthetic Data in 3 Lines of Code
- andrewnc 4y agoWe've been working hard to make it super simple to get started with useful synthetic data. If you want to know how you would go and use this for your own problems check out some of our other posts https://gretel.ai/blog/how-to-safely-work-with-another-companys-data https://gretel.ai/blog/how-to-safely-work-with-another-compa...
- niviksha 4y agoMy use case is generating a very high rate (10k e/s up to 100k e/s) of JSON-NL events from samples of JSON-encoded log data (JSON-NL to be exact). Is this supported in OSS Gretel? FYI, I'd built a hand-crafted generator using JSONNet templates and Golang, but I really wanted something that could model source data distributions accurately. The use case is large-scale load testing of customer workloads without requiring actual data.
- andrewnc 4y agoWe're currently beta testing something that fits this use case directly. The models we have today are really great at capturing the original distribution, but they're not always the fastest. This new stuff will change that, feel free to reach out (maybe on our slack?) and we can see if we can get something working
- andrewnc 4y agoBlog post is out now https://gretel.ai/blog/introducing-gretel-amplify https://gretel.ai/blog/introducing-gretel-amplify They get 43,300 records per second on this example, which seems to the right order of magnitude for you
- jmole 4y ago...and you too can learn the biases and weights matrix of our synthetic data generator!
- BobbyJo 4y agoModels are trained on input data to generate synthetics data similar to the input. It's not so much 'a' synthetic data generator. It's more like a platform for creating your own synthetic data generators using your own data and ML.
- johnwatson11218 4y agoDoes anyone know if there are deep learning libraries that can model the relations between table based data? I see many that work on a table or a data frame but where I work our db is over 1000 tables and nobody can understand it. I feel like the next frontier is a tool that you point to your oracle or sql server and it compresses the table space. Whether you consider it a kind of PCA dimensionality reduction or the logical extension of the db "normal" forms ... it is just compression.
- rockemsockem 4y agoI believe that random forests are still the primary way data scientists work with tabular data. Deep learning hasn't cracked tabular data like it has other areas.
- johnwatson11218 4y agoEverytime I look into this stuff it is just one table/dataframe. Nobody is modelling the relations between things. I think there is a huge opportunity for a product that can look at Customer, Items, Orders, Returns, Payment_Methods etc. and first of all show me when things tend to co-occur, so that it could generate synthetic customers that have the statistically correct number of registered payment methods where those payment details are also synthetically generated. Next would be the ability to decompose or factor my entire db into the subcomponents and make suggestions for combining tables. The use case would be a legacy enterprise system that has grown so complex and tangled that devs are afraid to do basic db refactoring. From where I'm sitting and working this is the next gold rush, apply DL methods to the bread and butter computing, log flow analysis, etc.
- iansane 4y agohi john, your first paragraph is my phd (2021). its a massively under researched area because relational data has more... dimensions to it and is understandably not as exciting to most. happy to discuss (email in profile)
- andrewnc 4y ago
- niviksha 4y agoWow, this is great. I built my own synthetic time series data generator for benchmarking, could have saved myself a bunch of trouble with this.
- agolshan 4y agoI was just looking for something that would allow me to create my own synthetic time series data generator for benchmarking. This is a fantastic solution! Excited to try it further!
- Raziarazzi 4y agoRandom forests, in my opinion, are still the most common way for data scientists to work with tabular data. Deep learning has yet to crack tabular data as it has in other areas.