3 ms·
Has anyone ever done a project like this that builds test databases from larger production databases? I mean something that can look at a prod db and model the
by johnwatson11218 13y ago
Has anyone ever done a project like this that builds test databases from larger production databases? I mean something that can look at a prod db and model the data in the columns and the relationships between tables and produce a much smaller test db that has the same statistical properties?
For numeric columns you could just fit a statistical distribution and sample from that.
For names I'm thinking you could look at the frequency charts for first, second, third ... letters and sample according to that. I believe it is also called Markov Chains.
In Andrew Ng's Machine Learning class he talked about taking labeled images and expanding the set by inverting, shearing, flipping, distorting them etc. He called the technique 'data synthesis'.
The test data problem has been hampering my team's ability to create maintainable automated tests.
- femto113 13y agoI've been working on such a project for a couple years as a solo/on-the-side thing. It's not available as a turnkey product yet, but the technology is working and can build arbitrary amounts of realistic test data (based either on your private data sets, or public sources like census data). I'd love to begin working with others on how exactly to integrate this into their test and development workflows. Email me at ken.woodruff@gmail.com if you'd like discuss further.