4 ms·
I would also agree that in general, there is certainly room to grow in dataset discovery. However, that will require a drastic improvement in the metadata assoc
by gervase 8y ago
I would also agree that in general, there is certainly room to grow in dataset discovery. However, that will require a drastic improvement in the metadata associated with many (most?) datasets, which I think is probably a large contributor to discovery difficulty.
Regarding your data, at an individual level, it will be very difficult to find that information publicly available, for the simple reason that the organizations that have that information (US Census, IRS) are under extremely strict privacy-preservation requirements whose goal is to prevent exactly these types of linkages.
You could instead try an analysis on aggregated data, i.e. correlations between (zip, average age, average household size) and (zip, average income) tuples. That data is available: [0, 1]
[0]: https://toolbox.google.com/datasetsearch/search?query=us%20income&docid=7wVMbp6OGT6LFlyWAAAAAA%3D%3D https://toolbox.google.com/datasetsearch/search?query=us%20i...
[1]: https://toolbox.google.com/datasetsearch/search?query=us%20census%20data&docid=LzZ6gQGw8DFk%2BliqAAAAAA%3D%3D https://toolbox.google.com/datasetsearch/search?query=us%20c...