3 ms·
yes basically. i worked at a firm that used mattermark, datafox, etc. the amount of inaccurate, duplicate, and missing data was very high. all of these service
by maxpain 12y ago
yes basically.
i worked at a firm that used mattermark, datafox, etc. the amount of inaccurate, duplicate, and missing data was very high. all of these services are like a mashup of data sources that are not very good.
there's a reason bloomberg, thomson, etc. focus on public companies.
- birken 12y agoIt is comical how often sources like crunchbase (and even the press) are wrong. Though it makes sense, private companies basically have no incentive to release any sort of data unless it plays to some narrative they are trying to create. For example, the date the company was founded is often wrong in Crunchbase (it is reported as more recent than it actually was). Many companies raise money and don't disclose it for months (staying under the radar). Many acquisitions never publicly release the price (so the founders can claim it was a "success" regardless of whether it was). It is just a tough space to be in to try to make something of the data that is available. Public companies are of course much easier.
- zeeshanm 12y agoData accuracy is important but sometimes the bigger problems to solve are discovery, advanced filtering, and readability. I've played around with Mattermark a little bit and I think it does a fantastic job in solving the later problems. From a hacker's perspective you can do curl/wget, awk, etc to aggregate and filter data from various sources yourself. But to make the same tooling available to a wider audience is what gets you to the market. Dropbox is a solid example, for instance, that made rsync super simple to use for anyone.
- dmritard96 12y agoagreed, we did some number crunching against crunchbase to look for the right investor profiles, making that into a usable product is much more involved
- 7Figures2Commas 12y ago> Data accuracy is important but sometimes the bigger problems to solve are discovery, advanced filtering, and readability. Discovery, advanced filtering and visualization are increasingly solved problems. Open source solutions like Elasticsearch Kibana[1] make it incredibly easy to analyze and visualize large amounts of data, and to do it better than many paid services. For high-value commercial use cases, like those in financial services, data accuracy and completeness is all-important because your ability to identify the best opportunities and make good decisions is almost always proportional to your knowledge of the market. Take a seemingly simple venture capital use case: I want to identify pre-Series A fintech companies in California and New York founded in the past 3 years that have raised between $100,000 and $1 million and last raised funds 4-8 months ago. If funding data and corporate information is incomplete or inaccurate, and/or funding events have not been properly categorized, the list of companies surfaced will likely exclude companies that meet the criteria, and include companies that don't meet the criteria. For many use cases, it doesn't take many false positives or false negatives to render a data set effectively useless. [1] http://www.elasticsearch.org/overview/kibana http://www.elasticsearch.org/overview/kibana