5 ms·
> I have mined and sold data to hedge funds Are you able to tell more about what types of data this was? I'd be interested in hearing more. I love the story of
by hellogoodbyeeee 9y ago
> I have mined and sold data to hedge funds
Are you able to tell more about what types of data this was? I'd be interested in hearing more. I love the story of hedge funds using satellite imagery of parking lots to predict retail store strength
- dsacco 9y ago> Are you able to tell more about what types of data this was? I mined data in the real estate, QSR, automotive and airline sectors (and a few peripherally related ones). We would identify a source of data that was a demonstrably strong proxy for a specific company's revenue (that is to say, if we broke out a naive timeseries of the data it would map nearly 1:1 to earnings results each quarter). Then we would collect this data en masse, writing software over a custom crawler and infrastructure, bespoke to each case. During ingestion we'd process and normalize the data and dump it in a database. Once we had a significant amount of data we'd build a forecasting model, which could be very straightforward ("how many products have they sold this quarter", for simpler companies) to very complex ("can we reverse engineer the amount of business which is now online" or "can we determine the undisclosed amount of business being done in this specific region"). This analysis (and not the data itself) constituted the product - we never even provided the data directly to the hedge funds. Once the crawling and analysis was mature, we'd incubate it for a few quarters while building interest; then, after showing a very strong mapping over several quarters, offer it as a data product. It was not uncommon to achieve <5% or even <1% margin of error for forecasting earnings announcements on a quarterly basis. The data was always public and legal, though challenging to identify and difficult to effectively crawl. It was typically derived from very uncommon sources, and would range in complexity from the extremely simple ("we can crawl sequentially incremented integers from this third party that appear to map to products sold by this company") to the very sophisticated and complex (e.g. "we've collected a zero sum distribution of ad partners and their spend on the network, are more partners churning off, and is this a better or worse outcome for the company?"). I've spoken about this in comments here before. This is still very doable, and I still do this sort of work for personal research (and profit). However, the work I do is now much more quantitative - I'm currently taking a similar approach for baskets of equities with the goal of forecasting macroeconomic trends (e.g. subprime auto loans) instead of the earnings results of individual equities. To give a very specific throwaway example (because it doesn't violate an NDA and it no longer works): it used to be possible to pretty accurately forecast large tech retailers' product sales each quarter (like Apple) by reverse engineering FedEx and UPS tracking numbers.
- defen 9y ago> it used to be possible to pretty accurately forecast large tech retailers' product sales each quarter (like Apple) by reverse engineering FedEx and UPS tracking numbers. That's great. Were tracking numbers vendor-specific in some way? So you could, let's say, order an iPhone once a week and get an idea for how many iPhones (or total Apple products) were sold in that time period, just by using the tracking numbers? Was it dependent on geography in some way?
- dsacco 9y agoYes, yes, no.
- WalterBright 9y agoReminds me of the British figuring out how many tanks the Germans made in WW2 by a sampling of the serial numbers.
- rmc 9y agoThat's genius! But then you have to buy a lot of iPhones, right? I suppose you can sell them on, and if you're getting ~$10k per quarter, it's OK. Or did you use a different approach? What changed BTW? Did the format of tracking numbers change?
- falsedan 9y ago> But then you have to buy a lot of iPhones, right? Just two per carrier would do: one at the start of the quarter, and one at the end. Earnings dates happen weeks after the quarter ends, so you'd have heaps of time to peddle your analysis around.
- opportune 9y agoWould you be able to describe how your data was normally priced? I'm sure it is very context dependent, but I'm very interested. For example, let's say you have information that allows you to forecast the earnings of a publicly traded company to within 1% MoE; for how much could you sell this information (either one time, or periodically) to a financial company? I once briefly chatted with a man who ran a company that used satellite imagery to predict crop yields. I found it amusing how rather than providing this information to farmers ("It looks like you're going to produce only 60% of what you did last year, time to cut some expenses!"), there was so much more money in providing this information to hedge funds and trading firms. Morality aside, the potential for technology like this is exciting. It's kind of cool how the combination of big data, scraping, statistics, etc. allows this small niche of analytics to thrive. I have a decent background in most of these subjects, and this seems like a very freelance-able job, so I'd love to dabble in this area and learn how it works.