3 ms·
When studying data science we had a major assignment that involved stock predictions. I came out on top of the class because I told my model to only bet when it
by roxgib 4y ago
When studying data science we had a major assignment that involved stock predictions. I came out on top of the class because I told my model to only bet when it was >90% sure of the outcome, discarding the result otherwise. The rest of the class didn't even think of that, and lost despite many of them have much better models than me when all the results were considered.
They were too focused on the technicalities of the problem too see the big picture, n issue I see a lot in data science students who often struggle to grapple with the problem they're supposed to be solving.
- largepeepee 4y agoDepending on results that give "90% sure" outcomes only work under specific constraints like if capital is low or there is near unlimited opportunities. Sounds like the rest of the class didn't read the fine print if that was the model that won.
- PaulHoule 4y agoThat's exactly what I'm pointing out. If you look at Kaggle answers or arXiv papers you might not even think calibration was a thing, but it's a key tool when going from academic ML to industrial ML. I got into calibration through text retrieval where the TREC methodology rates relevance functions by rank in such a way that you don't get points for being calibrated or well-calibrated. There are numerous things mainstream text retrieval systems don't do, most notably it is hard to build an alerting feature without calibration. Practically you have to set some relevance threshold to avoid getting too much irrelevant stuff. I calibrated a rather good search engine by fitting a curve to the score and found that the best it would ever give is p=0.7 to be relevant and that it gave that very rarely. So even with a calibrated score you wouldn't be able to set a very high threshold or expect to get many documents. Watson gets at this problem by having a large number of question answering models each which has a high p to be right when it is right but each of which also has very low recall, something you can do when using p as a universal score.