4 ms·
Nice writeup. It seems like a supervised learning approach to fraud detection. I have a question: Where does the is_fraud variable come? Is it done by humans
by xtacy 12y ago
Nice writeup. It seems like a supervised learning approach to fraud detection. I have a question: Where does the is_fraud variable come? Is it done by humans?
- msherry 12y ago(I don't work for Airbnb, but work in a similar space) Yes, this variable is usually set after the fact. For instance, a given transaction may have led to a chargeback, or may be done by a known fraudster. These models are usually trained on historical data, so we can know with some certainty which transactions are fraud. It could be that a transaction was fraudulent but has not yet led to a chargeback (maybe the real cardholder hasn't yet seen their statement?), so there's still some uncertainty, but hopefully that approaches a minimum after some time passes.
- segmondy 12y agowhat software are you using?
- msherry 12y agoI don't want to hijack this thread too much from the original post, but we use some of the same software as Airbnb (scikit-learn, randomforest models, etc.) as well as some stuff developed in-house. Credit card fraud has been one of our biggest issues, and we've developed some pretty robust systems to fight it. Contact me privately and I'd be happy to talk about it -- this stuff is what I do for a living.
- zephyrnh 12y agoThanks! It varies depending on the model, but in most cases, yes, the confirmation of fraudulent activity is manual since having the ground truth be correct is critical
- xtacy 12y agoOK good to know. :) Would you be writing another blog post on what features you use for fraudulent activity, or do you consider this a business secret?
- zephyrnh 12y agoNo, features are definitely a business secret. My goal was to share as much as possible that might be helpful to others combating fraud, without giving anything away that would hinder the effectiveness of our systems.