3 ms·
I've put a LOT of thought into address cleaning! And yep - levenstein distance seems to be the way to go. My current stack is: 1. Send addresses to https://sm
by bpchaps 8y ago
I've put a LOT of thought into address cleaning! And yep - levenstein distance seems to be the way to go.
My current stack is:
1. Send addresses to https://smartystreets.com/ https://smartystreets.com/ - They gave me a year's worth of unlimited geocoding for free. They also tokenize the addresses, but I had about a 50% success rate with them.
2. Tokenization raw addresses with https://github.com/datamade/usaddress https://github.com/datamade/usaddress.
3. Use a normalized levenstein distance algo to get ratio of difference.
4. Compare all of the addresses' levenstein distances with each other.
5. Apply logistical regression/gradient ascent algo to tickets by chaining heavilytypo'd addresses to less-typo'd and eventually to a static list of verified-correct addresses.
It works surprisingly well, but there are still a lot of problems that can't easily be solved:
1. Street types (st/ave/blvd/etc) are missing. So, when two addresses have the same street name, it's difficult to pair the two. It's still possible with some probability stuffs and matching the ticketers' paths to the nearest street.
2. Addresses have a LOT of one-off situations. For example, there's a street name called "Avenue A". The street name here is "Avenue", and the street type (usually st/ave/etc) is "A".
3. Lots of four letter streets make levenstein distance very difficult.
Glad you enjoyed it!
- kioleanu 8y agoI did enjoy it, yes, and I'm following your idea for my town also (it's open data here). Lucky for me, it's a little bit prettier (I think they have autocomplete on their devices for the addresses). I already have some preliminary data - in a city with 350k inhabitants, they gave 150k fines last year, totaling 2.5 mil EUR. I can't wait to search for the hotspots
- bpchaps 8y agoLet me know how it goes!