13 ms·
If you're dealing with international addresses, point 3 becomes very challenging. The 'tokens' of an address are called different things everywhere, take differ
by tomlagier 7y ago
If you're dealing with international addresses, point 3 becomes very challenging. The 'tokens' of an address are called different things everywhere, take different forms, and sometimes don't make much sense to compare. Figuring out a good balance of usability and generality can be really tricky.
Here's nearly 100 pages on the subject (from the perspective of addressing mail for USPS): http://www.columbia.edu/~fdc/postal/ http://www.columbia.edu/~fdc/postal/
- crgwbr 7y agoI’ve been through this before: storing a couple addresses for a few hundred different countries in an otherwise very normalized RDBMS. We couldn’t figure out anyway to do it other than EAV tables and a defined attribute pattern for each country. Even though that project is long in the past for me, I’m really curious how other people do this. Anyone come up with something better?
- pbowyer 7y ago> Even though that project is long in the past for me, I’m really curious how other people do this. Anyone come up with something better? I've always gone the other way. Bland fields (Lines 1-4, city, state, postcode, country) or - when I didn't need to validate it, a textarea for free form entry. Most of the time validation has been required by the shipping provider, so freeform hasn't been possible. As the shipping providers haven't got anything 'better', we stick to their format and field lengths (which are terrible for out-of-country addresses)
- Benjammer 7y agoAt that point you might as well just make an "Addresses" or "Address" table with a raw text field for the address. At the end of the day this thing is something that a human will read and map to the real world location.
- crgwbr 7y agoWe went the EAV route instead of just a free form text field because we still wanted to expose discrete fields to in the UI and have some front-end validation. But sure, from a data integrity standpoint it’s no better than a text field.
- uk_programmer 7y ago> At the end of the day this thing is something that a human will read and map to the real world location. That is quite an assumption. I personally would use a field with a structured format e.g. XML (most RDBMS systems can deal with XML easily) and different types of address field e.g. postcode, zip code etc. Then you could use a strategy for each country when dealing with it in your application.
- tomlagier 7y agoAt a past job, we basically did "both" - we'd store our best guess at a normalized address, and we'd store the text representation. For stuff like shipping labels, we'd use the text. For analytics, we'd lean on our best-guess normalization (and understand that there's some potentially significant error). You can also lean on various APIs to normalize, but nothing that I'm aware of does this well on a global scale.
- Benjammer 7y agoI think programmers over complicate things too much when it comes to addresses. When you're dealing with a global system that is able to deliver a package to "the 3rd house on the right, past the pond, with the red fence," or whatever, in rural anywhere, you probably should just treat it like we all do in the real world, as just a string to be interpreted in context.
- _jal 7y agoThis fails as soon as you talk to an API that requires normalization, need to aggregate your data, or, really, do anything other than try to deliver it.
- Benjammer 7y agoLmao, you're really going to hate me when I tell you I don't think we should be using addresses for anything other than delivering physical things. I know it's a cop out answer. But I don't have a better one.
- aerbherhber 7y agoSome counter-examples include Route Planning or determining where to locate a Consolidating Freight Service. If your addresses are not normalized, you lose the ability to see two addresses listed as "from city center, three lefts" and "from city center, one right" take you to neighboring locations.
- waste_monk 7y agoWhy not just store GPS coordinates and do geospatial queries? which gives you nice features like "show me stuff within X kilometres of a shop" and "is cutomer X within bounding box Y?".
- mgkimsal 7y agoyou could have 500 people all at the same lat/lon, with different floors, no?
- CameronNemo 7y agoI think city, postal code, state/province (optional), and country would be fairly standard. At least in North America, Japan, Australia, and the UK. Of course for country you would want to use standardized country codes, which are somewhat political.
- clon 7y agoThanks, this is a fascinating resource! My logic with the address field is to try to consider what is the purpose of the address in the application. If it is used for someone to send you a letter in the mail (for example: university admission) then you are better off with a big box where the user can write the magic blob of text that usually gets their mail delivered to the right place, up to instructions to the local postman. If you need to pull off some analytics, or need to estimate shipping, then just extract the fields that your shipping agent demands: country, zip code etc. Leave all else intact in the "magic blob".
- pelliphant 7y agoi don't mind web-pages (or whatever we are talking about) having lots of fields for address, as long as i don't have to type something in all of them. But it's super annoying when I have to make up things for fields that doesn't make any sense where I live.
- oftenwrong 7y agoIf you are mailing something, using a plain formatted address string is usually fine. Here's a harder problem from a previous job: Given two addresses, tell me if they both point to the same "place". (As you might imagine, the goal of this was joining various datasets.) There really is no good way to do this, but to come close, you might consider a few dimensions of the problem: 1. Formatted address - a plain string. What you would write on an envelope. 2. Tokens (what I call "address components") - the bits of information that make up an address. Something like, {"street_number": "10", "street": "downing st", "locality": "london", ...}. You can get these back from most location-focused APIs. 3. Locations. Coordinates of a point or area in space. 4. Time. "Places" change over time. For example, a new business can open in place of another. Buildings are built and demolished. If you are working with data over a long time period, this is more of a concern. Unfortunately, there are many-to-many relationships between all of these. For any address, there are many possible locations. For any location, there are many possible addresses. For any set of address components, there are many possible formatted addresses. For any address, there may have been multiple different "places" found there over time. Et cetera. Consider an office building, for an example. That building probably has at least one address corresponding to its main entrance, but it might have other entrances. An office within the building can sometimes use the building's main address as its own address, or the address of its own entrance, or an address with a floor number, or an address with a mailroom number. An office building may have a café (usually off of the lobby) that has its own addresses. Offices in the building can be split and merged over time. I have even seen multiple buildings merged into one. You might also have a co-working space with individual addresses for rooms within it.