13 ms·
In Defense of OpenStreetMap's Data Model
- deleted 4y ago[deleted]
- gennarro 4y agoSkip about 1/3 of the way down and the OSM article starts. “This is why bad design is everywhere...”
- nkozyra 4y agoSeriously, there's a good article buried in there. Are we going to see the recipe site SEO anecdotes propagate.
- MontyCarloHall 4y agoThe entire article can be summed up as: “OSM stores maps as graphs, in flat files where each line is either a node, an ordered list of nodes, or metadata. The graph nodes can be arbitrarily ordered in OSM files, which leads to computational complexity when parsing them. This is not a bad thing, since it means that the spec for OSM files can be extremely simple, which makes it easy for people to contribute to OSM. Other mapping formats optimized for parsing speed require a lot of irrelevant fluff that makes them much harder to understand by human contributors.” Ironically, 95% of this article is irrelevant fluff that does not make it any easier for the reader to understand.
- TuringTest 4y ago> OSM stores maps as graphs, in flat files where each line is either a node, an ordered list of nodes, or metadata. The graph nodes can be arbitrarily ordered in OSM files, which leads to computational complexity when parsing them. This is not a bad thing, since it means that the spec for OSM files can be extremely simple, which makes it easy for people to contribute to OSM. That's actually a sensible design. Treat user-facing stored data as user interface. If you need efficient processing of that data, such as fast parsing, you can always build it elsewhere, such as by caching that data into an intermediate structure that is recompiled whenever the user data changes.
- SteveCoast 4y agoSomeone gets it :-)
- seoaeu 4y agoWait, the proposed solution to a data format being slow to parse is to work around the bad performance by caching the already parsed representation? That seems like it has a clear flaw if you’re only accessing the data once…
- TuringTest 4y agoWhere's the flaw in that? If you're only accessing the data once, why does it matter how fast or slow it is? And, you're suggesting that user-facing data should be harder to work with only to make it faster to parse?
- seoaeu 4y agoAccessing any given data once. When you have a total dataset size in the 10-100s of gigabytes range, having to download any significant fraction of it to do data processing is really unfortunate. But seriously what's up with this total disdain for anyone trying to build applications with OSM data? You don't seem to care whether parsing is near instant or as other commenters have mentioned, literally a majority of total processing time for certain compute jobs
- xg15 4y ago> Treat user-facing stored data as user interface. Are you telling me, the main mode of contributing to OSM should be to edit XML files and put in GPS coordinates by hand? That would be about the most user-hostile UI for map editing I could think of.
- photochemsyn 4y agoThanks! I read the article, I read the post the article is responding to, I read all the comments and still I had no real idea what it all was about until I read your comment. It could be an example of an author assuming a general audience already knows the insider information but then I don't know who the target audience really was. This is the kind of thing that probably should have been spelled out in the introduction, with a link to something like this: https://labs.mapbox.com/mapping/osm-data-model/ https://labs.mapbox.com/mapping/osm-data-model/
- jtbayly 4y agoI made it about 1/3 through and still didn’t know what it was about. That is poor design for sure.
- Sujan 4y agoI think those are the important bits: > The Engineering Working Group (EWG) of the OSMF has “commissioned” (I think that’s OSMF language for paid) a longstanding proponent of rules and complexity to, uh, investigate how to add rules and complexity to OSM. > [...] > Let us pray that the EWG is just throwing Jochen a bone to go play in the corner and stop annoying the grownups. It's a "response" to https://blog.openstreetmap.org/2022/06/02/announcement-data-model-study/ https://blog.openstreetmap.org/2022/06/02/announcement-data-...
- everybodyknows 4y agoThis seems an important bit to me: > Facebook solved this in a beautifully OSM-like way: daylight. Daylight is a sanitized, consistent and cleaned up map based on OSM https://daylightmap.org/ https://daylightmap.org/ https://registry.opendata.aws/daylight-osm/#usageexamples https://registry.opendata.aws/daylight-osm/#usageexamples https://gist.github.com/jenningsanderson/3e42a99dcb8f760038ad8aa47ea38ce8 https://gist.github.com/jenningsanderson/3e42a99dcb8f760038a...
- RicoElectrico 4y agoThe proposed improvements would obsolete a bunch of problems such as broken polygons [1] which happen regularly. They would also make processing OSM more accessible without needing to randomly seek over GBs of node locations just to assemble geometries which takes a significant runtime percentage of osm2pgsql. For me Steve Coast lost his credibility when he joined the closed and proprietary what3words. [1] https://wiki.openstreetmap.org/wiki/OSM_Inspector/Views/Multipolygons https://wiki.openstreetmap.org/wiki/OSM_Inspector/Views/Mult...
- TuringTest 4y agoBut you don't need to complicate the storage format to fix a problem like that. You can build validation tools that will check whether the stored data conforms to the correct specified geometry, and only emit valid polygons to later tools in the pipeline when they do. "Be liberal in what you accept and strict in what you send" is still a good principle. The problem with rejecting invalid structures at the data storage format instead of a later validation step is that it hurts flexibility and extensibility. If later on you need a different type of polygon that would be rejected by the specification, you'll need to create a new version of the file format and update all tools reading it even if they won't handle the new type, instead of just having old tools silently ignoring the new format that they don't understand.
- RicoElectrico 4y agoThe best thing is not to allow invalid geometries to begin with. Any validation would need to be done in an off-line fashion for a number of reasons (such as needing to retrieve any referred OSM elements), and by that time you can't automatically revert offending changes as any revert carries a chance of an object version conflict.
- TuringTest 4y ago> The best thing is not to allow invalid geometries to begin with. The best thing for whom? The developer? Certainly not for the end user, who needs to have invalid geometries while the drawing is being made and the data is still incomplete. Having a file format that won't admit that temporary state means that either the user can't save incomplete draft work, or that an entirely different format will be needed to represent such in-process work. The article is rightfuly critizising that such incomplete way of thinking, that doesn't take into account the full picture nor the systemic effects of a change, is pushed forwards only because they seem "the right thing" from an incomplete understanding of all the concerns and the needs from all stakeholders. The right technical decision *must* include them to be correct, and the best design might involve a solution other than "update the file format so that it doesn't accept inconsistent geometry (acording to the set of rules that we understand as of today)". But to assess what the right decision is, you need to know how people is using the system in real use-cases beyond classic comp-sci concerns of data storage and model consistency; and to learn those, you need to talk to end users and perform field research to inform your decisions and designs.
- kawsper 4y agoI’ve started playing with data from OpenStreetMap. It started with me trying to fetch all the places where I could get water when moving around Copenhagen, which turned out not to be as easy as first envisioned, because OSM seems to have a lot of different ways to categorise available water, which makes sense, OSM and the tagging system isn't there to support only my usecase, and describing my idea doesn't fit 1:1 with the model. I identified the following tags to look out for: amenity=drinking_water, https://wiki.openstreetmap.org/wiki/Tag:amenity%3Ddrinking_water https://wiki.openstreetmap.org/wiki/Tag:amenity%3Ddrinking_w... man_made=water_tap, https://wiki.openstreetmap.org/wiki/Tag:man_made%3Dwater_tap https://wiki.openstreetmap.org/wiki/Tag:man_made%3Dwater_tap amenity=water_point, https://wiki.openstreetmap.org/wiki/Tag:amenity%3Dwater_point https://wiki.openstreetmap.org/wiki/Tag:amenity%3Dwater_poin... drinking_water=*, https://wiki.openstreetmap.org/wiki/Key:drinking_water https://wiki.openstreetmap.org/wiki/Key:drinking_water It's a tough problem to map out the world and describe it, especially when everyone can add or modify the data, but anything that could improve the experience of importing like osm2pgsql would be welcome.
- Aachen 4y agoI don't understand how this doesn't fit your use case. The tags are for different things, e.g. > for places where you can get larger amounts of "drinking water" for filling a fresh water holding tank, such as found on caravans, RVs and boats versus > a man-made construction providing access to water, supplied by centralized water distribution system (unlike in case of man_made=water_well [...]). The tag man_made=water_tap is used for publicly usable water taps, such as those in the cities and graveyards. Water taps may provide potable and technical water, which can be specified with drinking_water=yes and drinking_water=no. And another tag for when you're not mapping a separate water point, but indicating whether a given feature has drinking water (for example a well or mountain hut). You're saying that it's tough when anyone can mess with the data rather than working in a structured way, but these tags have distinct definitions and seem perfectly sensible to me (there are much worse examples like highway=track, which spawned huge discussions in various places within the community). How do these tags not match your use case to select which tags you need and display those in the way you want (e.g. as list or map)?
- tinus_hn 4y agoWhat is stopping users who have a problem with the model from transforming the data into a form that is better for their use case?
- RicoElectrico 4y agoIt is surprisingly difficult to say which closed ways are areas and which are not. This depends entirely on tags of the way and is only solved by heuristics. https://github.com/tyrasd/osm-polygon-features https://github.com/tyrasd/osm-polygon-features
- matkoniecz 4y agoIn addition, it is common to have objects that are both area and line at once. Or area according to one tool/map/edtor and line according to another. And many, many multipolygon relations are in inconsistent state and require manual fixup. Also, complexity of entire area baggage makes explaining things to newbies more complex. You can either try to hide complexity (used by iD in-browser-editor) leaving people hopelessly confused when things are getting complex or present full complexity (JOSM) causing people to be overwhelmed. See https://wiki.openstreetmap.org/wiki/Area#Tags_implying_area_status https://wiki.openstreetmap.org/wiki/Area#Tags_implying_area_... for a start of a complexity fractal.
- maxerickson 4y agoThat's mostly what people do. The current format stores locations and references to locations, so for example, a line feature only stores references to locations, so to realize it on a map, you have to go through the data and find all the locations it references and build up the actual geometric feature. So people do caching and so on, for sure. The proposed changes would make that sort of data transformation easier and less resource intensive.
- stevage 4y agoIn the first part of the article I was thinking, oh, maybe Steve Coast isn't such a jerk after all. Then I got to the meat of it. Oh dear. As one of the many many people who has had to deal with OSM data, I curse people with this attitude that the mess is somehow desirable or necessary. It's not. There is a long spectrum between totally free form and completely constrained, and OSM's data model is painfully down the wrong end, and causes enormous harm to all kinds of potential reuses of the data. It also causes harm to the people creating data. Try adding bike paths and figuring out what tags are appropriate in your area. Try working out how to tag different kinds of parks, or which sorts of administrative boundaries should be added or how they should be maintained. It puts many people off, me included. Bah.
- delusional 4y agoI tend to agree with you, having done a fair bit of cursing at the OSM format as well. Yet they've made an open source map, and I haven't. The data tells me that I'm wrong.
- maxerickson 4y agoFor a crowd sourced dataset, a strict ontology anyway wouldn't work. Instead of messy tag definitions you'd have tag use that didn't align with the definitions. I don't mean that as an argument against improving the tagging! The biggest friction point is probably that people resist rationalization of tagging schemes that have demonstrated themselves to be problematic. The tagging system in the iD editor tries to address the issue, supporting search terms and suggesting related tags and so on. The article is more about the underlying storage of the geometries (I don't think there is the same level of interest in changing the basic approach to tagging/categorization).
- agumonkey 4y agoMaybe there's a need for bridging app. Something to aid,suggest,review. So people could spend their energy slowly but surely ? an OSMCAD
- 4y ago
- JackFr 4y agoI know it’s nothing to do with the main thrust of the article, but the author fundamentally misrepresents KYC. Know-your-customer is a facet of anti-money laundering and anti-corruption regulation. It has nothing to do with talking to users.
- myself248 4y agoPerhaps an existing term was co-opted by financial legislation...
- bornfreddy 4y agoMaybe, but not likely. The quoted text fits the common term definition: > The answer, as any product owner will tell you, is to get close to the customer. To talk to them. To understand them. To feel their pain. The {big short}: > Deutsche Bank had a program it called KYC (Know Your Customer), which, while it didn't involve anything so radical as actually knowing their customers, did require them to meet their customers, in person, at least once.
- londons_explore 4y agoThe exact same argument that praises OSM's super flexible tagged node data model should also praise MS Excel for the number of things that can be achieved in the world of business with just a grid of boxes. Both have been hugely successful, and both have the same pile of downsides.
- TuringTest 4y ago> Both have been hugely successful, and both have the same pile of downsides. Exactly. And the solution should not be to throw away spreadsheets completely and turn them into relational databases, but to create new tools to alleviate the downsides and reduce their impact (possible by exporting the spreadsheet information into a relational database, but without taking away the user's option to continue working with it.)
- Doctor_Fegg 4y agoNo one is suggesting throwing away OSM's data model completely. The current suggestion is basically "maybe we should think about a point release to properly address an ugly hack we invented in 2007".
- pramsey 4y agoRight? The emotion of the response seems completely out of scale to the ambition of the reform proposed. "Maybe polygons?"
- sp8962 4y agoNot really even that. What is currently on the table is simply a way to cleanly differentiate between closed ways that are polygons and actual closed ways. Example roundabout enclosing a park. The problem is that right now this relies on determining this from the tagging. This could well be implemented as a flag on the existing way type and not as an actual new datatype. There is at this stage no intention to revamp the way how we model areas that are more complex than the single polygons from above, that is with multi-polygon relations. The more controversial topic is giving OSM way objects partially or fully their own geometry. The former would have for all practical purposes no noticeable contributor effect outside of geometry changes always creating new versions of ways, contrary to the current behaviour which can be somewhat puzzling for newbies. The later would be quite drastic, but would provide more benefits for at least some kinds of processing, for others not, as then topology would have to be inferred. In any case the 90% of the discussion on this topic fretting about tagging is completely misplaced as literally nobody is even remotely considering changing that.
- zigzag312 4y agoAuthor states: >And that’s the point, rules and complexity have completely unknowable downsides. Downsides like the destruction of the whole project. With each rule and added complexity you make the system less human and less fun. You make it a Computer Scientists rube goldberg machine while sterilizing it of all the joy of life. While too much rules and complexity can certainly be bad, some basic amount of standardization can actually reduce complexity and really doesn't cause a "destruction of the whole project". As a counterpoint, too much flexibility can also increase complexity. For example, without defined rules, 5.6.2022 can mean 5. June 2022 or 6. May 2022. Nor user, nor parser can know for sure what it means, if standard isn't defined. This kind of flexibility certainly isn't fun. Example from OSM wiki for "Key:source:date": > There is no standing recommendation as to the date format to be used. However, the international standard ISO 8601 appears to be followed by 9 of the top 10 values for this tag. The ISO 8601 basic date format is YYYY-MM-DD. https://wiki.openstreetmap.org/wiki/Key:source:date https://wiki.openstreetmap.org/wiki/Key:source:date Just define some essential standards. It won't lead to destruction of the project! And while you are making breaking changes, please fix the 'way' element. Maps are big. Storing points in ways as 64bit node-ids, while coordinates in nodes are also 64-bit (32bit lon and 32bit lat), just leads to wasted space and wasted processing time. There are billions of these nodes and nearly all of these nodes don't have tags, just coordinates. There is no upside for this level of indirection. And in case tags are needed for a point, this can already be solved with a separate node and a 'relation' element. OSM data format could certainly be improved and it would benefit end users, as better tools/apps could be made more quickly and easily.
- atoav 4y agoThe date example is a good one. No one has fun by choosing their own date format. This is putting the burden of choice onto the user. They might like to think about some map stuff and now they have to think about data format stuff. Of course projects like these have to strike a balance between the strictest bureaucratic nightmare and such a structure so loose that people are overburdened by the available options at every corner. I think a lot of that complexity can (and should!) live in the tools themselves. Who cares about a date format, when the tool that creates it offers a date picker or extracts the correct date from the meta data of an image? The date format in the backend should be fixed and then you should offer flexibility in the frontend for user input.
- matkoniecz 4y ago> The harder you make it for them to edit, the less volunteers you’ll get. And that is why dedicated area type (rather than representing areas with lines or special relations[0]) could help new mappers and new users of data. There would be very significant transition costs, but maybe it would be overall beneficial. It is possible to have objects that are both area and line at once. Or area according to one tool/map/edtor and line according to another. And many multipolygon relations are in inconsistent state and require manual fixup. Also, complexity of entire area baggage makes explaining things to newbies more complex. You can either try to hide complexity (used by iD in-browser-editor) leaving people hopelessly confused when things are getting complex or present full complexity (JOSM) causing people to be overwhelmed. See https://wiki.openstreetmap.org/wiki/Area#Tags_implying_area_status https://wiki.openstreetmap.org/wiki/Area#Tags_implying_area_... for a start of a complexity fractal. [0] https://wiki.openstreetmap.org/wiki/Area https://wiki.openstreetmap.org/wiki/Area
- pramsey 4y ago1000 times yes! I am a spatial data expert but only a some-time OSM editor and I still have yet to figure out how to create a polygonal feature more complex than a single building footprint. The theoretical advantage of a unified topology model of just nodes/edges where polygons and lines share core geometry is nullified by cultural rules that say "don't do that" to editors (I had a bunch of parks that shared a boundary with a road reverted with nasty notes). The current setup is not just hard for processors, it's hard for non-experts to understand and therefore a higher barrier than a simple polygon model would be.
- matkoniecz 4y ago> how to create a polygonal feature more complex than a single building footprint In ID (default editor) you can mark area and area inside or select two disjointed areas and press right click on the and select "Merge". Or press "c" while selecting areas for combining. In JOSM there is equivalent "create multipolygon" (or "update multipolygon") https://wiki.openstreetmap.org/wiki/Relation:multipolygon#How_to_map https://wiki.openstreetmap.org/wiki/Relation:multipolygon#Ho... > parks that shared a boundary with a road FYI, that is because highway=* road line represents centerline of carriageway - and unless park somehow ends in the middle of road and includes half of its surface it will be not correct. It also makes future editing quite nasty.
- matkoniecz 4y ago> Let us pray that the EWG is just throwing Jochen a bone to go play in the corner and stop annoying the grownups. That is neither helpful, not useful, nor making me more likely to treat this diatribe more seriously.
- NelsonMinar 4y agoIt is the kind of disrespectful rhetoric that defines the OSM community though.
- matkoniecz 4y agoI do not consider it as defining and definitely nor desirable or improving ones standing. For reference: I am extremely active in OSM community. On channels that I moderate this would result in user being warned/kicked (but not banned, unless in case of repetitive insults).
- uptime 4y agoI have been looking into improving kerb and traffic_signals data for some bboxen. It is daunting, and I figure I need to work backwards - try to find out what the accessibility map apps look for and use those pairs. If I know what target I am shooting for I guess it will be alright. This is like my first week looking into this and I hope to find these targets soon.
- matkoniecz 4y agoWhat exactly you are trying to do? Import data? Map something manually on your own? Something else?
- uptime 4y agoAdd keys to existing nodes mostly. Possibly using tasks.openstreetmap.org and/or possibly doing something in a batch if I can get data from the city to use. These structures seem well defined, thankfully. And the crossings and signal locations look to be complete.
- matkoniecz 4y agoIn this case I would strongly encourage to start from manual mapping. StreetComplete Android app may be useful here (disclaimer: I am involved in making it). See also https://wiki.openstreetmap.org/wiki/Import/Guidelines https://wiki.openstreetmap.org/wiki/Import/Guidelines before importing data
- uptime 4y agoThank you!
- chaps 4y agoHeh, their data model is 99% of the reason why I don't use OSM. It's scattered all over the place with so many tables! It's such a nice project, but damn is it impossible to work with programmatically, let alone poke around it to discover what's all in there.
- matkoniecz 4y ago> so many tables ? There are various complaints about OSM data model but this is a new one to me. In OSM basically everything is mixed together and there is no real separation into layers. What you mean by "many tables"?
- peace4all 4y agoHe probably means what the osm2pgsql import tool creates.
- matkoniecz 4y agoOK, then it is about default osm2pgsql data model that in part is independent from OSM data model. "many tables" is definitely osm2pgsql design decision
- npteljes 4y agoWould the Overpass API (have) fit your use case?
- seoaeu 4y agoThe claim that a dataset with billions of users has only dozens of people able/interested in doing data processing on it is a damning admission that the format is too hard to deal with
- Doctor_Fegg 4y agoI maintain some OSM data-mangling code - moderately popular perhaps, but certainly not core - and even that has 840 github stars. I'd take the "dozens" as poetic licence really.
- ciphol 4y agoDozens of open source volunteers who are interested in volunteering their free time to do software development using the format. In addition to the innumerable developers in Facebook, Apple, and other corporations who are paid to do the data processing and actually bring the data to those billions of users.
- an9n 4y agoThere's an amusing paradox I've seen many times amongst GIS managers - their complaints about how dreadful OSM data is are pretty much the direct opposite of their enthusiasm to use it!
- teddyh 4y agoWhat happened to the title? It used to be “In Defense of OpenStreetMap's Data Model”, which is the literal blog post title. Someone has now changed it to the boring-sounding “OpenStreetMap's Data Model”, probably resulting in fewer clicks.
- mring33621 4y agoSo, come up with an improved format that has/enforces various 'rules' and also provide a conversion program for moving between the new/old format, as desired. Slowly deprecate the nastiest parts of the old, as people get used to the new format.
- boredumb 4y agoStreet-Complete on android is a neat way to contribute to OSM.
- anticristi 4y ago> The harder you make it for them to edit, the less volunteers you’ll get. I'm not sure I like OSMs obsession with the data model. I guess this has to do with its business model, but then let's not pretend the product manager is tasked to optimize for end-users. I'd like it to focus on the UI. The easier it is to input a geographic thingie and the easier it is to visualize the geographic thingie, the better OSM both for users and volunteers. Two issues to strengthen my point: 1. The osm-tag mailing list regularly discusses how tags are visualized in various renderers when recommending which to choose. 2. Quick mobile-based correction are nearly impossible with OSMAnd. I'd love to take a picture and write a quick note like "speed limit changed", so that someone (perhaps a bot) can pick this up and update the data model. Same with restaurant opening hours. Or various POIs.
- matkoniecz 4y ago> I'd love to take a picture and write a quick note like "speed limit changed" This can be done with StreetComplete (Android app) - long press on map to create note, photo can be added. > Same with restaurant opening hours. Or various POIs. You can also outright survey this with StreetComplete. See https://github.com/streetcomplete/StreetComplete https://github.com/streetcomplete/StreetComplete Note: speed limit needs to be enabled, it is disabled by default. It is also unavailable in USA due to horrific default speed limit system which requires massive work to support. Disclaimer: I am one of people working on StreetComplete > I'm not sure I like OSMs obsession with the data model. Given effort that went into various parts - fundamental data model has not received any changes for a long time. I would not describe it as obsession. > I guess this has to do with its business model OSM do not really have business model, it is not a business
- xg15 4y agoThe author is ranting a lot about OSMF's recent decisions, gives all kinds of reasons why they will undoubtedly lead to horrible consequences and grants sage advice what should have been done instead. The only thing I'm missing is any indication that OSMF's course actually did cause any problems in reality.
- pete_nic 4y ago>If you won’t or can’t talk to customers then at least make the thing simple. The simpler something is, the more people will use it. Great advice
- westnordost 4y agoJust for context, here is the current OSM data model in a nutshell: There are three data types - nodes, ways and relations. All of the three can have any number of "tags" (i.e. a map<string, string>), which define their semantic meaning. For example, a way with `barrier=fence` is a fence. This stuff is documented in the openstreetmap wiki. A node is a point with a longitude and a latitude. A way is a sequence of nodes. A relation is a collections of any number of nodes, ways or other relations. Each member of this collection can be assigned a "role" (string). Again, the semantics of what each role means is documented in the openstreetmap wiki. To modify data, simply new versions of the edited data are uploaded via the API. --- The most prominent point that stands out here is that only nodes have actual geometry. This means that... 1. to get the geometry of a way (e.g a building, a road, a landuse, ...), data users first need to get the locations of all the nodes the way references. For relations, it is even one more step. 2. in order to edit the course of a way, editors actually edit the location of the nodes of which the way consists of, not the way itself. This means (amongst other things) that the VCS history of that way does not contain such changes