6 ms·
But you don't need to complicate the storage format to fix a problem like that. You can build validation tools that will check whether the stored data conforms
by TuringTest 4y ago
But you don't need to complicate the storage format to fix a problem like that. You can build validation tools that will check whether the stored data conforms to the correct specified geometry, and only emit valid polygons to later tools in the pipeline when they do.
"Be liberal in what you accept and strict in what you send" is still a good principle. The problem with rejecting invalid structures at the data storage format instead of a later validation step is that it hurts flexibility and extensibility. If later on you need a different type of polygon that would be rejected by the specification, you'll need to create a new version of the file format and update all tools reading it even if they won't handle the new type, instead of just having old tools silently ignoring the new format that they don't understand.
- RicoElectrico 4y agoThe best thing is not to allow invalid geometries to begin with. Any validation would need to be done in an off-line fashion for a number of reasons (such as needing to retrieve any referred OSM elements), and by that time you can't automatically revert offending changes as any revert carries a chance of an object version conflict.
- TuringTest 4y ago> The best thing is not to allow invalid geometries to begin with. The best thing for whom? The developer? Certainly not for the end user, who needs to have invalid geometries while the drawing is being made and the data is still incomplete. Having a file format that won't admit that temporary state means that either the user can't save incomplete draft work, or that an entirely different format will be needed to represent such in-process work. The article is rightfuly critizising that such incomplete way of thinking, that doesn't take into account the full picture nor the systemic effects of a change, is pushed forwards only because they seem "the right thing" from an incomplete understanding of all the concerns and the needs from all stakeholders. The right technical decision *must* include them to be correct, and the best design might involve a solution other than "update the file format so that it doesn't accept inconsistent geometry (acording to the set of rules that we understand as of today)". But to assess what the right decision is, you need to know how people is using the system in real use-cases beyond classic comp-sci concerns of data storage and model consistency; and to learn those, you need to talk to end users and perform field research to inform your decisions and designs.
- purple_turtle 4y ago> Having a file format that won't admit that temporary state means that either the user can't save incomplete draft work, or that an entirely different format will be needed to represent such in-process work. Saving such temporary state is very rarely needed in OSM and should be never uploaded to the OSM database. In addition, in almost all cases it can be simply saved as area of shape that is not yet matching intended one.
- TuringTest 4y ago> Saving such temporary state is very rarely needed in OSM and should be never uploaded to the OSM database. Maybe, but you're missing the other use case - that in the future you'll need an extension requiring geometries that are considered invalid by the current set of rules, forcing you to update all tools processing the file format to acommodate the new extension. Keeping storage and validation as two separate steps is a more flexible design, preferable on platforms where data is entered by a large number of users in a complex domain that is not easy to model inambiguously. Think of Wikipedia and what would have happened if its text format had only supported grammatically correct expressions without spelling mistakes, and without letting you save templates with any errors. The project would never have attracted the volume of editors it took to create the initial version with millions of articles, and the product would never have taken off. In an open project with data provided by the general public, keeping user data validation in the same layer as the automatic processing model is a design mistake.
- matkoniecz 4y agoI think that noone serious proposes to include rules like > You could have rules that say you can’t link Finland to Barbados. in the data model. That is a red herring. But rules like "area must be a valid area" are a good idea, in the same way as Wikipedia is requiring article code to be a text and is not allowing saving binary data there.
- Archelaos 4y ago> Maybe, but you're missing the other use case - that in the future you'll need an extension requiring geometries that are considered invalid by the current set of rules, forcing you to update all tools processing the file format to acommodate the new extension. I think the way to go is to define several layers of correctness. A data set might then be partially valid. In such cases a tool might, for example, support transitions from a complete valid state A to a complete valid state C by an intermediate partially valid state B. (As databases with referential integrity may allow intermediate states in a transaction where referential integrity is broken.)
- maxerickson 4y agoThe "expression" layer of the data model has had 20 years to evolve and has largely been static for a decade. Making everything slower and harder to retain flexibility you don't need isn't a great tradeoff.
- TuringTest 4y agoWhy do you need to change the data format to make it faster (at the cost of making it harder to work with to end users)? The data is the same as it was at the beginning, it doesn't justify a technical redesign. Why not just create accelerators based on an intermediat format?
- maxerickson 4y agoI guess I don't follow your analysis. People doing mapping tasks will use an editor and not really see the change. People consuming the data will also mostly use tools, tools that likely run much faster. I've written some code to chop up overlapping gis areas into ways and relations (to match the current data model of references to shared nodes). The input to that code is pretty close to the proposed data model, so not going to be more difficult to do that processing (as an example of a task that doesn't just use 3rd party tools).
- zigzag312 4y agoEnd users don't work with data format directly. They use tools and these tools could be better, if data format was improved.
- NateEag 4y agoSome of us do, though. Vespucci is a really handy Android app for making contributions to OSM, but it's hard to use without knowing something about the tagging conventions.
- matkoniecz 4y agoIt is plausible that data format can be made better for everyone at cost of very significant redesign cost of software interacting with it. > Why not just create accelerators based on an intermediat format? making things easier for mappers by introducing new data format requires changing format used by mappers
- matkoniecz 4y ago> You can build validation tools that will check whether the stored data conforms to the correct specified geometry, and only emit valid polygons to later tools in the pipeline when they do. It is not helping at all when the problem is that important areas disappeared. It is also not helping at all other mappers or confused newbie.
- arccy 4y agopostel was wrong https://tools.ietf.org/id/draft-thomson-postel-was-wrong-03.html https://tools.ietf.org/id/draft-thomson-postel-was-wrong-03....
- seoaeu 4y ago> "Be liberal in what you accept and strict in what you send" is still a good principle. No, it is a terrible principle which produces brittle software and impossible to implement standards. The problem is that no one actually follows the “be strict in what you send” part, and just goes with whatever cobbled together mess the other existing software seems to accept. Before long, a spec compliant implementation can’t actually understand any of the messages that are being sent > just having old tools silently ignoring the new format that they don't understand. This sounds like another headache. I don’t want my tools silently breaking.
- TuringTest 4y ago> This sounds like another headache. I don’t want my tools silently breaking. Yet here you are, posting your comment through a web browser on a web page. And the new standard that was intended to make web pages display catastrophic failure and stop processing with each error (XHTML) was never widely adopted. Makes you wonder why? Maybe the nature of an open data platform for human consumption has something inherent to it so that it's better to accept a certain degree of inaccuracy and inconsistencies in its stored data?
- seoaeu 4y ago> Maybe the nature of an open data platform for human consumption has something inherent to it so that it's better to accept a certain degree of inaccuracy and inconsistencies in its stored data? That's complete nonsense. The only reason that web browsers accept malformed webpages is because there were already orders of magnitude too many webpages that violated the relevant specs when XHTML was introduced. If web browsers had enforced XHTML from the start, then everyone would have damn well followed it.
- danShumway 4y agoI don't have horribly strong opinions here, but the argument feels circular to me: - The format should be kept simple to encourage more people to build tools on top of it, and users will be more likely to work with it. - We should deal with the emergent complexity of bad validation by making tools more complicated and having them detect errors on their end. If users are going to use a validation tool to work with data, then they can also use a helper tool to generate data. And if the goal is to make it easier to build on top of data, import it, etc... allowing developers to do less work validating everything makes it easier for them to build things. I'm going over the various threads on this page, and half of the critics here are saying that user data should be user facing, and the other half are saying that separate tools/validators should be used when submitting data. I don't know how to reconcile those two ideas; particularly a few comments that I'm seeing that validation should be primarily clientside embedded in tools. Again, no strong opinions, and I'll freely admit I'm not familiar enough with OSM's data model to really have an opinion on whether simplification is necessary. But one of the good things about user facing data should be that you can confidently manipulate it without requiring a validator. If you need a validator, then why not also just use a tool to generate/translate the data? To me, "just use a tool" doesn't seem like a convincing argument for making a data structure more error prone, at least not if the idea is that people should be able to work directly with that data structure. ---- > you'll need to create a new version of the file format and update all tools reading it even if they won't handle the new type, instead of just having old tools silently ignoring the new format that they don't understand. Again, not sure that I understand the full scope of the problem here, and I'm not trying to make a strong claim, but extensible/backwards-compatible file formats exist. And again, I don't really see how validation solves this problem, you're just as likely to end up with a validator in your pipeline that rejects extensions as invalid, or a renderer that doesn't know how to handle a data extension that used to be invalid or impossible. Wouldn't be nicer to have a clear definition of what's possible that everyone is aware of and can reason about without inspecting the entire validation stack? Wouldn't it be nice to not finish a big mapping project and then only find out that it has errors when you submit it? Or to know that if your viewer supports vWhatever of the spec that it is guaranteed to actually work, and not fall over when it encounters a novel extension to the data format that it doesn't understand or that it didn't think was possible? Personally, I'd rather be able to know right off the bat what a program supports rather than have to intuit it by seeing how it behaves and looking around for missing data. Part of what's nice about trying to do extensions explicitly rather than implicitly through assumptions about data shape, is that it's easier to explicitly identify what is and isn't an extension.