4 ms·
Another problem I see is that we have a snapshot of data from friday. We cannot really link our data back to any of the original OSM data. So if we want to upgr
by robmil 14y ago
Another problem I see is that we have a snapshot of data from friday. We cannot really link our data back to any of the original OSM data. So if we want to upgrade our dataset, we have to throw everything away that we have and start a new import.
This is the biggest hurdle to overcome, in my experience. A custom data format is typically essential (most location databases arrive as CSVs or XML, which are useless for real-time querying), but imports can take forever.
It's sometimes, counterintuitively, been more worthwhile to concentrate on the performance of importing than of querying; the out-of-the-box query performance you get with (no)SQL often isn't terrible, but your import script usually starts out pretty awful.
- klaustopher 14y agoYeah, it basically is ... We use osmosis to preparse the data and than just parse the XML ... It is a pain. I think what we can try is to import the data in the same format they have it in OSM and use smart indexes within the database to issue queries quickly. I think this will be the major investigation when going forward.
- robmil 14y agoSounds interesting — it would definitely be worth a follow-up post if you do manage to get that working!
- michaelt 14y agoLast weekend I did some work with the OSM planet file - the thing with the XML format is it took several hours just to decompress it - even though I was reading it from RAM on an EC2 m2.2xlarge instance. And after that it still took an age to parse all the XML. All told it took 24 hours just to decompress the file and do a three-pass parse. With the benefit of this experience, I decided it was worth switching to OSM's alternative 'PBF' format [1]. It's a dense binary format that doesn't require additional compression. It's also reportedly 6 times faster to read than gzipped XML. Honestly it seems very complicated to parse, but if you're willing to work with Java or C there's a parser already available. [2] [1] http://wiki.openstreetmap.org/wiki/PBF_Format http://wiki.openstreetmap.org/wiki/PBF_Format [2] https://github.com/scrosby/OSM-binary https://github.com/scrosby/OSM-binary
- klaustopher 14y agoYeah, next time we will use the PBF, but there's currently no ruby parser for this. If we continue to regularly parse the data, we will have to write a wrapper around the C parser. Thanks for the link
- pygy_ 14y agoOne of the best protobuf parsers is Haberman's upb. There are no Ruby bindings, but it is built with dynamic languages in mind. There are already Lua and Python bindings, you could use them as example. https://github.com/haberman/upb/wiki https://github.com/haberman/upb/wiki https://github.com/haberman/upb/tree/master/bindings https://github.com/haberman/upb/tree/master/bindings
- jandrewrogers 14y agoThe export formats of geospatial data tend to be extremely inefficient to parse. Unfortunately, the very high computational cost of parsing is multiplied by the very large size of the data. Formats like XML, JSON, and CSV are only convenient when the absolute size of the data being parsed is relatively tiny. I consider efficient export formats for geospatial data (or similar rich/complex data sources) to be a bit of an unsolved problem. It is not difficult to design storage formats that are literally a couple orders of magnitude faster to process but the formats most people are using were designed for files small enough that parsing efficiency doesn't matter. Consequently, at my company we spend time designing highly optimized parsers for inherently inefficient formats and designing non-standard internal export formats that nothing else understands but which are nonetheless vastly faster to use at scale. It is a big problem that it seems like it should be solved by now.