3 ms·
One thing not mentioned in the article is the lossy nature of XML when coming from data structures with concepts of arrays and maps. A typical array in XML look
by deckeraa 6y ago
One thing not mentioned in the article is the lossy nature of XML when coming from data structures with concepts of arrays and maps. A typical array in XML looks something like:
<foo>
<bar>First element</bar>
<bar>Second element</bar>
</foo>
which can lead to issues if the second element isn't in there: should bar just be a property of foo or should it be an array?
(I'm definitely open to correction on this if someone knows a better way; my experience with this particular issue is from writing a parser to take XML and turn it into a MUMPS-native data structure. I mostly had to have the parser take an educated guess and then let API users specify which elements were arrays if they wanted it to parse the same way against all inputs).
This is analogous to how saving Clojure sets (e.g. #{:a :b :c}) to a JSON document in CouchDB is also a lossy operation because JSON doesn't have sets, just maps and arrays.
- cmwelsh 6y agoSeems like you’re more interested in encoding your implementation details into a text document. This seems like something that could be left to the parser. Provide the target types for the deserialization. This isn’t Python pickle or Java serialization - it’s a text string.
- deckeraa 6y agoFair enough. Though wouldn't that work better with languages like Java where the programmer typically ends up making a type or class for most things? With languages like Clojure it's fairly typical to declare no classes or types and to simply have your data structure be a composition of maps, sets, and vectors, etc. I suppose you could use your Clojure.spec keywords as the tag names in XML, but that's probably going to be more work than (clojure.edn/read-string (slurp "my-data.edn")). Could be an interesting option if you had to use XML as an interchange format.
- rcxdude 6y agoProblem is XML is extremely frequently used to represent such types (probably even more so than it's used to represent marked up text). And it's not very good at it.
- sk5t 6y agoIt's the same problem you'd get by looking at a JSON document out of context ("is this array-like thing an array, list, set? ordered or unordered? type bounds on members? are these integers or decimals with no remainder?") You gotta refer to a schema document or human-readable document or whatever and not guess.
- deckeraa 6y agoI'd agree that things like JSON and EDN don't convey all the type information that you might want (things you mention like type bounds, orderedness, number format). Though at least in a JSON string like {foo: [elementOne elementTwo]} you know that foo is a list of things, which is something that the XML document doesn't encode for you. So perhaps I should make my statement more precise and say that XML data, in the absence of the schema information, is lossy compared to some formats like JSON or EDN. [Edit: to answer the potential question of why I'd be using XML without have schema information, I was being provided the XML by an API that called web APIs and converted the result of the web call into XML (even if it was JSON) before passing it to me.]