5 ms·
So... for the one occasion out of a million where somebody needs to "debug" a piece of data, it's necessary to suffer the bloat of a text format for every other
by strags 15y ago
So... for the one occasion out of a million where somebody needs to "debug" a piece of data, it's necessary to suffer the bloat of a text format for every other piece of data we transmit?
How about we just standardize on a binary data representation (eg. MessagePack), and use common tools to export/import to/from a human-readable format? Best of both worlds.
And, as an aside - why are we using XML? It's ok as a markup language, I guess, but as a container for data? We could hardly have picked a worse format:
It's crazy verbose - even when compared to other text formats (eg. JSON). Compare:
values: [1,2,3]
with:
<values>
<value>1</value>
<value>2</value>
<value>3</value>
</values>
Its verbosity makes it hard to read, and hard to edit.
It has a poor mapping to the structures we actually use while programming - it has no built-in notion of arrays. It has superfluous node "attributes" that don't map well to common run-time constructs.
- specialist 15y agoplaintext + gzip is preferable to binary. The early adopters of XML saw their choices as semistructured text vs nice parseable XML. The third choice, creating grammars, was ignored. Grammars for most configuration or data transfer or protocols are trivial. Certainly since ANTLR 2.x. Much more trivial than any equivalent XML-based parse and validation tool stack. FWIW, ASN.1 is worse that XML. Not a defense of XML; all uses of XML are incorrect. For my own work, I use a format descended from VRML that I call ARON (A Righteous Object Notation). It concisely describes groves (trees annotated with key/values) and supports most commonly used datatypes. So it's a bit more concise that JSON or YAML and a lot more strongly typed. Here's an example test file: http://code.google.com/p/aron/source/browse/trunk/test/cronk/test1.aron http://code.google.com/p/aron/source/browse/trunk/test/cronk... I use ARON for all my own projects, as you can see, it's not really polished enough for others (yet). As this example shows, I mostly use it to loft Java object graphs. I haven't reimplemented VRML's prototyping (DEF / USE) functionality in this branch (yet).
- strags 15y ago>>> plaintext + gzip is preferable to binary. I'm not sure I agree with you on that. You're imposing a load of extra CPU (and, I suspect some bandwidth) overhead where it's not necessary for all except the most infrequent cases, and you're inheriting all the weaknesses of text as a data format. Plus, since your data is now gzipped, it's no longer human-readable on the wire. In order to read it, you need to pipe it through a decoder (gunzip) - why not use a sensible binary protocol, and pipe it through the decoder for that?
- guard-of-terra 15y agoIt is human readable if you use firebug or vim. And that's what you use.
- derleth 15y ago> In order to read it, you need to pipe it through a decoder (gunzip) - why not use a sensible binary protocol, and pipe it through the decoder for that? I don't have your decoder. I have gunzip. gunzip is not threatened by a patent. gunzip doesn't cause a Drama Meltdown. gunzip won't be a proven attack vector for remote execution exploits. gunzip does not require a contract in my hand or money in my bank. While your decoder is being debugged, gunzip will be live. (To the tune of "The Revolution Will Not Be Televised")
- strags 15y agoOh, I'm just advocating using a standard binary serialization format like MessagePack that is far more efficient, can easily represent binary data, and has a far more obvious mapping to runtime data structures.
- seanp2k2 15y ago>"standard binary serialization format" Mmmm, yes, good luck with that :)
- strags 15y agoA man can dream.
- stock_toaster 15y agoIn most parsing comparisons I have seen, tnetstrings[1] is probably one of the faster textual formats to parse (and rather "safe" too). I have started using tnetstring (python module) for some backend message formats where I commonly use json, and it has been pretty great so far. (esp with gzip, snappy, bzip, lzma, or whatever compression fits your needs/tradeoffs) [1]: https://github.com/j2labs/cerealization https://github.com/j2labs/cerealization
- guard-of-terra 15y agoXML has XPath. JSON and YAML do not have xpath. They have those lame language-specific constructs that aren't recursive, aren't traversable and throw null pointer errors when they can't match. You can format XML with xmllint and query it with xmlstarlet. You can't do that for json - good luck if you've got unformatted json. Tooling is everything.
- specialist 15y agoI added lightweight path expressions to LOX (Lightweight Objects for XML). It's a modern XML object model, that's easy to inspect, with built in support for xpath-like expressions (globbing, really), without all the suckage. http://code.google.com/p/lox/ http://code.google.com/p/lox/ Here's some example of path expressions. http://code.google.com/p/lox/source/browse/trunk/test/lox/test/Expression.java http://code.google.com/p/lox/source/browse/trunk/test/lox/te... Works fantastically. I use LOX for all my own XML work. My ARON project doesn't support path expressions in the same way. ARON's grammar only supports drill down dot notation, like "parent.child.grandchild". And, thus far, I've only used ARON to loft Java object graphs, so I haven't needed an object model with path expressions. I'd love to see a high level, statically typed language with built-in path expression. Groovy's GPath is closest to the mark that I've seen. (I really need to polish these open source side projects. And publicize them.)
- guard-of-terra 15y agoMaking your own is cool, but where's the value added compared to XPath?
- specialist 15y agoThanks! It was a lot of fun and proven very useful. Biggest reason is conciseness (less code), followed by ease of debugging. When using JXPath or Jaxen, you have to use contexts. Pseudo code (from memory): Node n = new Node( "ugh" ); // add some children here JXPathContext c = new JXPathContext( n ); NodeList list = c.eval( "child" ); for( int i = 0; i < list.length(); i++ ) { Node m = (Node) list.get( i ); ... } Whereas LOX does it like this: Element e = new Element( "ugh" ); // Add some children here for( Element e : e.find( "child" )) { ... } I do a lot of ETL work (XML -> SQL). The ability to debug (interactive inspection) speeds development. I'm sure you've tried to debug XPath expressions. Not easy. Happily, all of the LOX's object implement toString() method to render their XML content. And the evaluation of path expressions, while not easy, is feasible. Whereas with JXPath or Jaxen, it's damned near impossible. (I've written a few object adapters for JXPath; getting them right is black magic.)
- 6ren 15y agoI think there are two issues with XML/JSON/etc vs. roll-your-own format/grammar. The first issue is the difficulty of creating it, which you describe as "trivial". I think it's an "easy once you know how", but not easy beforehand, not easy for everyone, and not every aspect of it is easy even then. Another way of looking at this is that designing and writing and testing and debugging and rewriting and documenting a grammar is always going to be more work than using pre-existing code. For example, although you've put in enough work to be able to use ARON in several projects, it's not ready for others. IMHO it's a non-trivial undertaking. (BTW: does ANTLR 2.x automatically take care of left-recursive grammars?) My experience with parsing is that there be gotchas; this is one reason I'm very impressed with the principles behind XML Schema. (It's a shame it's so soul-destroyingly horrific to actually use. Also XML namespaces: righteous concept, diabolical realization.) The second issue is familiarity/standardization vs. specific-use formats (a kind of DSL), purely in terms of ease-of-use and adoption. I think familiar formats really are easier to read - your eye and brain have internalized short-cuts for interpreting them, so you can quickly work on them. Similar to your finger-memory for your editor. An alternative must be significantly better to overcome this unfair advantage of the incumbent. However, counter-example: for both these reasons (the first is a kind of mechanized version of the second), I thought it would be impossible to displace XML, much as ASCII seemed impossible to displace (there used to be competing character encodings; the present standard, unicode, can be seen as a superset of ASCII). Although JSON seemed wonderfully better (I first saw it in the ActionScript version of ECMAscript), I thought it couldn't overcome XML's incumbency. Yet, JSON is making inroads, and not just in web-clients, but in APIs. It's interesting that JSON lacks analogs to XML Schema, XSLT, XPath; although proposals have been made, they aren't adopted. I wonder if this is because people are now figuring out how to do without these extras? Or if, when they do need them, they just use the XML stack... I think new formats/grammars have the best chance of adoption in niches, when they are so specialized to the task that their benefits far outweigh the advantages of the standardized incumbent - in that particular use-case. For example, most programming languages define their own grammars e.g. python, ruby, java, etc. BTW: in your other comments with code, just indent each line by two spaces for nicer formating. :-)
- specialist 15y ago(Love your metaphors/adjectives.) Grammars aren't easy, but not currently no harder than XSD. Further, DSLs are all the rage. Personally, I prefer to use a pro tool like ANTLR to some in language voodoo with Scale (or some such). LALR (lex/yacc) kicked my ass. Just could not understand it. LL(k) (ANTLR 2.x, JavaCC) was a struggle. But I slogged thru it. LL(*) (ANTLR 3.x) is a pleasure to work with. I really enjoy it. For example, the delta between SQL's (idealized) BNF and my SQL ANTLR grammar is very small. ANTLR 4.x, currently in progress, will do left recursion, which I haven't played with yet. I'm not too worried about adoption. My tools make me more productive. If my competitors choose to use rocks for hammering, I'm okay with it. Lastly, I indent with tabs. Old habits die hard. :)
- bo1024 15y agoI think there is a key distinction here between messages we pass between applications and data we store on disk. These problems have different tradeoffs, and a conflict probably occurs when the same bits are used for both purposes. I think the fundamental difference is that messages seem to have these sorts of properties: - temporary and transient; exist primarily to move information from one program to another - may be merely copies/representations/serializations of information that already exists in a running program - may specify implementation ("the following is a hashmap of strings to strings consisting of k1,v1;k2,v2;,...") - may use a common format/language (e.g. XML) to encode arbitrary data structures - therefore, can use commonly available parsers - not usually read/edited by humans It seems like binary data is very often a good choice here, especially since performance seems to matter. But I think the author's use case is very different. S/he's dealing with code. (Data, code, tomato, tomato.) It has different properties: - permanent: stored to disk, possibly copied to other disks, but not "created and destroyed" - is the real true representation of the actual object in question; as the object is changed, the data is edited/overwritten on disk - usually descriptive of problem rather than giving implementation details ("line (7,2) -> (8,4)" rather than "<class>[line] <member>[start] <value>[<pair>[7,2]],...") - is usually domain-specific rather than an arbitrary common language like XML - thus, tends to require a custom parser to read - primitive values transparently map to changes in the object In this latter case (think HTML, LaTeX, and all programming language files), we get tons of good benefits from plain text. It can be read in any editor, it's universally and trivially portable, it can be manipulated with tools like grep and find/replace, it can be generated or altered with simple scripts and programs, etc. And finally, it's primary purpose is to be compiled into a representation which is presented to the user (such as a webpage or pdf document or so on). Those are the two extremes, I think. Data vs messages. So we can argue about which option is better for cases that seem somewhere in between, but this is the landscape as I understand it.
- strags 15y agoYour comparison of data vs. messages is interesting, although I think what you really should contrast is messages/data vs. markup. HTML and LaTeX are clearly markup languages (or "code"). It's entirely reasonable to expect somebody to edit HTML and LaTeX in a plain old text editor. Totally agree with you on this. Where I disagree, however: The OmniGraffle file in the article, isn't code - it's a serialization of an internal data structure. The primary method for editing this data is not a text editor, nor should it be. While it's cool (I guess) that the author was able to hack around inside it, I don't think XML is a good choice here, for the reasons I described earlier. Nor do I think it's a good choice for most object serializations (including both transient messages, and persistent data). Now, the OmniGraffle file in question is pretty simple, so you could argue that maybe XML isn't so terrible. But, consider cases like http://en.wikipedia.org/wiki/COLLADA http://en.wikipedia.org/wiki/COLLADA . Storing 3d object vertex data in XML is, if you ask me, insane. If you have an object with tens of thousands of vertices, you will never edit this file by hand. What is XML gaining anybody here? Yay - you can use an off-the-shelf XML parser! But you then have to copy XML's graph into your own vertex structures in order to do anything useful! So, you really haven't gained anything except vastly increased memory and CPU usage. See http://collada.org/public_forum/viewtopic.php?f=12&t=25&start=0 http://collada.org/public_forum/viewtopic.php?f=12&t=25&... for some discussion on this.
- seanp2k2 15y agoThe biggest problem with XML is that people don't understand how to implement it properly and end up with stuff like you talk about above. Namespaces in particular can help with the problem, but I see tons of APIs failing to implement them in ways that actually alleviate the bloat, and they instead opt for insanely verbose Java-style naming conventions that do implement namespaces and attributes that ultimately just make concise XPath a nightmare.
- nradov 15y agoWe use XML as a container because it supports namespaces and XPath / XQuery. There are no comparable JSON standards which are widely interoperable across different languages and systems. Advanced database engines like IBM DB2 can store XML documents internally in a compressed binary format suitable for fast searches. So there's no extra overhead for tags.