7 ms·
A brief XML rant
- paulnechifor 14y agoThis isn't really criticism of XML, though. You can do a good job of screwing up in any language or format.
- adamtaro 14y agoIt was not intended as a criticism of XML at all. XML is a perfectly cromulent standard. It is a criticism of amateurish use of XML.
- egeozcan 14y agoWhich is everywhere (xhtml, anyone?)
- rimantas 14y agoThere are very few websites where xhtml is served with a proper MIME type. MIME type triggers xml parsing mode in browsers, so in most cases xhtml is treated exatly as deserved: just a tag soup.
- sageikosa 14y agoThe "X" is for "extreme", right?
- Kequc 14y agoYea but I've been using haml for quite a while to generate markup. XML is horribly inefficient by comparison and prone to mistakes.
- dscrd 14y agoXML is so complex and obtuse that one can hardly blame the practicioners for misusing it.
- sageikosa 14y agoXML is obtuse only when you try to read definitions of XML dialects written in XML dialects themselves; even then, it's understandable, though it is a "high-art" discipline of schema-world semantics.
- dscrd 14y agoI agree that <tag attribute="value">data</tag> is simple, but XML is, unfortunately, much more than that. And this article complains about that "much more" part.
- sageikosa 14y agoI'll agree with your agreement; the human legibility of XML data often leads the novice programmer into making bad assumptions about the simplicity of implementing XML. While it is possible to implement "well-formed" XML easily enough, validating against schema is another matter. In this particular article's case, the "well-formedness" isn't even there.
- jarman 14y agoFor one-off, transport xml it's not much more. It's proper escaping, declaration with character set and not using features you do not know how to use. First two are solved by using proper library, third - by common sense
- masklinn 14y ago> do not use template languages to generate XML. Small correction: do not use text template languages (Jinja, moustache, erb — which seems to be the one used here considering `%= display_date %>`, raw PHP, smarty, freemarker, what have you) to generate XML. There are templating languages whose primary use case is to generate markup (including XML)[0] and (unless they're broken to uselessness) they should guarantee the output is valid XML. > Schema-design-wise, the content:encoded and excerpt:encoded element names are deeply suspect, as if someone looked at RSS 2.0, squinted, shrugged, and invented their own ad hoc analogous namespace prefix, rather than understanding the role of elements in XML. They seem to be using Wordpress's WXR import/export format, hence the wp-namespaced elements. The "content" and "excerpt" namespace garbage comes straight from there according to http://ipggi.wordpress.com/2011/03/16/the-wordpress-extended-rss-wxr-exportimport-xml-document-format-decoded-and-explained/ http://ipggi.wordpress.com/2011/03/16/the-wordpress-extended... > <content:encoded> Is the replacement for the restrictive Rss <description> element. Enclosed within a character data enclosure is the complete WordPress formatted blog post, HTML tags and all. > <excerpt:encoded> This is an unknown elementThis is a summary or description of the post often used by RSS/Atom feeds.. Considering the cottage industry of wordpress interaction, it was probably a good move to shoot for interop (should allow posterous exports to be directly imported into wordpress?). Not sure they succeeded though. [0] genshi for instance http://genshi.edgewall.org/ http://genshi.edgewall.org/
- adamtaro 14y agoAgreed on your stipulation on Genshi and the like. And thanks for the further reverse engineering of the likely intent of the export. I wouldn't disagree with most of the WP-centric design choices. But attempting to run through a real XML parser might've been a good choice as well. (And I note there's a fair bit of complaint on the WP forums about the difficulty of using the data for import.)
- timdorr 14y agoThere are templating languages whose primary use case is to generate markup (including XML)[0] and (unless they're broken to uselessness) they should guarantee the output is valid XML. Since they are using Rails, they should be using Builder for this: http://api.rubyonrails.org/classes/ActionView/Base.html#label-Builder http://api.rubyonrails.org/classes/ActionView/Base.html#labe... https://github.com/jimweirich/builder https://github.com/jimweirich/builder
- fpgeek 14y ago> Get off my lawn, you kids. Isn't that what they were doing?
- dylangs1030 14y agoUpped for giving me a chuckle in the midst of some very heated XML discussion :)
- TheAnimus 14y agoI'd just like to take a moment to mention Nested Comments. Oh if I had a £1 for every time I'd had to sift through lines and lines of code, because I can't just comment an element. I just can't comprehend why they'd need to reserve -- inside a comment.
- masklinn 14y ago> I just can't comprehend why they'd need to reserve -- inside a comment. It's because the feature was inherited from SGML, first for commenting in element declarations (e.g. <!ELEMENT -- this is an element>) and then generalized to the whole document: in SGML, the grammar for a comment is comment declaration = MDO ("<!"), (comment, ( s | comment )* )?, MDC (">") comment = COM ("--"), SGML character*, COM ("--") HTML — as an SGML application — theoretically inherited this feature (most UA don't really implement it correctly so it's not exactly safe to use sequences of dashes inside a comment, browsers may or may not toggle commenting). See http://www.howtocreate.co.uk/SGMLComments.html http://www.howtocreate.co.uk/SGMLComments.html for a more extensive explanation especially in relation to browsers (SGML-compliant comments handling used to be part of early ACID2, before being removed because it was a stupid idea) Meanwhile XML took half of it, threw the rest away, and called it a day.
- tlarkworthy 14y agoIt would take 0.5 days work to get that into any format you desire so I don't think it fails its purpose.
- jerf 14y agoNo. Once you screw up encoding, the information is generally gone. It's not just a matter of munging, it's often a matter of having to grovel over the entire file, by hand, correcting things. Programmers seem to love to think that encoding errors are a joke, but they aren't. The data is gone. That's a big deal. Why are you even writing a program in the first place if it's just going to output unrecoverable gibberish? So you can throw the onus on the user to figure it out? And that's to say nothing of trying to recover the date.
- mikeash 14y agoIt drives me bonkers. Use UTF-8. Use other encodings only when talking to systems that require it, and use those other encodings only when actually reading or writing the data. Translate to UTF-8 at the earliest opportunity, and translate from UTF-8 at the last possible moment, and only if you must. This isn't the 90s. This stuff is basically solved now, except people can't be bothered to use the solution.
- ohwp 14y agoLets say you could have earned $100 per hour instead of writing your own "parser". Then suddenly 0.5 days is $400.
- westi 14y agoTo be fair to the Posterous Team they are doing a good job of fixing the bugs in the export as they are reported to them. Hopefully they will get all of them fixed before the final close down. If you want an easy way to get your Posterous Export file cleaned up and into a more Valid XML file then feel free to use the Import from Posterous option over at WordPress.com - http://en.support.wordpress.com/import/import-from-posterous/ http://en.support.wordpress.com/import/import-from-posterous... We've spent some time on writing code which cleans up the XML file so that it can be imported into WordPress successfully. You can then export a clean WXR file and import elsewhere much easier - http://en.support.wordpress.com/export/ http://en.support.wordpress.com/export/
- gizzlon 14y agoHe has a few valid complaints (by a few I mean one), but this is really not that bad compared to a lot of the XML floating around. No reason to be shocked "There are no namespace declarations. No self-respecting XML parser will have anything to do with this XML data." I don't get this comment. I have never seen an XML parser that would refuse to parse XML without a namespace.. Am i missing something? Or is that just mindless hyperbole?
- masklinn 14y ago> Am i missing something? Or is that just mindless hyperbole? Note that the document uses namespaces but does not declare them. In Python, both ElementTree and LXML will blow up parsing when they encounter the first undeclared prefix (dc, from dc:creator)
- gizzlon 14y agoAh, you're right, I did miss something =) Still nothing to be "shocked" about though ..
- masklinn 14y agoTFA wasn't "shocked" (I suspect he was being slightly hyperbolic) at the sole invalidity-through-broken-namespacing, broken templating also had a hand in it: simply exporting a post and proof-reading the output is sufficient to catch the latter. Then again, you just have to put the output through any XML parser (it's not hard to find) to realize the document is completely broken, but...
- bambax 14y agoYou can't process XML that uses namespaces without a namespace declaration. A namespace prefix is just a shorthand for the namespace itself. prefix:name-of-element doesn't mean anything by itself, you need to know what 'prefix' stands for. As it is, this XML is not parsable; it's not well-formed and therefore it shouldn't even be called XML; it's just text with random tags thrown in. It is, indeed, quite shocking.
- nanoscopic 14y ago"There are no namespace declarations. No self-respecting XML parser will have anything to do with this XML data." I would argue that any self respecting xml parser should parser it just find and shouldn't demand the namespaces to be defined at all. "...invented their own ad hoc analogous namespace prefix, rather than understanding the role of elements in XML" I don't think you understand the base concept of XML much. It is meant to be a generic container to hold whatever you want. XML in and of itself doesn't enforce node naming. Sure if you are talking about the official spec it does, but people pretty much globally use whatever node names they want. Don't have a cow. "I haven’t been able to determine the intended encoding of the files" Well maybe you should look into a parser that just parses as is without attempting to use some specific encoding. Check out XML::Bare on cpan for perl. It will parse pretty much anything you throw at it, in any encoding. It leaves it up to you, the user, to decide what to do with the data after parsing.
- masklinn 14y ago> I would argue that any self respecting xml parser should parser it just find and shouldn't demand the namespaces to be defined at all. The XML Namespaces specification unambiguously requires that a namespace be declared: > The namespace prefix, unless it is xml or xmlns, MUST have been declared in a namespace declaration attribute in either the start-tag of the element where the prefix is used or in an ancestor element (i.e., an element in whose content the prefixed markup occurs). A self-respecting XML parser would follow the spec. A namespace-aware XML parser must fault on undeclared namespaces. Most XML parsers are namespace-aware. > I don't think you understand the base concept of XML much. Pot, meet kettle. > XML in and of itself doesn't enforce node naming. Sure if you are talking about the official spec it does Don't you feel like you're contradicting yourself a bit there? > Well maybe you should look into a parser that just parses as is without attempting to use some specific encoding. So he should look into parsers which do not parse XML and have no issue mangling the content? What are they going to do, assume the encoding is ascii-compatible anyway and go to town? How wonderfully anglo-centric. > Check out XML::Bare on cpan for perl. XML::Bare is an XML parser in the same sense that xhtml interpreted as text/html is an XML document: not in any way, shape or form. And if that's what you're shooting for, don't pretend to suggest an XML parser and suggest a recovering "soup" parser instead, something like html5lib or BeautifulSoup. But herein remains the issue: I expect Posterous advertised their export as XML files, not as "encoding-deficient tag soup" (which it apparently is). I'm sure TFA would have had no expectations if he'd been told he got garbage in, and would have relied on tagsoup-parsing and encoding-guessing (using whatever libraries for doing so are available in his language of choice). As it stands, he did have the pretty basic and undemanding expectation that he could shove supposedly-XML files into an XML parser and get data.
- stblack 14y agoI don't see any problem with this XML that can't be easily overcome. The comment about GMT-offsetting the date is particularly pithy, Assuming the blog in question isn't about ephemerides. By and large, blog posts have dates. If you desperately need an hour-offset from GMT, one might suggest this is your edge-case because, by and large, it doesn't matter. Count me among those who would argue that the omission of a schema is a blessing. I've wasted whole f*cking days of my life wrangling with so-called "non-amateur" XML. Invariably this was over-bloated XML with schemas that did nothing to help the discoverability and the processing of the data. Plain and simple, XML is over-spec'd and many data publishers, aided by their inflexible toolsets, pushed their XML beyond reason. Be careful what you wish for. I would take this XML, map-it, iterate it, done! End of story. I don't think there's much to complain about here.
- masklinn 14y ago> Count me among those who would argue that the omission of a schema is a blessing. TFA didn't ask for a schema, TFA asked for namespace declarations. Because they're kind-of necessary to parse namespaces with a namespace-aware XML parser. That's got 0 relation with a Schema. He only mentioned in passing because `content:encoded` and `excerpt:encoded` make very little sense... schema-wise (not "in an XML-Schema document"). > I would take this XML You can't "take this XML" because it's not XML. Once you know it's not XML you can "take this tag soup", shove it into a tagsoup library (maybe with some encoding-guessing beforehand) and hope things come out about right at the other end — with no insurance that this is the case, you're deep in GIGO land at this point — but you can't "take it and map it"
- obviouslygreen 14y agoAs someone who has used BeautifulSoup very happily without considering its etymology... is "tag soup" an actual term or just a very apt description you're using? [edit: A quick search, which I should have conducted instead of posting this, shows this has at least been used before, and enough not to be deleted by Wikipedia editors for lack of notability. That's pretty funny.]
- adamtaro 14y ago
- bazzargh 14y agoOne thing that bugs me about this is the use of CDATA. CDATA sections are just-about ok in hand-crafted xml, but in machine generated xml, they are absolutely pointless, and usually hint that the coder doesn't know what they're doing. For example, the author thinks that the content inside the CDATA is escaped, but in fact, it isn't necessarily - eg in this case they're including chunks of html which may contain more CDATA sections, and of course they don't nest (you need to terminate and restart the CDATA section). I've also seen examples where the enclosing encoding and the encoding of the CDATA section were incompatible. The worst thing is specs with CDATA sections in examples. Junior devs bend over backwards to use things like xsl's disable-output-escaping to get a character-for-character match in test results, and then wonder why their code breaks in production.
- gav 14y agoOutside of a few special cases (such as wanting to make embedded content in XML human editable) CDATA should be treated as a big warning flag that the author of the code that generated the XML doesn't really understand what they are doing. There's always the issue that one day ']]>' will somehow sneak in and everything will break. The key is using a tool to generate the XML that will transparently handle things like escaping correctly instead of using templating tools designed for text or HTML output.
- mnarayan01 14y agoI'm not sure making "XML human editable" should really be considered a special case.
- daGrevis 14y agoProbably they used regexes to parse it. :)
- _kst_ 14y agohttp://stackoverflow.com/a/1732454/827263 http://stackoverflow.com/a/1732454/827263 for those who haven't seen it.
- niggler 14y agoIt's ironic how many problems (large and irritating enough to justify blog posts or public spates) could have been avoided if someone bothered to test beforehand. If someone did a trial export he would immediately see the missing dates.
- icedchai 14y agoYes, it's crap, but it would take a few minutes to clean this up with a couple of sed scripts to turn ns:tag into ns_tag or something to make it parseable. Or you could prepend some fake namespace declarations.
- Sami_Lehtinen 14y agoSo it seems that we prefer XML which is easy to read. I have seen those files way often. Like: <xml><item><key>1</key><value>Something</value></item><item....></xml> Then you have to combine what ever keys and values are in item tags. I found out these to be very annoying files to handle. Especially when key is X3 and value is 83d, you have to look for every combination from some kind of mapping, because non of those tells you absolutely nothing directly. At least its easy to create files that full fill the schema, because the complexity is pushed out of XML level. Often these files are created by "upgrading" CSV to XML. Let's just call column key # and then put what ever is in that column to the value tag. Yes attributes could be used, but often aren't. Then you have to know that if key X contains value Y then you also need to look for key Z and hopefully it does contain value N or what ever.
- kaoD 14y agoWho uses XML in 2013 anyways?
- icebraining 14y agoAnyone who wants to interoperate with software not written in 2013?
- kaoD 14y agoShame on them.
- duaneb 14y agoWhat would you recommend to replace XML that handles arbitrary trees, namespaces, attributes, and tools that are built on this, e.g. XSLT? I don't think XML is amazing, but it still has its place.
- kaoD 14y agoPut your torches out, it's just a joke :)
- function_seven 14y agoSketchers (http://www.skechers.com/ http://www.skechers.com/). Go View Source on that.
- rrreese 14y agoIn 2013 XML is widely used. What alternatives would you suggest?
- peterkelly 14y agoThere's nothing wrong with invalid XML - why is everyone complaining? C compilers should similarly take a stab in the dark about what the programmer meant if they encounter invalid syntax as well. And those linking errors always annoy me - it should just pick the closest matching symbol if the specified one can't be found.
- p3d4nt 14y agoNo. Like this <rant> What the hell good is XML anyway, all that damn bloat and why, WHY? so I can put a style sheet in it? So I can nearly parse it. screw XML. </rant>
- LoneWolf 14y agoAm I the only one bothered by the extremely oversized xml snippets? Or is it just me? Chrome 25.0.1364.97 m
- adamtaro 14y agoIt's a pretty new redesign. I use Chrome myself, but shoot me a screenshot? hello at article_domain