6 ms·
This was for fun (as he says.) > Remember: if you really need speed, do not use XML. He defines his own schema language, in XML (not XML Schema), which the pa
by 8ren 16y ago
This was for fun (as he says.)
> Remember: if you really need speed, do not use XML.
He defines his own schema language, in XML (not XML Schema), which the parser needs. Nothing wrong with that. You could even transform (some subset of) XSD's to it using XSLT. (see "2. defining the schema" http://tibleiz.net/asm-xml/tutorial.html http://tibleiz.net/asm-xml/tutorial.html and "defining the schema" http://tibleiz.net/asm-xml/documentation.html http://tibleiz.net/asm-xml/documentation.html)
He stores attributes in an array, to give O(1) access (... if you assume attribute order is significant, which it isn't according to the XML spec and most tools... even though humans find them more readable if in an expected order. So, to find a specific attribute, you'll have to step through them.)
XML parsers are unbelievably inefficient.. (but c'mon... it's XML... all other sins pall to insignificance before the big one, as fireflies at dawn.) He makes the nice point that once you have the XML parsed, you can just use those strings directly instead of copying them again.
IBM research came out with a this idea a while back, of combining the two, but it doesn't seem to have gone anywhere. I'd guess they patented it like crazy, because, after all, they are IBM. http://portal.acm.org/citation.cfm?id=1135777.1135796 http://portal.acm.org/citation.cfm?id=1135777.1135796 (that's just the abstract - sorry, couldn't find the pdf.)
But again, all this cleverness wouldn't matter except that it's fun (if you really need speed, do not use XML.). Oh... and except in "XML appliances": http://en.wikipedia.org/wiki/XML_appliance http://en.wikipedia.org/wiki/XML_appliance
DISCLAIMER: despite my tone, I am pro-XML. It's very useful. Some of its limitations facilitate some parsing ideas that escaped everyone else because of their cost in terms of performance and expressiveness. Just as some people are so obsessed with performance that they miss modular elegance, some people are analogously obsessed with expressiveness. There! I think I've insulted everybody.
- marcinw 16y agoOh my, that IBM research probably went into their DataPower appliances >> http://www-01.ibm.com/software/integration/datapower/performance.html http://www-01.ibm.com/software/integration/datapower/perform... My biggest gripe with XML is namespaces followed by XSLT. Working with XML at that point just leaves me feeling disgusted. edit: here's the full text >> http://www2006.org/programme/files/xhtml/5011/p5011-mendelsohn.html http://www2006.org/programme/files/xhtml/5011/p5011-mendelso...
- masklinn 16y ago> My biggest gripe with XML is namespaces followed by XSLT. XSLT I agree. Namespaces I quite like, the 3 issues I have being: * I can't write namespaced documents in Clark's notation. Clark's notation rocks. * Some XML tools/libraries/whatever manage not to understand namespaces correctly. * It's backwards-compatible with non-namespace-aware parsers. This is what makes XML namespaces broken and lets people avoid understanding them. They're not even hard to get, you just have to understand a namespaced name is a shortcut for a pair of (namespace-uri, local-name), and that a namespace is scoped to the node it's declared on. And you're set, you're done with namespaces. But instead, you have numbskull who tell you ElementTree 1.2 is broken because it has a very good handling of namespaces but doesn't let you configure a default namespace or customize namespace aliases on serialization (it just sets them as ``ns\d``). Oh yeah, now that I think about it there is one completely broken thing about XML namespaces (as far as I'm concerned): an element with no specified namespace lives in the current default namespace (``xmlns=uri``), but an attribute with no specified namespace lives min the "null" namespace instead. That is stupid and annoying.
- 8ren 16y agoNamespaces, as an abstract idea, are great. But I find XML's version confusing in practice. It's partly the several different ways of configuring defaults (four I think), and partly the syntax. You'd think they'd be as easy as Java: set-and-forget defaults (eg. java.util.*); or explicitly qualify everything (in import or in use). In practice, collisions are rare. Is it harder in XML because arbitrary XML documents are commonly nested? (which can't happen to Java source code)
- masklinn 16y ago> It's partly the several different ways of configuring defaults (four I think) Which ones? I only know of using the xmlns attribute to set a default namespace, though as I mention above (in my formatting-broken-by-HN text) attributes playing by completely different namespace rules is... annoying, to say the least. > Is it harder in XML because arbitrary XML documents are commonly nested? That's about the only difference, and I don't think it matters much, in my opinion (nor do I think Java's namespaces are very good). And XML namespaces are actually unique, unless you do very stupid things you can't have namespace collisions though you can have namespace-alias collisions (and as I noted above, using Clark's notation is a great way to understand how things actually work)