5 ms·
Ah, the wonders of XML. Somehow somebody decided to use XML to store some text data. Because after all you already have a parser for this in the approved tool
by fafner 13y ago
Ah, the wonders of XML. Somehow somebody decided to use XML to store some text data. Because after all you already have a parser for this in the approved tools. So why bother using something different?! And look it does all the cool stuff like namespaces and validation and amazingly it can even fetch remote DTDs by loading IE...
- simias 13y agoI hate XML as much as the next guy but in this case I would blame a poorly designed (or maybe very misused) API. There's no reason why any file parsing library would end up fetching remote data without being explicitly asked to do so. Actually, it shouldn't even be the library's concern to fetch those files, an XML library has no business with networking. It's a security concern and a maintenance hell.
- acqq 13y ago> There's no reason why any file parsing library would end up fetching remote data It's not the API, it is the part of many of the specs, as it was thought to be a good idea once, specifically many standardizations involved additional definitions located at the http servers, example: http://msdn.microsoft.com/en-us/library/aa468557.aspx http://msdn.microsoft.com/en-us/library/aa468557.aspx Which was "poised to play a central role in the future of XML processing, especially in Web services where it serves as one of the fundamental pillars that higher levels of abstraction are built upon." So you had to implement it to be "conforming," and then to avoid overheads as "optimizations." Ironically, "/optimize" feature isn't optimized. The reason this it's not discovered earlier is that the Visual Studios which contain that option were priced more thousands of dollars (I don't know the what the currently cheapest version containing "/optimize" is -- anybody knows?).
- pwg 13y ago> There's no reason why any file parsing library would end up fetching remote data without being explicitly asked to do so. I would agree. But the XML spec. authors clearly disagreed with both of us: External Entities: XML 1.0: http://www.w3.org/TR/2008/REC-xml-20081126/#sec-external-ent http://www.w3.org/TR/2008/REC-xml-20081126/#sec-external-ent XML 1.1: http://www.w3.org/TR/2006/REC-xml11-20060816/#sec-external-ent http://www.w3.org/TR/2006/REC-xml11-20060816/#sec-external-e... This is "XML speak" for an "include" statement to something like cpp, with the exception that this "include" could end up performing remote network fetches to acquire that which is being included. So, technically, to be a proper, standards compliant XML parser, the parser has to at least submit requests to "fetch" these entities to the higher level code using the library, and let that code decide what to do about the "includes". As to why Microsoft's implementation is the way it is, absent a Raymond Chen blog post explaining the why, we can only guess.
- acqq 13y agoEver tried to look in the documents saved by Libre Office or MS Office? They all use XML now. The ODT document with only two words in it has at the start of the content.xml inside of ODT this beauty: <?xml version="1.0" encoding="UTF-8"?> <office:document-content (...) xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:math="http://www.w3.org/1998/Math/MathML" xmlns:ooo="http://openoffice.org/2004/office" xmlns:ooow="http://openoffice.org/2004/writer" xmlns:oooc="http://openoffice.org/2004/calc" xmlns:dom="http://www.w3.org/2001/xml-events" xmlns:xforms="http://www.w3.org/2002/xforms" xmlns:xsd="http://www.w3.org/2001/XMLSchema" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:rpt="http://openoffice.org/2005/report" xmlns:xhtml="http://www.w3.org/1999/xhtml" xmlns:grddl="http://www.w3.org/2003/g/data-view#" xmlns:officeooo="http://openoffice.org/2009/office" xmlns:tableooo="http://openoffice.org/2009/table" xmlns:drawooo="http://openoffice.org/2010/draw" It's not Microsoft specific bug that ate some brains.
- yuhong 13y agoThese are namespace directives, not references to external DTDs.
- acqq 13y agoThe problem is, whenever you have some link anywhere and it is assumed that it should be refreshed sometimes, how can you know that you shouldn't load the more current version? If you write something like a DLL or library why not leave it to the expert: let the IE try to fetch it, and if it already fetched, it will return it from its own cache! Brilliant, problem solved! Except when that happens from 144 instances all the time and IE needs some windows creations at the start which is what Bruce seems to manage to trigger.
- shawnz 13y agoTo expand on what the parent commenter was saying: namespace directives aren't meant to be accessed; they are just used as a unique identifier.