19 ms·
JSON vs. XML
- Finnucane 3y agoWhat a pointless debate. I've worked with XML manuscript archives and I can be certain that if I'd had to do it in JSON I'd have killed myself.
- adamgordonbell 3y agoThis is about JSON being created or discovered and Doug struggling to convince people it was relevant when everyone was so bought in on XML. Are you saying you think JSON shouldn't exist and everyone should use XML for everything? Tooling around XML was certainly more established, but man there was a lot of complexity built up around it.
- daveslash 3y agoThe Complexity of XML reminds me of something from Adam Bosworth's ISCOC04 Talk [0]. To me, the big takeaway is that HTML succeeded because of it's limitations, not despite of them. JSON seems very simple compared to XML. XML seems to be very powerful, but also very complex - it's like, if all you need to do is pick your kids up from Soccer Practice, you don't need the powerfullness (complexity) of the Space Shuttle in your vehicle. In 1996 I was at some of the initial XML meetings. The participants� anger at HTML for �corrupting� content with layout was intense. Some of the initial backers of XML were frustrated SGML folks who wanted a better cleaner world in which data was pristinely separated from presentation. In short, they disliked one of the great success stories of software history, one that succeeded because of its limitations, not despite them. I very much doubt that an HTML that had initially shipped as a clean layered set of content XML, Layout rules – XSLT, and Formatting- CSS) would have had anything like the explosive uptake. https://adambosworth.net/2004/11/18/iscoc04-talk/ https://adambosworth.net/2004/11/18/iscoc04-talk/
- HarHarVeryFunny 3y agoBut you don't have to use any more of the XML-related standards than you want to. You can ignore schemas, and add-on technologies like XPATH and XSLT and just use XML as a hierarchical tag-value format, just like JSON. At this level they are both about equal in complexity: JSON has data types that XML doesn't, and XML has attributes and CDATA that JSON doesn't. JSON syntax is more succinct, but XML syntax is more regular.
- rhdunn 3y agoXML is good for documents that don't have a regular markup (XHTML, DocBook, JATS, MathML, etc.) where you can mix content elements -- e.g. italic annotations. JSON is good for structured data/records such as serialized data structures found in RPC protocols. They both have their own pros and cons that make them suited to different use cases. Choose the one that best suites your data model and use cases.
- bayindirh 3y agoNo. JSON is great as Javascript's serialization format, but it's not as readable and robust as XML, period. I use both extensively, and for bigger objects and definitions, XML is a very clear winner. I'm a big believer in horses for courses type of approach, and my personal gripe is the push to replace one thing with another. These data types can coexist, and can be used where they shine. XML can be read and written stupidly fast, so it's way better as a on disk file format if people gonna touch that file. YAML and JSON are not the best fit for configuration files. JSON is good as an on-disk serialization format if humans not gonna touch that. XML is the best format for carrying complex and big data around. TOML is the best format for human readable, human editable config files.
- sirwhinesalot 3y agoXML is great except at being a configuration format, a messaging format, a serialization format, or any other purpose really. It's not insane like YAML I'll give it that. I'll take XML over that garbage any day.
- tehbeard 3y agoWhat specifically about XML means it can be read/written "stupidly fast"? It's still a text bound serialization format, you still have to parse a tree for it. Is it just particularly mature libraries?
- bayindirh 3y agoIt is primarily mature libraries, but also XML is more straightforward to parse, because there are not many data types and tags makes it very deterministic. By "stupidly fast", I mean I can read a 120K XML file, parse it, create the objects which generated from that file definition under 2ms. The library I use (RapidXML [0]) can parse the file almost with the same time cost of running strlen() on the same file. That's insane. [0]: https://rapidxml.sourceforge.net/ https://rapidxml.sourceforge.net/
- dralley 3y agoBeing a maintainer of the fastest XML library for Rust, I strongly disagree that XML is inherently fast to parse, and I question any such claim which comes with no evidence. Especially when it has remained unchanged on their page since (at least) 2008 [0]. Have you actually tested that claim or are you taking it at face value? IME the XML spec is so complex that you either end up with a slow but compliant parser or a fast one that doesn't implement the spec completely. JSON, unlike XML, is minimal enough that writing an entire compliant parser with SIMD intrinsics [1] is actually practically feasible. That library claims 3 GBps parsing speed, which could theoretically process your 120kb of data in 1/25000th of a second instead of 2/1000ths of a second. I would wager that JSON is faster to parse, on balance. [0] https://web.archive.org/web/20080209172554/https://rapidxml.sourceforge.net/ https://web.archive.org/web/20080209172554/https://rapidxml.... [1] https://github.com/simdjson/simdjson https://github.com/simdjson/simdjson
- irrational 3y agoDebate? Did you even read the article? It was about the history of how JSON came around. I didn't read a debate (despite what the title implies).
- adamgordonbell 3y agoThis quote is funny: Douglas: The first time I saw JavaScript when it was first announced in 1995, I thought it was the stupidest thing I’d ever seen. And partly why I thought that was because they were lying about what it was. A bigger more interesting thing though is how his company failed, in part, because they used hand-rolled JSON for messaging. Douglas: And some of our customers were confused and said, “Well, where’s the enormous tool stack that you need in order to manage all of that?” “There isn’t one, because it’s not necessary”, and they just could not understand that. They assumed there wasn’t one because we hadn’t gotten around to writing it. They couldn’t accept that it wasn’t necessary. Adam: It’s like you had an electric car and they were like, “Well, where do we put the gas in?” Douglas: It was very much like that, very much like that. There were some people who said, “Oh, we just committed to XML, sorry, we can’t do anything that isn’t XML.” I started my career during peak XML crazy and while I liked parts of it at the time, the number of things it was used for was quite insane. I had to maintain a system once where a major part of it was XSLT, when could have just been a simple imperative algo with some config settings. Anyhow, hope you like the episode!
- justin66 3y agoI really respect that you provide transcripts. It's terribly important for accessibility and for getting what people have said into the various search engines. I was sad to hear that Crockford is not aiming to be the author of "the next language" anymore, but I wonder how sincere that really is. His thoughts on actor-based languages are interesting.
- adamgordonbell 3y agoThanks! Crockford's thoughts on actors are really interesting. I tried to pull them apart but I didn't get very far and ended up not including them in the episode. What he is envisioning is not exactly like Erlang but not exactly like Scheme. He said that Carl Hewitt had a lot of ideas and they were hard to unpack. If you're interested though, I would reach out to him. He is very approachable and excited to talk to people with ideas for new ways of making things simple.
- hk1337 3y agoHonestly, I would relegate XML to application configuration. Trying to communicate with it with something like HTTP requests/responses is absurd.
- bayindirh 3y agoWhen you can get the data as XML, verify via its schema externally and then transform it via XSLT, it's not. Also, it's way better in transferring/storing big, complex intricate data like 3D objects.
- mattacular 3y ago> Also, it's way better in transferring/storing big, complex intricate data like 3D objects. Curious how come?
- bayindirh 3y agoI have a project which can work on 3D objects imported from STL files. My file format has more metadata and much more detailed information w.r.t. a standard STL file, and all the data and the metadata can be written in a way which is both readable and modifiable by a human if needed be. Having the same tags many times means the file can be nicely compressed, it's being XML means it can be verified independently with a schema (and the schema can be defined as a remote location over HTTP if needed be), too. You can always store data more efficiently with binary formats, but XML DOM parsers allows to access arbitrary parts of the tree instantly, so working with it is both easy and fast at the same time.
- mattacular 3y agoThank you!
- Kuyawa 3y agoI remember playing with an invoice data in XML sending it via email and opening it in a browser being beautifully shown in all it's visual glory using an XSLT directive in just one line at the top of the data file. Absolutely amazing. I wondered all the implications and applications of just transmitting data that knew how to present itself to the user
- user3939382 3y agoIf it's not obvious, the issue is that standardizing a data format is going to have trade offs. Interoperability, leveraging tooling universally so all effort is going in the same direction, awesome. The problem is that some uses cases for the format are going to be insanely complex, which will make the standard and tools unnecessarily complex for the simple cases. JSON is simpler and easier for many cases, but then you lose the interoperability. Go try to make an app right now dealing with Federal government systems or finance, you're going to end up translating JSON<->XML which isn't fun. There's not going to be a silver bullet solution to this problem, it's not completely solvable.
- Sohcahtoa82 3y ago> you're going to end up translating JSON<->XML which isn't fun. Not fun? It's not even possible in the general sense. If you have XML that looks like: <meal type="breakfast"> <eggs count="3"> <topping>cheese</topping> </eggs> </meal> How would you convert that to JSON without knowing how the JSON consuming application expects it to be formatted? Where do you put the "breakfast" and "count" attributes? You'd need to manually write a translator for each potential translation.
- user3939382 3y ago> You'd need to manually write a translator Yep, therein lies the “not fun”. You write a bunch of super complex, brittle code. Unfortunately because XML is entrenched in certain domains, you have to decide between writing these converters or doing everything in XML which also sucks, especially if you’re trying to write a modern app with a modern stack.
- dvh 3y agoFor me 3 killer features of JSON are: 1. Parsing JSON doesn't require adding new firewall rules 2. There are no comments, so nobody will try to invent their own meta format or annotations in comments and instead they will put data in the JSON as they should 3. (When compared to JS) someone finally had the balls and picked one type of quotes, this makes making parser so much simpler.
- oblio 3y agoJSON could really use schemas as part of the main implementation. Local schemas, not crazy remote schemas. Or some sort of way to bless an "official" schema format.
- Ygg2 3y ago> There are no comments, so nobody will try to invent their own meta format or annotations in comments and instead they will put data in the JSON as they should It also means it's worse format for configs where you sometimes need to annotate a few nodes with comments.
- lttlrck 3y agoYep. "comment": gets littered across the JSON... or temporary changes are copied and the original property name is invalidate with a prefix. The simple structure is gone, replaced with adhoc workarounds. Similarly when you want to use a type not supported by JSON such as datetime or binary data, you might end up with "type":"binary" and use base64 or whatever in the value (shoehorning attribs) - when it really needs a schema to follow during parse and stringify. Or OpenAPI, which is hardly lightweight and really doesn't match the simplicity of JSON.
- hiccuphippo 3y agoFor configuration files I add an extra key "_" and a string with the comment as the value. I even add multiple "_" keys to the same object and never seen something break.
- IshKebab 3y ago
- sanitycheck 3y agoI have huge respect for Doug Crockford, and I never imagined I would disagree with him. However I think by now we've seen that a lot of that "unnecessary" XML complexity was not, in fact, entirely unnecessary. These days we use JSON for everything, but now we've got JSON Schema, Swagger/OpenAPI, Zod, etc etc. It's not really simpler and there's a lot of manual work - we might as well be using XML, XSD & SOAP/WSDL.
- adamc 3y agoJSON is great, but I surely wish it supported comments. That's the nature of its failings: too minimal.
- beached_whale 3y agoLuckily a good number of parsers support extensions to JSON like comments and trailing comma's.
- wongarsu 3y agoThat depends on what you want it to be. For a data interchange format, having no comments is arguably a strength. For a config file format, having no comments is a big weakness.
- sgtnoodle 3y agoMany people use text format protobufs for config. It vaguely resembles yaml, and the schema is enforced by virtue of requiring a message definition in order to parse it.
- hot_gril 3y agoJust do { "someSetting": true "comment": "TODO change to false when ready" } Though really text-based protobufs are better for config.
- sixbrx 3y agoProblems are that some tools will rewrite the file and reorder the "comment" away from what it's meant to comment on. Also might complain the "comment" item isn't expected there. I seem to remember package.json suffering from both of these under the control of npm.
- mproud 3y agoTo me, Douglas Crockford is the unofficial grandfather of JS. He is amazing, and I love hearing him speak!
- HideousKojima 3y agoWhat if I hate both formats? XML is overly verbose, while JSON isn't specific enough or precise enough for a lot of my needs. - This message brought to you by TOML gang
- dralley 3y agohttps://github.com/ron-rs/ron https://github.com/ron-rs/ron
- rewgs 3y agoAbsolutely agree. TOML is far and away the best for config files.
- mdaniel 3y agoYou can't be serious; I can't stand having to guess what kind of crazy markup is required to express things in toml. As a concrete example I converted my local kubeconfig (which is yaml) and here are the completely random characters indicating some kind of hierarchy apiVersion = "v1" current-context = "" kind = "Config" [[clusters]] name = "my-cluster" [clusters.cluster] certificate-authority-data = "LS0tL..." server = "https://example.com" [[contexts]] name = "context0" [contexts.context] cluster = "my-cluster" user = "my-user" [[contexts]] name = "context1" [contexts.context] cluster = "my-cluster" user = "my-user" [[users]] name = "my-user" [users.user] [users.user.exec] apiVersion = "client.authentication.k8s.io/v1beta1" args = ["eks", "get-token"] command = "aws"
- jsmith45 3y agoIt is really just arrays of objects/dicts that get complicated with TOML. For dictionary properties not inside an array, you can either fully specify the path `foor.bar.baz = 7`, or use a header like `[foo.bar]` and specify `baz=7`. Also if I was handwriting that I would probably make more use of doted property names implying dictionaries like so, which though it has a little bit more repetition in property names, seems easier to read: apiVersion = "v1" current-context = "" kind = "Config" [[clusters]] name = "my-cluster" cluster.certificate-authority-data = "LS0tL..." cluster.server = "https://example.com" [[contexts]] name = "context0" context.cluster = "my-cluster" context.user = "my-user" [[contexts]] name = "context1" context.cluster = "my-cluster" context.user = "my-user" [[users]] name = "my-user" user.exec.apiVersion = "client.authentication.k8s.io/v1beta1" user.exec.args = ["eks", "get-token"] user.exec.command = "aws" If k8s was designed with TOML in mind, it probably have been structured differently, such that "Contexts" for example might be just a dictionary mapping names to an object that has the values from the "context" property (The existing pattern of an array of objects where each object has a name, but store most of their properties in a property whose name matches the object's type is already weird, but doesn't look terrible in yaml.) Such a redesigned to be a more TOML friendly schema would then look like this: apiVersion = "v1" current-context = "" kind = "Config" [clusters.my-cluster] certificate-authority-data = "LS0tL..." server = "https://example.com" [contexts.context0] cluster = "my-cluster" user = "my-user" [contexts.context1] cluster = "my-cluster" user = "my-user" [users.my-user.exec] apiVersion = "client.authentication.k8s.io/v1beta1" args = ["eks", "get-token"] command = "aws"
- simonw 3y agoMy favourite Douglas Crockford quite, from a debate back in 2006 about why JSON was reinventing the wheel when XML already existed: > The good thing about reinventing the wheel is that you can get a round one. https://simonwillison.net/2006/Dec/21/crock/ https://simonwillison.net/2006/Dec/21/crock/
- Devasta 3y agoNamespaces, schemas, custom elements, client side templating, XML has so much stuff that the web threw away, so now its forced to reinvent worse versions of it every few years, a shame. Abandoning XML was the webs biggest mistake.
- Zamicol 3y agoXML is oversized for the majority use case. It's easier to extend a simple standard than to amputate a behemoth with unneeded appendages.
- recursive 3y agoThe problem with extending a standard is that there are so many ways to do it.
- giantrobot 3y agoWhenever XML gets discussed here it's interesting to see what people complain about. In my completely unscientific assessment most things people hate(d) were the overwrought "Enterprise" uses/systems. Very unfortunately for everyone XML came up at the same time as peak "Enterprise" moat building. No design pattern went unused everything was built with mind numbing "configuration". XML got used heavily in that space because it allowed massive "Enterprise Objects" (local branding varies) to be serialized in a way another system might have a chance to read. Meanwhile the features you mention got thrown out with the bath water because everyone hated Enterprise style architectures. While I don't love, for instance, everything about XSLT it's built directly into browsers as native code. How many person hours, megabytes of JavaScript, and wasted CPU cycles have been spent reinventing client side templating using JSON? XSLT is already right there and will happily convert serialized data to your presentation format. You also get the ability to have comments in the data and a built in schema validation. On my current project I'd much weather be emitting and consuming XML rather than JSON. But alas everyone hated Enterprise XML so we're stuck with JSON and the inability of some parsers to handle trailing commas and ambiguous definitions of numerics and not a comment to be found.
- Zamicol 3y agoFrom another interview: >The best thing we can do today to JavaScript is to retire it. Twenty years ago, I was one of the few advocates for JavaScript. Its cobbling together of nested functions and dynamic objects was brilliant. I spent a decade trying to correct its flaws. I had a minor success with ES5. But since then, there has been strong interest in further bloating the language instead of making it better. So JavaScript, like the other dinosaur languages, has become a barrier to progress. We should be focused on the next language, which should look more like E than like JavaScript. - https://evrone.com/douglas-crockford-interview https://evrone.com/douglas-crockford-interview One of the traits that makes Douglas great is being willing to say the obvious even if it is politically unpopular.
- butt_____hugger 3y ago[dead]
- irrational 3y agoThe biggest impedances I see to replacing JS are: 1. You've got to keep JS around for backwards compatibility for the billions of websites already using it. 2. You will need to two engine teams, one to maintain JS and one for the new language. 3. Now you have a whole new vector for security issues. You've made the threat surface much broader. So, you will probably need to hire additional people. 4. You need to coordinate with all the other browser makers so everyone rolls out their new engines more or less concurrently. Other than experiments, nobody is going to start using it unless it works on all the major browsers and platforms.
- hajile 3y agoThat depends on the language you choose. If we went to a scheme dialect as originally intended, we could have just ONE language for all the things. Legacy JS? Just compile it into Scheme and run it. HTML? Use S-expressions and support legacy HTML syntax by compiling it into them. Now you get all the power people want from template languages, but baked right into main language itself. CSS? No more weirdness like adding sin() or calc() to make up for shortcomings. Once again, you get the power of the full Scheme language right there.
- taeric 3y ago> Turned out JavaScript was the first language to give us lambdas, and that was an amazing breakthrough. I mean... with charity I can see the context and get it. But. What!? Overall fun read through history, even if definitely from Doug's perspective only. (As evidence by JavaScript being an originator of lambdas...) I do find the idea that JSON was as novel as history says it was kind of odd. I remember inlining javascript objects years before "JSON" was a thing. Making it a subset of what javascript could already do seems straight forward and a good execution. Getting rid of comments feels asinine to me. (I'll also note that the plethora of behaviors you get from JSON parsers shows that it is effectively CSV. Sure, there may be a "standard" out there, but by and large it is a duck typed one.) I'm also a bit on the camp that XML is better than JSON. Being able to have better datatypes, for a start. Schemas that allow autocompletion. Is also easier to see as a markup language (per the name). That said, they clearly went too far with entities and despite making sense for markup, attributes versus children are more than a touch awkward. I also recall that what killed XML and WSDL files in general, was the complete shit show that was getting a single document to work with both MS and non-MS clients.
- slaymaker1907 3y agoThe current XML standard is hot garbage since it completely disallows null characters even via "