19 ms·
Introducing TJSON, a stricter, typed form of JSON
- mnarayan01 10y ago> All base64url strings in TJSON MUST NOT include any padding with the '=' character. This seems like it makes a streaming parser's job (slightly) more of a headache, without any serious advantage. Which seems particularly odd to me given that this seems heavily focused on binary stuff.
- bascule 10y agoPadding is redundant when base64url is encapsulated in a quoted string. If you're writing a state machine-based parser which is processing a quoted base64url it will, in amortized time, be able to find a close quote token faster than it will be able to find valid close padding.
- pg_is_a_butt 10y agois the author isn't japanese, this is racist. SWEEP THE LEG! you're all idiots.
- marianoguerra 10y agoisn't there a way to extend the types to specify our own and register constructors for them? like transit? otherwise we will be in the same place of json in terms of extension where our own types are second class citizens.
- cdmckay 10y agoAgreed. Just adding some fixed types doesn't really help that much. Something like EDN for JSON would be cool: https://github.com/edn-format/edn https://github.com/edn-format/edn
- fnordsensei 10y agoIsn't Transit basically EDN for JSON in that it adds types and whatnot, and encodes to JSON? Or do you mean, you want a format that's sort of halfway between EDN and JSON?
- marianoguerra 10y agotransit works great except that it's unreadable with current tools (for example browser devtools or attaching listeners to kafka). I know it's a tool problem but I don't see the whole world embracing transit. If this format gets adopted with extensible types we get a readable format that has what transit provides and if there's no tooling support we can still read it with standard json tools or none at all.
- unlogic 10y agoTransit is unreadable exactly because it has to work around the limitations of JSON (like string-only keys) to deliver its primary features: true maps, tagged collections etc. TJSON only has tags for primitives, so yeah, it's not much different from JSON this way, the tooling is happy.
- bascule 10y ago> Just adding some fixed types doesn't really help that much. It brings the set of scalar types you can express in a JSON message on par with other serialization formats like Protobufs: https://developers.google.com/protocol-buffers/docs/proto3 https://developers.google.com/protocol-buffers/docs/proto3
- kccqzy 10y agoThe problem I see is that everyone has their own favorite type systems. Functional people may consider sum types (tagged unions) indispensable, while OOP people might want their types to have notions of inheritance. Another functional programmer might want existential quantification, higher-kinded types that most people outside the functional niche have never heard of, but a lisp programmer might want actual code as data (quote/eval) so the type has to involve functions, etc. Extending the types beyond the basic primitives is difficult because there are so many different ways of doing that.
- marianoguerra 10y agoit's not about specifying a type system, just letting users specify a tag for a type and then register a constructor for that tag, then inside it you can have whatever type system thing you like, the serialization format doesn't care, for example {"#Some:myOption": "s:value"}, the decoder will call the constructor registered for Some passing the value and not care about your type system.
- tiglionabbit 10y agoWe could just write a JSON Schema for it. It allows you to specify a "format": http://json-schema.org/latest/json-schema-validation.html#anchor104 http://json-schema.org/latest/json-schema-validation.html#an... So you can write a schema like: {"type": "string", "format": "email"} or: {"type": "integer", "format": "uint64"} There's no spec for what is allowed as a "format", so you have to decide on your own values and write your own validators, but someone could come up with a standard spec for this. Swagger formally specifies some values of "format" in this document: http://swagger.io/specification/ http://swagger.io/specification/
- bascule 10y agoThe purpose of TJSON is to be self-identifying and schema-free. If you want a schema, use Protobufs or the myriad JSON schema languages.
- marianoguerra 10y agoI don't want a schema, I want to preserve types between serialization and deserialization thus avoiding conventions or having to specify those types "out of band", the same way you want to make it clear that an int is an int and a date is a date, I want to tag an object to tell that that object is a city, a person or something else, each program should register a function to rebuild the actual object but at least it's not a convention anymore.
- kr0 10y agoWhy don't float types use a tagged string? It says "tagging is mandatory" in the initial document, but floating point types are then omitted in the official spec
- jerf 10y agoFloating point types are tagged by the use of the floating point grammar. It would require the standard to be clear that the only way to indicate integers is via "i:288", though, or there will be ambiguity. I don't know if that circle can be squared, either; if you require integers to use the tagged string, it isn't really backwards compatible any more. If you don't, the floats remain ambiguous. Given that the text of the blog post suggests, probably correctly, that new parsers will be necessary to use this format, I'm not convinced that trying to reuse JSON's grammar is that advantageous. If I'm switching parsers, the competition is no longer JSON, it's the full range of possible replacements, including Protocol Buffers, Cap'n Proto, XML, BSON, and everything else. If you're willing to replace parsers there's probably already something out there for you.
- bascule 10y agoIt would require the standard to be clear that the only way to indicate integers is via "i:288", though, or there will be ambiguity. The spec does this here: https://www.tjson.org/spec/#rfc.section.4.3 https://www.tjson.org/spec/#rfc.section.4.3 4.3. Floating Points All numeric literals which are not represented as tagged strings MUST be treated as floating points under TJSON. This is already the default behavior of many JSON libraries. If I'm switching parsers, the competition is no longer JSON, it's the full range of possible replacements, including Protocol Buffers, Cap'n Proto, XML, BSON, and everything else. As noted in the post (which names a similar list of binary formats), TJSON is intended to be supplemental to binary formats, not a "replacement"
- jerf 10y agoThank you. I skimmed over that accidentally. Good.
- deleted 10y ago
- drawkbox 10y agoWhy muddy up the actual values where you will have to parse that value with "t:" where t is type? Why stuff it in one key/val? why not separated where it looks to see if type is present, if so it converts to it/validates against it (you can also place other validations/constraints on it like min/max values, length etc -- that will fall apart if you are trying to stuff it all in one key/value). Like this: { "val":"Hello, world!", "type":"string", "validation": "[regex]" } Instead of: { "s:string":"s:Hello, world!" } This is typically how we type fields in JSON when needed as there is no parsing needed on the value. If you need to check type and it is present you can act on it.
- jerf 10y agoStoring validation next to the type like that is a bad idea in general. If you can't trust the incoming data to be valid, then for the same reasons, you can't trust the incoming data's claim for what would make it valid.
- drawkbox 10y agoPossibly out in the wild but if it comes from a server you control and systems you validate then both this and TJSON or any JSON type system would have that same issue. Typically typing/schemas are system to system and not necessarily filled by users or in areas they can be edited. Same issue with XML validation, any schema info needs to be enforced by the server/backend/api.
- wojcikstefan 10y agoThat's a lot of extra bytes you have to send over the wire. Also, I don't think validation makes sense. When sent by the server, it's too limited (would lead to situations where you're doing half the validation in TJSON and half in the client code). When sent by the client, it can't be trusted anyway.
- drawkbox 10y agoTrue if validation is on there. I just put it in to show you could have other easily added validations aside from type (TJSON is locked to just type as it is concat/mashed in one value colon separated). If you just take the "val" and "type" it is really no extra bytes or very minimal but cleaner. { "val":"Hello World", "type":"string" } OR { "s:string":"s:Hello, world!" } Pretty much the same. I guess my personal preference is I don't like to mash values and parse values out of key/value values. In the end all validation is done on the server anyways so types/schemas for JSON are really just a nice to have and should not be relied on unless you control both ends of the pipe.
- lillesvin 10y agoI feel like there's a missed opportunity in not calling it TySON or something like that. That aside, wouldn't it make more sense to fix the JSON parsers instead? They are the ones having issues parsing e.g. 64 bit integers, JSON has no problem holding them.
- Kinnard 10y agoYes! Name should totally be changed. I hope they see this.
- rpedela 10y agoI was confused by the claim that JSON parsers do not handle 64-bit integers. If the parser is written in Javascript, then it has a problem because Javascript does not support 64-bit integers. But I have not seen that problem in any other language. For example, Postgres's JSON parser can handle whatever the maximum size of PG numeric is and Python can handle extremely large numbers as well.
- bascule 10y agoFrom RFC 7159 section 6. Numbers: https://tools.ietf.org/html/rfc7159#section-6 https://tools.ietf.org/html/rfc7159#section-6 Note that when such software is used, numbers that are integers and are in the range [-(2**53)+1, (2**53)-1] are interoperable in the sense that implementations will agree exactly on their numeric values. You can't depend on interoperable support for 64-bit integers in JSON. Furthermore many JSON libraries convert all numbers to floats, so this problem doesn't affect only JavaScript. TJSON requires conforming parsers to support the full 64-bit signed and unsigned ranges. This will involve using bignums in JavaScript.
- bpicolo 10y agoI'm still waiting on xml with curly braces instead of angle brackets. As far as I can tell that's all that's holding us back
- falcolas 10y agoYup, we already have schema validation, JSONRPC, and transformations, all that's really missing is namespaces and comments. Then we can go full WSDL and SOAP.
- bpicolo 10y agoDon't worry, we have comments https://hjson.org/ https://hjson.org/ and our scientists are hard at work on namespaces: http://www.goland.org/jsonnamespace/ http://www.goland.org/jsonnamespace/
- mianos 10y agoI find this comment extremely offensive. Microsoft is going to fix the namespaces in their SOAP for .net real soon now. In the meantime all you have to do is put a few patches in your non .net SOAP code to deal with the badly formed namespaces. Besides, those 20 patches have only been needed for ten years now.
- andrewf 10y agoI'm looking forward to JSON-I becoming the recommended intersection of standards to interoperate with others.
- msoad 10y agoIt's amazing how many people are trying to reinvent protocol buffers! Every time I see something like this I think the developer didn't do their research or maybe they wanted to make a hobby project anyway. Stuff like this is dangerous to use in production. Even JSON as simple as it looks had a lot of bugs that are now. If you want typed data structure transfer, use protocol buffer.
- bascule 10y agoDid you even read the post? First paragraph: "Its primary intended use is in cryptographic authentication contexts, particularly ones where JSON is used as a human-friendly alternative representation of data in a system which otherwise works natively in a binary format." See also the "Content-Aware Hashing" section: The goal of this format is to enable content-aware hashing which produces the same digest for data encoded either as TJSON or a binary format such as Protobufs. I am using it in conjunction with Protobufs.
- mianos 10y agoSo they are inventing ASN.1 again, for the third time. Next thing they will invent distinguished encoding rules so data can be hashed without decoding.
- bascule 10y agoNo, far from "inventing ASN.1 again", TJSON could potentially be a very useful format for representing equivalent structures to ASN.1, similar to: https://github.com/google/der-ascii https://github.com/google/der-ascii
- dewitt 10y agoI take it you don't know who either of the two authors are? - https://en.wikipedia.org/wiki/Ben_Laurie https://en.wikipedia.org/wiki/Ben_Laurie - https://github.com/tarcieri https://github.com/tarcieri They know about protocol buffers.
- 10y ago
- rubyfan 10y agoIt's funny that people expect JSON is a serious data interchange format.
- matt_wulfeck 10y agoBased on ubiquity do you disagree?
- rubyfan 10y agoYes. Ubiquity doesn't mean it is a fit for purpose. Especially for most things these extensions try to overcome. In this case ubiquity is largely a product of the primary consumer of much JSON data is a web browser. Likely much of that data is simple enough that it does not require more than what JSON provides.
- bpicolo 10y agoUbiquity is a very useful property of interchange formats (APIs are a big deal). That said, for internal-only things where I have a lot of control (and I'm writing in supported languages - wtb elixir), I'd probably be using grpc
- drawkbox 10y agoThe next guy who inherits your internal code would prefer you to just use JSON. There is a reason it is ubiquitous, simplicity. I wonder how long though with all these type systems and XJSONs.
- bpicolo 10y agoMmm, disagree. It's not really the interchange format that's the only useful part of GRPC. (Though protobuf is pretty standardized these days). This was just the context of "I'm a big org standardizing on microservice tradeoffs", so maybe slightly out-of-context
- kevinSuttle 10y agoWhat about https://amznlabs.github.io/ion-docs/ https://amznlabs.github.io/ion-docs/ ?
- bascule 10y agoIon is a superset of JSON: not all Ion documents are valid JSON documents. TJSON can be viewed as a subset of JSON: all TJSON documents are valid JSON documents, and parsed by existing JSON parsers. Consuming TJSON documents as JSON will involve stripping the tags, but as noted in the post, people already do these sort of transformations on parsed JSON to e.g. extract binary data.
- Kinnard 10y agoReminds me of Tyre – Typed regular expressions: https://news.ycombinator.com/item?id=12292389 https://news.ycombinator.com/item?id=12292389
- bruth 10y agoAll of the keys in JSON must be strings, so they should not need tags for themselves. Instead why not put the tag of the value assigned to the key in the key: { "s:string":"Hello, world!", "b64:binary":"SGVsbG8sIHdvcmxk", "i:integer":42, "f:float":42.0, "t:timestamp":"2016-11-02T02:07:30Z" } This prevents having to mess with the values in general and integers don't need to be encoded as strings. EDIT: I see this constraint: Member names in TJSON must be distinct. The use of the same member name more than once in the same object is an error. which is still satisfied, however you could have `i:foo` and `s:foo` which would result in redundant keys in the resulting JSON document. This constraint could be clarified that, untagged key names must be unique. Another question, is a mimetype planned for this? `application/tjson`?
- nemothekid 10y agoI think you may still want to encode integers as strings anyways if you are encoding/decoding in Javascript.
- bascule 10y agoThat is not the case. Binary data is also allowed as the keys of objects (see https://www.tjson.org https://www.tjson.org or the spec). As noted in the "Content-Aware Hashing" section, an intended future feature is to support redaction, so tags on keys are needed to support this feature. Finally, if you were to do it that way I think it would make more sense to place the type tags on the values, not the keys, both visually and semantically.
- bruth 10y ago> That is not the case. Binary data is also allowed as the keys of objects What is the value of binary key? A key is just the name for a value, it should not contain any data itself. > I think it would make more sense to place the type tags on the values, not the keys, both visually and semantically. Tags on keys are like types for columns or any other schema. I would rather not have to pre-process the values. To be pedantic, this would require copying all string-based values just to add a prefix.
- zeveb 10y ago> Its primary intended use is in cryptographic authentication contexts, particularly ones where JSON is used as a human-friendly alternative representation of data in a system which otherwise works natively in a binary format. The author might care to take a look at canonical S-expressions, a format from the 90s which attempted to do the same thing for many of the same reasons, and has the advantage of being rather more elegant. E.g: { "s:string":"s:Hello, world!", "s:binary":"b64:SGVsbG8sIHdvcmxk", "s:integer":"i:42", "s:float":42.0, "s:timestamp":"t:2016-11-02T02:07:30Z" } could be: (string "Hello, world!" binary [b]|SGVsbG8sIHdvcmxk| integer [i]"42" float [f]"42.0" timestamp [t]"2016-11-02T02:07:30Z") Which is a perfectly valid encoding, but can use the canonical encoding (useful for cryptographic hashes): (6:string13:Hello, world!6:binary[1:b]13:Hello, world!7:integer[1:i]2:425:float[f]4:42.09:timestamp[1:t]20:2016-11-02T02:07:30Z) Which can be encoded for transport as: {KDY6c3RyaW5nMTM6SGVsbG8sIHdvcmxkITY6YmluYXJ5WzE6Yl0xMzpIZWxsbywgd29ybGQhNzpp bnRlZ2VyWzE6aV0yOjQyNTpmbG9hdFtmXTQ6NDIuMDk6dGltZXN0YW1wWzE6dF0yMDoyMDE2LTEx LTAyVDAyOjA3OjMwWik=} Granted, 'elegance' is in the eye of the beholder, but I like it. I also think that there's a deeper concern with any shallow notion of types. An application doesn't care so much about 'some integer' as it does about 'a valid integer for this domain,' and that concern is what leads to schemas and profiles and things like that. Just encoding the machine type of a value is insufficient: one has to encode the domain type, which means conveying the domain, which means assuming some sort of shared knowledge.
- bascule 10y agoS-expressions are great, and I'm a big fan of SPKI/SDSI, which used S-expressions in a security context. However, they have generally not gained favor in the greater programming ecosystem, whereas JSON has. TJSON is trying to tap into the greater ecosystem of people who are familiar with JSON to some extent. Hence its backwards compatibility with JSON, and not adding a backwards-incompatible type syntax, as Amazon Ion did.
- jasonkostempski 10y ago"underspecification has lead to a proliferation of interoperability problems and ambiguities." So TJSON has a perfect spec and everyone, now and forever, will interpret it perfectly?
- falcolas 10y agoHuh. Thought that's what XML with namespaces and schemas was supposed to do. Only being a little sarcastic...
- brazzledazzle 10y agoOn the other hand you're "better off with a diamond with a flaw than a pebble without". Perfect is the enemy of good and all that.
- bascule 10y agoNo, but it has a set of machine-readable examples which are intended to cover JSON's present underspecified edge cases: https://github.com/tjson/tjson-spec/blob/master/draft-tjson-examples.txt https://github.com/tjson/tjson-spec/blob/master/draft-tjson-...
- shitgoose 10y agojson became so popular in first place because of its simplicity, i.e. no schemas, namespaces, attributes, less bizarre notation than xml. let's keep it this way.
- bascule 10y agoTJSON doesn't add any of the things you just complained about
- shitgoose 10y agoit doesn't. instead, it takes it to new heights: "s:id":"i:11" this illustrates, what in my mind is main problem with contemporary software development. in old days, first, there was a problem, for which we had to find a tool that is good enough. nowdays there are plenty of tools, for which we are hoping someone will find a problem.
- tiglionabbit 10y agoWhen have you ever written a program that doesn't know ahead of time what type of data it's going to be operating on? Especially if you're using a statically typed language. Whether you validate incoming payloads in JSONSchema or not, you will always have some understanding of what the shape of the incoming JSON is supposed to be, down to the most concrete types. You'll probably receive many JSON payloads that all conform to the same schema. So why bother redundantly describing that schema in every individual payload? If you want strict types, write a JSONSchema. If you need to know specific sub-type information, start specifying what should go into the "format" field in JSONSchema. They did it in Swagger: http://swagger.io/specification/ http://swagger.io/specification/ Since the article complains about JSON parsers not knowing how to handle certain situations, perhaps people should start writing JSON parsers that allow you to pass in a JSONSchema document at parse time so they're sure to handle each field type correctly.
- 3pt14159 10y agoPlenty of times, either when I'm taking other people's JSON or when I'm coding for future-me or for arbitrary JSON traversal or search. But I agree that TJSON rubs me the wrong way. The simplicity of JSON is what I like and I can code around it when I need to.
- dvdhnt 10y ago^ this
- lobster_johnson 10y agoThere are many use cases where you don't know the shape of the data. Many apps need to index or store or transform arbitrary key/value pairs, but without knowing anything about those keys or values mean. JSON is a schemaless interchange format, so those situations arise pretty much by default. Not that I love this format -- fixing JSON needs a bit more effort, especially on syntax.
- colanderman 10y ago
- partycoder 10y agoCan you have a typed array too?
- bascule 10y agoTBD: https://github.com/tjson/tjson-spec/issues/23 https://github.com/tjson/tjson-spec/issues/23
- emmelaich 10y agoThose labels in the example are confusing. Instead of string, binary, integer, float,timestamp please use something like name, password, age, height, sessiontime. Using string and binary is worse than using foo and bar.
- gengkev 10y agoI'm a bit confused that TJSON only allows UTF-8 strings. The only way to escape Unicode characters in JSON is \uXXXX. But to encode astral characters with this syntax, UTF-16 surrogate pairs must be used. How does TJSON handle this, if strings must be encoded with UTF-8 only?
- rurban 10y agoJSON is defined to use surrogate pairs to encode these. TJSON must do nothing here. e.g. \ud8a4\uddd1 => U+391d1
- colanderman 10y agoSix things: 1) "Lack of full precision 64-bit integers" is bullshit. Numeric precision is not specified by JSON. If a parser can't deal with 64-bit integer values, it's a poor parser. 2) "s: UTF-8 string" What does this mean? JSON strings are strings of Unicode code points; JSON itself may be encoded as UTF-8, -16, or -32. So does this mean "encode the string as UTF-8, then represent as Unicode code points"? That makes no sense. Does this mean "encode the string as UTF-8 and output directly regardless of the encoding of the rest of the JSON output"? That makes no sense either. So I'm guessing the author just conflated "UTF-8" with "Unicode", which is concerning given that he is attempting to define an interchange protocol. 3) "i: signed integer (base 10, 64-bit range)" What does this mean? (-2^64,2^64)? (-2^63,2^63)? [-2^63,2^63)? 4) "t: timestamp (Z-normalized)" What does that mean? There are literally dozens of timestamp formats. Does he mean full ISO 8601, restricted to UTC? 5) What is the point of TJSON anyway? When you deserialize, you still have to check that the data is of the type you expect. At best this saves a bit of parsing, since the deserializer can do that automatically. Various JSON schema languages already exist, which give you this richer typechecking. The only use case I can think of for this is exactly what the author mentions further down the article: canonicalization for content-aware hashing. But this only works if the only types you care about fall into the small handful he thought of. What about, say, IP addresses? Case-insensitive strings (such as e-mail addresses)? 6) If we're talking about canonicalization, TJSON does not say how to canonicalize decimal numbers. I suppose this stems from the author's mistaken belief that numbers in JSON are IEEE floats (they're not, regardless of what common broken parsers do). I hate to be so negative, but this really comes off as half-baked. EDIT: Looking at the spec [1] it seems to address some of these, but still indicates a strong confusion between data types (Unicode, rational numeric) and data representations (UTF-8, IEEE double). [1] https://github.com/tjson/tjson-spec/blob/master/draft-tjson-spec.md https://github.com/tjson/tjson-spec/blob/master/draft-tjson-...
- bascule 10y ago1) Numeric precision of integers is (under)specified in RFC 7159 section 6: https://tools.ietf.org/html/rfc7159#section-6 https://tools.ietf.org/html/rfc7159#section-6 Note that when such software is used, numbers that are integers and are in the range [-(2**53)+1, (2**53)-1] are interoperable in the sense that implementations will agree exactly on their numeric values. There is no contract that JSON integers give you full 64-bit precision. TJSON has such a contract, and tests for support for full-precision 64-bit integers (and expected failure in the boundary cases) is specified in the canonical test cases/examples file: https://github.com/tjson/tjson-spec/blob/master/draft-tjson-examples.txt#L210 https://github.com/tjson/tjson-spec/blob/master/draft-tjson-... 2) Please see https://github.com/tjson/tjson-spec/issues/27 https://github.com/tjson/tjson-spec/issues/27 3) Yes, these specific ranges are covered in the spec: https://www.tjson.org/spec/#rfc.section.3.3 https://www.tjson.org/spec/#rfc.section.3.3 4) Z-normalized RFC3339. See: https://www.tjson.org/spec/#rfc.section.3.4 https://www.tjson.org/spec/#rfc.section.3.4 5) TJSON provides a repertoire of types which approximates what's available in the scalar types of a format like Protobufs: https://developers.google.com/protocol-buffers/docs/proto3#scalar https://developers.google.com/protocol-buffers/docs/proto3#s... What about, say, IP addresses Simple solution for that case: IP addresses have canonical representations as strings, so use their string representations. Or, if you prefer, represent them as a TJSON object. 6) objecthash provides an alternative to canonicalization: we can use a "content-aware" hash algorithm to produce a digest of the content rather than trying to arrange the content into a canonical form. See: https://github.com/tjson/tjson-spec/issues/24 https://github.com/tjson/tjson-spec/issues/24 I hate to be so negative, but this really comes off as half-baked. As far as I can tell, you didn't read the spec. All of your perceived ambiguities are addressed.
- DiabloD3 10y agoSo, why would I use this instead of actual JSON (== browser support), BSON (binary JSON), or Capn Proto (I control both ends of this)?
- jasoncchild 10y agoDoes a time zone key trigger the enforcement a specific ISO standard format for the value?
- rxbudian 10y agoWhy not just have a separate metadata file. It will keep the json file lean.
- novaleaf 10y agoand still no ability to have comments. one reason I strongly prefer JSON5 http://json5.org/ http://json5.org/
- jaimex2 10y agoThis is literally protobuffs.
- aboodman 10y agoIt's actually the opposite of protobufs. This format is self-describing - the type information is carried along with the data. Protobufs aren't self-describing. You need to have the type information out-of-line in order to make any sense of serialized protobufs.
- drdaeman 10y agoI'm sort of nitpicking, but Protobufs have wire-level type tags (so old app versions are able to handle newer schemes, with fields they don't know yet). They're limited, but they exist.
- amorphid 10y agoI've been writing a JSON parser when I have a few minutes here and there. I was surprised by the lack of specificity in defining numbers, specifically floats. If floats are know to lose precision after a few decimal places... iex> 1.5555555555555555 1.5555555555555556 ...why not just specify a max precision? You can always say "if you need a more precise number, just store it as a string". If I wanted a room for interpretation, I'd use YAML!
- mitchtbaum 10y agoThis looks similar to msgpack with saltpack for crypto parts. Right? http://msgpack.org/ http://msgpack.org/ https://saltpack.org/ https://saltpack.org/
- justin_vanw 10y agowow, this looks awful and painful. There's no reason to tag the type of a field when you have a typed syntax. The real problems with JSON aren't at all addressed by this: keys have to be strings lack of 'attributes' like xml, which means you have to make a document convoluted from the start. For example, lets say I am storing product data, I might do it like: {'title': "Billy goes to Buffalo", 'page_count': 193, 'author': "Ray Broadbunky"} But later I might want to be able to store attributes or metadata, in xml this doesn't change the schema of the document: <product> <title>Billy goes to Buffalo</title> <page_count>193</page_count> <author>Ray Broadbunky</author> </product> Can be extended to: <product> <title human_verified="false">Billy goes to Buffalo</title> <page_count human_verified="true">193</page_count> <author human_verified="true">Ray Broadbunky</author> </product> It's not beautiful but anything using this data will not have to change at all to add any metadata like this. However, with JSON you have to either add new data that can somehow be joined to the data originally, or more commonly you have to be very defensive and 'plan for' this stuff, greatly complicating the schema. You end up starting with: {'attributes': [ {'name':'title','value':"Billy goes to Buffalo"}, {'name':'page_count', 'value':193, ... so that you can add unanticipated things later without breaking consumers of the data but at least some are addressed: no standard way to store bytestrings lack of time type
- romanovcode 10y agoI'd rather use XML than this atrocity.
- rurban 10y agoThis argumentation is complete bullshit and even dangerous. > "Parsing JSON is a Minefield": From a strictly software engineering perspective these ambiguities can lead to annoying bugs and reliability problems, but in a security context such as JOSE they can be fodder for attackers to exploit. It really feels like JSON could use a well-defined “strict mode”. Not at all. This article just outlined the differences of the various implementations regarding the 2 specs. And then added a spec test suite, including all the undefined problems, with suggestions how to go forward. JSON is already strict enough. The problem are people like op to make it even not-stricter. The latest JSON spec RFC 7159 adds ambiguity by allowing all scalar values on the top level, which leads to practical exploitability. See e.g. https://metacpan.org/pod/Cpanel::JSON::XS#OLD-VS.-NEW-JSON-RFC-4627-VS.-RFC-7159 https://metacpan.org/pod/Cpanel::JSON::XS#OLD-VS.-NEW-JSON-R... "For example, imagine you have two banks communicating, and on one side, the JSON coder gets upgraded. Two messages, such as 10 and 1000 might then be confused to mean 101000, something that couldn't happen in the original JSON, because neither of these messages would be valid JSON. If one side accepts these messages, then an upgrade in the coder on either side could result in this becoming exploitable." What the op now suggests is adding the insecurity-mistake YAML took by adding tags to all keys. Here types don't add security, they weaken security! It is security nightmare as it is leading to exploits which are e.g. already added to metasploit (CVE-2015-1592). tagged decoders are always a problem, and currently JSON and msgpack are the only serializers safe from such exploits due to its strictness. I would suggest that the remaining JSON libraries first fix their problems by conforming to the specs. First the secure old variant (RFC 4627) as default, and then maybe the relaxed new RFC 7159 variant, but denoting the security problems with interop of scalar values. Currently only my Cpanel::JSON::XS library pass all these tests from the Minefield article. E.g. the ruby one, which the author complains about, not. The type problem is esp. problematic in dynamic languages like ruby, where classes are not finalized by default.