6 ms·
I agree that getting just the right amount of simplicity is actually a hard challenge. JSON over HTTP feels like it has won as it was so trivial to implement in
by mands 11y ago
I agree that getting just the right amount of simplicity is actually a hard challenge. JSON over HTTP feels like it has won as it was so trivial to implement in any language, however experience has taught me that the lack of a schema causes problems over time.
At StackHut (www.stackhut.com) we're using JSON with a lightweight schema. So far this is working very well, providing simple yet 'typed' remote interfaces into containers. However I often wonder about switching to XML/XML-schema/XML-RPC, protocol buffers/gRPC or some custom system in the future or to just keep it simple and not over-complicate things.
- hyperpallium 11y agoI think what happened is that people building a complex system, who knew they needed schema etc would just use the XML/ws-* ecosystem, because as repugnant as some might feel it is, it is all written, debugged and works. If people didn't need complexity, they used JSON. The XML option kept complexity out of JSON. However, its inevitable people starting with simple systems (using JSON) would become complex - perhaps partly due to dramtic success. And converting everything to XML would be time-consuming and error-prone and... repugnant. So. This is where the demand for json-schema comes from. And it does seem inevitable: the inertia of back-compatibility is one of the more predictable features of software. I can make the hopeful observation that things aren't quite as bad as the cynic take above: people do learn from some of the mistakes of previous technologies. There is some progress. I dislike JSON-schema because it's like a JSON version of XML schema. I think a simpler scheam would better serve JSON and its typical uses. Is your "lightweight schema" a simple schema written in json-schema? Or written in a lightweight schema language?
- mands 11y agoSorry for the delay - was on a break from HN Yes, I think you are right, the availability of XML/ws-* acted as a magnet for people who required extensive schemas, etc. I agree that JSON is a great starting place for the rest of us who don't have such immediate needs for complexity and get can by with it. But I think eventually software growth then pushes them to more complex interchange format, e.g. JSON schema. I think that there is progress, with JSON on the simpler side, and now newer formats like ProtoBufs, Thrift and mechanisms such as RPC - we do seem to be learning from the past. It does feel that perhaps we do swing from one extreme to another - first RPC was great and CORBA came and went, following this the perception was that it was utterly unsuitable for anything, until perhaps the introduction of ProtoBufs, Thift, JSON-RPC and so on. I personally think it can be incredibly useful, but deciding on just the right features to keep things manageable is incredibly difficult (more so than the tech itself I believe). We're no fans of JSON-Schema either, I've thought about it a few times but it feels over-complicated. Instead we've settled/forked a criminally overlooked system called Barrister RPC (http://barrister.bitmechanic.com/ http://barrister.bitmechanic.com/). This just supports basic JSON types, structs created from their aggregate, and optional nullability. It has worked great so far, although we may expand shortly to add more numeric types. You can try it live at http://www.stackhut.com http://www.stackhut.com (source at http://www.github.com/StackHut http://www.github.com/StackHut) - would love to hear your thoughts re the schema/RPC layer.
- hyperpallium 11y agoI read through all those links, and even got your example working (on a phone - no curl etc - so had to write a litle java http client). Consequently, this comment is long. I hope it's useful to you! BTW: I personally would like to see your idea usable on a phone - without a full local machine (iterating might be a pain, but your system sounds really fast). Not just serverless, but machineless! A currently underserved niche. > deciding on just the right features to keep things manageable is incredibly difficult (more so than the tech itself I believe). I agree implementation is the easier part, though we've been stuck so long, I think there must be a simpler way to look at the whole thing, involving some mathematical or algorithmic insight (as relational algebra did for databases). Barrister: I sometimes get lost in special cases, and forget the main point that provides the fundamental help to people: Barrister seems feature-full, but from reading the first paragraph or so, it seems to only output docs - because that's all they say it does! If the reader already knows what Thift etc do (i.e. something of an expert, across the field), they could guess that maybe Barrister does more... but the users who just want the main thing you do are easier recruits. Perhaps this is partly why it's overlooked... StackHut: I'm not sure you need a separate IDL, if it is generated automatically - why confuse the user with it? It seems like a more sophisticated customization tool, that you could leave aside for later? The IDL itself is pretty clear, and although I'd thought about (eg) java classes as defining a schema, I hadn't made the connection that IDL (as from CORBA) also do that. (1). RPC format. I'm familar with OO serialization and schema languages (unfinshed PhD, book chapters, a library and business), but less so for RPC - so maybe I'm off-base here. And standards - even nascent standards - may be worth complying with. But why not omit the meta stuff, and make it even simpler: { "stackhut/web-tools": { "renderWebpage": ["https://stackhut.com", 711, 393] } } I'm really just wondering if there's a strong reason for the meta data. It can be helpful to orient the reader, but here it is clear from context - the keys and how it's being used. Maybe there are other optional fields you sometimes need? But... for your use-case, perhaps it doesn't matter that much, as the bindings hide it from users (but the point of text protocols is human readable, eg for debugging; so the simpler the better). You could use XML, or a binary protocol. (2). JSON by example: This is my great idea for JSON schema, which I'm amazed no one has done yet: instead of another meta-format, do it by example: Because JSON primitive values are typed, you can use a value to signify type. An object therefore also implicitly defines its type (like a java class). For example, the above JSON can also be used as a schema, because it indicates the two nested objects required, and the types of the primitives (string, number, number). Though I suggest a convention of using "", 0 and false as values. Those zero values help convey that it's a type, not a value. It's temping to want to encode information in the value itself (as opposed to the type), such as a default value - but that maybe a mistake, because there is so much more than can be done with strings than with numbers. Keep it super simple. [ Arrays are usually not fixed-length, enabling the next trick to encode optional values and polymorphism. ] The JSON spec allows duplicate keys: so you include the same key with different types for all the polymorphic types. Specifically, you include the null value to indicate it is optional. e.g. version is an optional string: { "version": "", "version": null } This is a bit dodgey, because all JSON parsers simply return one value if duplicate keys are found. You need to write your own parser. But it is valid JSON - and more importantly, it looks like valid JSON. That's the key idea of "by example" - it looks like what it represents; there's minimal cognitive leap from type to instance. It's rare to want polymorphic primitive types (e.g. string and number), so this is more for polymorphic objects. Unlike the common trick of a "type" field for nominal polymorphism, this is structural polymorphism - where only the different fields distinguishes types. NB: there are some tricky cases here, when the fields overlap, and I'm not sure that client code would want to mess with it. Finally those non-fixed length arrays: polymorphism is represented by the types of a set of values in the array. In other words, the values aren't ordered, but just represent the permissible types. eg: [ { "image_url":"", "width":0, "height":0 }, { "text": "" }, { "link_url":"", "link_text":""} ] That's a schema for an arbitary-length list that can contain instances of those three types of objects, with those mandatory fields, of those primitive types. NB: for a schema of the RPC "header" above, the fixed length array schema represents a fixed length array - a special case. (3). primitive datatypes: This idea can be extended, in a second level, with explicit primitive datatypes - this is the next level of schema power that everyone wants. Every value is now a string, and looks something like: "url", "date", "email", and then gets closer to XS, with ranges like "int:1..31" and even those ridiculous regex defining valid values (great idea, awful in practice, like the regex for email). The key thing is it still is JSON and looks like JSON, since the datatype specification language is just a JSON primitive value (string), and the syntax and meaning is obvious and familar. BTW minor typo on your website: s/intergrated/integrated/ (like integer)