5 ms·
Sorry for the delay - was on a break from HN Yes, I think you are right, the availability of XML/ws-* acted as a magnet for people who required extensive schem
by mands 11y ago
Sorry for the delay - was on a break from HN
Yes, I think you are right, the availability of XML/ws-* acted as a magnet for people who required extensive schemas, etc.
I agree that JSON is a great starting place for the rest of us who don't have such immediate needs for complexity and get can by with it. But I think eventually software growth then pushes them to more complex interchange format, e.g. JSON schema.
I think that there is progress, with JSON on the simpler side, and now newer formats like ProtoBufs, Thrift and mechanisms such as RPC - we do seem to be learning from the past. It does feel that perhaps we do swing from one extreme to another - first RPC was great and CORBA came and went, following this the perception was that it was utterly unsuitable for anything, until perhaps the introduction of ProtoBufs, Thift, JSON-RPC and so on. I personally think it can be incredibly useful, but deciding on just the right features to keep things manageable is incredibly difficult (more so than the tech itself I believe).
We're no fans of JSON-Schema either, I've thought about it a few times but it feels over-complicated. Instead we've settled/forked a criminally overlooked system called Barrister RPC (http://barrister.bitmechanic.com/ http://barrister.bitmechanic.com/). This just supports basic JSON types, structs created from their aggregate, and optional nullability. It has worked great so far, although we may expand shortly to add more numeric types. You can try it live at http://www.stackhut.com http://www.stackhut.com (source at http://www.github.com/StackHut http://www.github.com/StackHut) - would love to hear your thoughts re the schema/RPC layer.
- hyperpallium 11y agoI read through all those links, and even got your example working (on a phone - no curl etc - so had to write a litle java http client). Consequently, this comment is long. I hope it's useful to you! BTW: I personally would like to see your idea usable on a phone - without a full local machine (iterating might be a pain, but your system sounds really fast). Not just serverless, but machineless! A currently underserved niche. > deciding on just the right features to keep things manageable is incredibly difficult (more so than the tech itself I believe). I agree implementation is the easier part, though we've been stuck so long, I think there must be a simpler way to look at the whole thing, involving some mathematical or algorithmic insight (as relational algebra did for databases). Barrister: I sometimes get lost in special cases, and forget the main point that provides the fundamental help to people: Barrister seems feature-full, but from reading the first paragraph or so, it seems to only output docs - because that's all they say it does! If the reader already knows what Thift etc do (i.e. something of an expert, across the field), they could guess that maybe Barrister does more... but the users who just want the main thing you do are easier recruits. Perhaps this is partly why it's overlooked... StackHut: I'm not sure you need a separate IDL, if it is generated automatically - why confuse the user with it? It seems like a more sophisticated customization tool, that you could leave aside for later? The IDL itself is pretty clear, and although I'd thought about (eg) java classes as defining a schema, I hadn't made the connection that IDL (as from CORBA) also do that. (1). RPC format. I'm familar with OO serialization and schema languages (unfinshed PhD, book chapters, a library and business), but less so for RPC - so maybe I'm off-base here. And standards - even nascent standards - may be worth complying with. But why not omit the meta stuff, and make it even simpler: { "stackhut/web-tools": { "renderWebpage": ["https://stackhut.com", 711, 393] } } I'm really just wondering if there's a strong reason for the meta data. It can be helpful to orient the reader, but here it is clear from context - the keys and how it's being used. Maybe there are other optional fields you sometimes need? But... for your use-case, perhaps it doesn't matter that much, as the bindings hide it from users (but the point of text protocols is human readable, eg for debugging; so the simpler the better). You could use XML, or a binary protocol. (2). JSON by example: This is my great idea for JSON schema, which I'm amazed no one has done yet: instead of another meta-format, do it by example: Because JSON primitive values are typed, you can use a value to signify type. An object therefore also implicitly defines its type (like a java class). For example, the above JSON can also be used as a schema, because it indicates the two nested objects required, and the types of the primitives (string, number, number). Though I suggest a convention of using "", 0 and false as values. Those zero values help convey that it's a type, not a value. It's temping to want to encode information in the value itself (as opposed to the type), such as a default value - but that maybe a mistake, because there is so much more than can be done with strings than with numbers. Keep it super simple. [ Arrays are usually not fixed-length, enabling the next trick to encode optional values and polymorphism. ] The JSON spec allows duplicate keys: so you include the same key with different types for all the polymorphic types. Specifically, you include the null value to indicate it is optional. e.g. version is an optional string: { "version": "", "version": null } This is a bit dodgey, because all JSON parsers simply return one value if duplicate keys are found. You need to write your own parser. But it is valid JSON - and more importantly, it looks like valid JSON. That's the key idea of "by example" - it looks like what it represents; there's minimal cognitive leap from type to instance. It's rare to want polymorphic primitive types (e.g. string and number), so this is more for polymorphic objects. Unlike the common trick of a "type" field for nominal polymorphism, this is structural polymorphism - where only the different fields distinguishes types. NB: there are some tricky cases here, when the fields overlap, and I'm not sure that client code would want to mess with it. Finally those non-fixed length arrays: polymorphism is represented by the types of a set of values in the array. In other words, the values aren't ordered, but just represent the permissible types. eg: [ { "image_url":"", "width":0, "height":0 }, { "text": "" }, { "link_url":"", "link_text":""} ] That's a schema for an arbitary-length list that can contain instances of those three types of objects, with those mandatory fields, of those primitive types. NB: for a schema of the RPC "header" above, the fixed length array schema represents a fixed length array - a special case. (3). primitive datatypes: This idea can be extended, in a second level, with explicit primitive datatypes - this is the next level of schema power that everyone wants. Every value is now a string, and looks something like: "url", "date", "email", and then gets closer to XS, with ranges like "int:1..31" and even those ridiculous regex defining valid values (great idea, awful in practice, like the regex for email). The key thing is it still is JSON and looks like JSON, since the datatype specification language is just a JSON primitive value (string), and the syntax and meaning is obvious and familar. BTW minor typo on your website: s/intergrated/integrated/ (like integer)
- hyperpallium 11y agoRe: schema/RPC layer Barrister semantically seems like the java type system with minor syntactic differences like field/type order. It uses extends for choice/sum/polymorphism. In addition to schema definition, it also has interfaces for RPC definition. The only major limitation I noticed is it lacks recursion (cyclic type references). Important for completeness, but I don't think it's needed that often in practice - I'm interested if this is true. One great thing about an IDL approach to schema definition, for RPC, is its similarity to objects/classes - because that is the domain it is used from. It's familiar to users, and minimal cognitive burdern to switch between. (Similar argument for JSON - it's an "Object Notation"). In contrast, xml schema looks utterly different from objects/classes. I think there's even a hypothetical argument for making the IDL look exactly like java - not because it's great but because everyone knows it, even the haters (or, maybe borrow Python's type hint system - or whatever is familiar to your target users). XML schema has a rich type system, both for structure (eg minOccurs) and primitives. But these aren't needed to represent objects, because programming type systems are poorer and lack those features. --- I think your main problem is a very old one, of transferring primitive values between different languages (like different languages - even different C implementations - having 16 bit vs 32 bit ints). Text representations like XML and JSON mostly solve this, but of course they can still be too big. You might look at old solutions, like ASN.1, to see the problems that arise, but I think you're better off solving much problems when (and if) they occur - because maybe they just won't come up in practice today because of conventions unconsciously established, even though theoretically they should come up. Just make it as easy to use as possible, to get the thing done that the user wants to do, and you'll get all the feedback/guidance you need. Maybe you would like to have a type system subset that works for all languages - but this would just hamper those in the one language, or between two languages that are close. OTOH some enterprises like future proof options, in case they need to support other languages. Would love to hear your thoughts and experience on all these issues.
- mands 11y agoYes, we've been very happy with Barrister so far, in that it provides a simple syntax for RPC definitions that looks very similar to Java so is not hard to learn. And yes, keeping it similar to standard function/object calls lowers the cognitive load. Hmm, yep recursion not supported atm, although we are thinking of extending the basic system shortly - there are a lot of other features we'd like to add - true sum types, a dynamic operator, and more. Python annotations are nice, and a lot of users have asked for a way to annotate RPC functions in their code. Doing this all in a cross-language way will be hard and so we're thinking of keeping the external definition in the meantime. -- Thanks so much for your advice. We're certainly looking at past attempts, and the research in transferring data is vast - I think it boils down to a trade-off between expression and usability - and for now JSON seems to have found pretty good sweet spot. Perhaps with just a simple schema layer on top that will be suitable for most use-cases. I imagine Thrift, ProtoBufs, etc will always beat it on speed, XML (+XML Schema) on expressibility, etc. For us, as you have guessed, we are just trying to make it as easy as possible for the majority of users – particularly web and mobile developers. This does also mean we're trying to stay language agnostic and find a type subset that works across most languages - JSON primitives let us get some of the way there. Thank you so much for all your advice - it's been so helpful and we're really happy you took the time out to try it all and record your thoughts. Would love to chat further - please feel free to drop me an email at HNusername @stackhut.com.
- hyperpallium 11y ago(this is my third reply) Because researched stackhut (HN, quora, your website), I feel compelled to offer (unsolicited) advice: I couldn't tell what your selling point was - as a commentor on one of your HN submissions also said. I see you've pivoted/changed it around it bit over the last 2-3 months, so that makes it harder. Here's a three-step exercise that might help: (1). What did you feel was cool about your basic idea (which seems to be to "wrap docker and kubernetics for easy microservice deployment in the cloud")? i.e. what excited you, what's really cool, about the technology idea, the original technological impetus? (2). What benefit would other people feel, if they were completely ignorant and had never heard of docker or kubernetics or heroku etc? i.e. aiming here at the basic benefit. To illustrate, "easier to use than kubernetics" doesn't mean anything to someone who's never heard of kubernetics or its difficulties. (Many enterprise customers would be in exactly this perfect state of ignorance - they have enough to deal with with their actual domain!) It wouldn't hurt to also consider what benefit that get from cloud hosting (assume they'd never heard of it, as unlikely as that is). Think of yourself as not converting fellow experts from other technologies, but as gaining adherents for the first time - they don't actually want your tech. They want the cloud and you are just the ladder. (3). Finally, think about target users again, now not as completely ignorant, but in terms of and their current situation: their needs (what their bosses/customers need), their plans, how they presently do things, and especially what technologies they have actually heard of and basically know what they are. [ for example, I've heard of the cloud and docker, but I don't really know what the big deal about docker is. I've heard "kubernetics", but I don't know what it is. ] Now craft a message that speaks to these people, given their present situation and perspective. There will actually be several different target customers, with different situations and knowledge, that would require different messages to really speak to them. For example, I'd like a really simple way to write a bit of code and get it hosted in the cloud, for free, to make a simple webservice (or website). It seems it could be (should be) simple, but it isn't. That would be sensational! And great publicity for paid users (freemium, as you're already using).
- mands 11y agoHi - thanks so much for all the feedback and advice. Have just seen it and will reply fully in the morning (we've been super busy pitching for our accelerator's demo day in the UK - which is an awful excuse!) Really really appreciate you taking the time out to reply and to investigate the tool. Would be great to get on a call sometime and discuss your thoughts further - my email is my HN username @stackhut.com. Thanks again!