10 ms·
Strongly Typed Events
- edejong 7y agos/json/XML/ s/json schema/XSD/ And we’re full circle. I was hoping for some theoretical event-based systems approach, using Pi-calculus to prove correct systems composition.
- ProfHewitt 7y agoSynchronized requesting in Communicating Sequential Processes (CSP) [Hoare 1978] proceeds as follows: “Such communication occurs when one process names another as destination [to receive a request] and the second process names the first as source for [the request] ... [in order that providing the request] is delayed until the other process is ready [to receive the request].” Synchronized requesting x with request r (i.e. x!r) can be implemented as follows using a 2-phase commit protocol: x.synchronize[Implements Provider ⟦provide ↦ r⟧] so that after x has received a synchronize message with parameter Implements Provider ⟦provide ↦ r⟧, x can get r from the parameter using a provide message (cf. [Knabe 1992]). Synchronized sending x a request r (i.e. x!r) can be algebraically reduced (which is a primary requirement of communication in the π-calculus [Milner 1993]) because x is provided with r without arbitration by Implements Provider ⟦provide ↦ r⟧. Synchronized requesting (i.e. x!r) has the following significant costs in time, communication bandwidth, and robustness by comparison with unsynchronized requesting (i.e. x.r): 1. The requester must wait for the receiver’s provide message in order to provide request r. 2. After receiving a synchronize message, the receiver must wait for the request r to be provided (meanwhile holding up processing of other requests). 3. Both the requester and receiver must be online concurrently for communication to take place. Unsynchronized requesting (i.e. x.r) cannot in general be reduced using an algebraic equation as in [Milner 1993] because in general, the request must go through arbitration in order to be received. Although algebraic reductions may be elegant mathematics, synchronized requesting is not widely used in large software systems because it is slower, uses more communication bandwidth, and is less robust than asynchronous requesting (especially for IoT).
- tempguy9999 7y agoI suppose the key to this is in the last line - you're implying something that doesn't wait is better. That'd be actors; your creation, yes? :) But out of interest, the syncing that you apparently dislike but which actors sidestep by having a receiving queue, said syncing allows for messages to be processed when the receiver is ready (obviously). That's a time overhead. If actors have unbounded queues then there's a memory overhead (and one which may grow infinitely on a finite machine if someone isn't processing their messages fast enough). How do actors handle that? If I'm talking rubbish some reading matter is welcome.
- ProfHewitt 7y agoActually, the point is that being faster is better. Ideally, a message sent to an Actor is never stored in persistent memory. A runtime system should never accept a unbounded backlog of communications for an Actor. Instead worst case, it should generate exceptions for further requests. Here are references that you requested: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3418003 https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3418003 https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3459566 https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3459566
- tempguy9999 7y agoMuch obliged, thank you.
- AdieuToLogic 7y agoA good article on the Actor model can be found here[0]. > If actors have unbounded queues then there's a memory overhead (and one which may grow infinitely on a finite machine if someone isn't processing their messages fast enough). Depending on the Actor implementation one employs, this concern can be mitigated. For example, Akka[1] provides BoundedMessageQueueSemantics[2] which can be used to provide hard limits. 0 - https://en.wikipedia.org/wiki/History_of_the_Actor_model https://en.wikipedia.org/wiki/History_of_the_Actor_model 1 - https://doc.akka.io/docs/akka/current/typed/index.html https://doc.akka.io/docs/akka/current/typed/index.html 2 - https://doc.akka.io/docs/akka/current/mailboxes.html#requiring-a-message-queue-type-for-an-actor https://doc.akka.io/docs/akka/current/mailboxes.html#requiri...
- DonHopkins 7y agoAs long as we're circling back, let's circle back to Relax/NG instead of XSD, please! http://www.drdobbs.com/a-triumph-of-simplicity-james-clark-on-m/184404686 http://www.drdobbs.com/a-triumph-of-simplicity-james-clark-o... Reposting: https://news.ycombinator.com/item?id=20384206 https://news.ycombinator.com/item?id=20384206 He's not that famous outside of hard core geek circles, but I am a huge fan of James Clark (not the SGI guy, but he's awesome too for different reasons), who wrote the Expat XML parser, Relax/NG XML schema language, and many other widely used standards and bodies of code that drive the Internet. http://www.jclark.com/ http://www.jclark.com/ http://www.jclark.com/bio.htm http://www.jclark.com/bio.htm He has such clarity of thought and mastery of diverse languages (some of which he invented and implemented). He saw exactly what was wrong with XML Schemas, and addressed it practically and elegantly with TREX, which he and others refined through humble constructive collaboration into Relax/NG. And he's a big proponent and creator (not just a talker like ESR) of free open source software. https://relaxng.org/jclark/ https://relaxng.org/jclark/ Here's a fascinating insightful DDJ interview of James Clark, "A Triumph of Simplicity: James Clark on Markup Languages and XML": http://www.drdobbs.com/a-triumph-of-simplicity-james-clark-on-m/184404686 http://www.drdobbs.com/a-triumph-of-simplicity-james-clark-o... >A Triumph of Simplicity: James Clark on Markup Languages and XML >If you peek under the hood of high-profile open-source projects such as Mozilla, Apache, Perl, and Python, you'll find a little program called "expat" handling the XML parsing. If you've ever used the man command on your GNU/Linux distribution, then you've also used groff, the GNU version of the UNIX text formatting application, troff. If you've ever done any work with SGML, from generating documentation from DocBook to building your own SGML applications, you've undoubtedly come across sgmls, SP, and Jade. >Whether you've heard of him or not (and mostly likely, you haven't), James Clark (below right) has made your life easier. In addition to authoring these and other widely used open-source tools (see http://www.jclark.com/ http://www.jclark.com/ for a complete list), Clark served as the technical lead of the original W3C XML Working Group and as the editor of the XSLT and XPath recommendations. He recently founded Thai Open Source Software Center (http://www.thaiopensource.com/ http://www.thaiopensource.com/). His latest project is TREX, an XML schema language. Clark sat down with Eugene Eric Kim to discuss markup languages, the standardization process, and the importance of simplicity. [...] >DDJ: You're well known for writing very good reference implementations for SGML and XML Standards. How important is it for these reference implementations to be good implementations as opposed to just something that works? >JC: Having a reference implementation that's too good can actually be a negative in some ways. >DDJ: Why is that? >JC: Well, because it discourages other people from implementing it. If you've got a standard, and you have only one real implementation, then you might as well not have bothered having a standard. You could have just defined the language by its implementation. The point of standards is that you can have multiple implementations, and they can all interoperate. >You want to make the standard sufficiently easy to implement so that it's not so much work to do an implementation that people are discouraged by the presence of a good reference implementation from doing their own implementation. >DDJ: Is that necessarily a bad thing? If you have a single implementation that's good enough so that other people don't feel like they have to write another implementation, don't you achieve what you want with a standard in that all implementations — in this case, there's only one of them — work the same? >JC: For any standard that's really useful, there are different kinds of usage scenarios and different classes of users, and you can't have one implementation that fits all. Take SGML, for example. Sometimes you want a really heavy-weight implementation that does validation and provides lots of information about a document. Sometimes you'd like a much lighter weight implementation that just runs as fast as possible, doesn't validate, and doesn't provide much information about a document apart from elements and attributes and data. But because it's so much work to write an SGML parser, you end up having one SGML parser that supports everything needed for a huge variety of applications, which makes it a lot more complicated. It would be much nicer if you had one SGML parser that is perfect for this application, and another SGML parser that is perfect for this other application. To make that possible, the standard has to be sufficiently simple that it makes sense to have multiple implementations.
- agumonkey 7y agoI see the same pattern with early php/python where it was all about ease and RAD and now 80s sweng/plt ideas are coming back. Social tissue has a resitance.
- eyjafjallajokul 7y agoPretty sure that the author knows about XML and XSD - https://en.wikipedia.org/wiki/Tim_Bray#XML https://en.wikipedia.org/wiki/Tim_Bray#XML . As an Invited Expert at the World Wide Web Consortium between 1996 and 1999, Bray co-edited the XML and XML namespace specifications.
- ProfHewitt 7y agoStrongly-typed events are axiomatized up to a unique isomorphism here: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3418003 https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3418003 https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3459566 https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3459566
- guitarbill 7y agoFirst CloudFormation adds a registry [0], now EventBridge [1]. Wonder if AWS always builds duplicates, and if we'll end up with a schema service. [0] https://aws.amazon.com/blogs/aws/cloudformation-update-cli-third-party-resource-support-registry/ https://aws.amazon.com/blogs/aws/cloudformation-update-cli-t... [1] https://aws.amazon.com/blogs/compute/introducing-amazon-eventbridge-schema-registry-and-discovery-in-preview/ https://aws.amazon.com/blogs/compute/introducing-amazon-even...
- thelittlenag 7y agoThinking about events to me is a sign that you are probably concerned with the wrong level of abstraction. Instead, you should be thinking about building protocols, of which events may be part. To me at least, a protocol combines notions of both the means of communication, as well as the semantics of that communication. In this context of this article, I would prefer to version a protocol. That version subsumes the particular schema of the events used for that version of the protocol. If your protocol is well-designed then many of the assumptions of semantic versioning can be applied and relied on, so that parts of your system can speak an older version compatibly with a parts that speak a slightly newer version. I really wish there was more support in the major ecosystems for creating protocols and a less obsessive focus on message encodings. Oh well.
- proc0 7y ago> so that parts of your system can speak an older version compatibly with a parts that speak a slightly newer version. I think this doesn't guarantee bug-free event code at all. The usual problems would still arise, I can see at least this would create repetition in logic which doesn't scale properly with large apps. Additionally versioning is needless if you have strongly typed events that get checked at compile time. Any time you need to add an abstract feature to the event architecture, the compiler would tell you what other part of the app broke as a result.
- thelittlenag 7y ago> I think this doesn't guarantee bug-free event code at all. Please don't misinterpret me. The goal is to have fewer issues, not zero issues. > Additionally versioning is needless if you have strongly typed events that get checked at compile time. Any time you need to add an abstract feature to the event architecture, the compiler would tell you what other part of the app broke as a result. I disagree. Most type systems, especially those describing data, are really only good for checking syntactic consistency and not very good at semantic consistency. That is, it can tell if field changes from an Int to a String, but not if the content of a String field changes it's semantic, i.e. from a user id to an email address. Lack of ability to validate content pervades wire-appropriate encodings.
- proc0 7y agoI'm new to Elixir/Erlang, but it seems that is what those languages are going for, an easy way to properly use event based architectures. Someone with more XP here can correct me.
- agentultra 7y agoVersioning is important. I wrote a library in Haskell for indexing a family of types by a natural number (representing the version). It includes some generics-based machinery for enabling type-correct migrations from type n to type (n + 1) [0]. I've used it for migrating schema-less documents to formats with a strongly-typed schema. I'm also using it for versioning event streams and it works well. It'd be really neat to be able to share the schemas of these messages in a wider context. Versioning events and the stream is a useful property to have. More types the better. [0] https://hackage.haskell.org/package/DataVersion-0.1.0.0/docs/Data-Migration.html https://hackage.haskell.org/package/DataVersion-0.1.0.0/docs...
- CoolGuySteve 7y agoI write a lot of low level C++ network/disk stuff. I need minimal decoding costs and rely on the type system to catch message differences asap. The following is an example of an append only format that uses the sizeof as a kind of version. In practice, not only is this infinitely faster performance-wise, but I find maintaining this is about the same amount of work as Protobufs or JSON but with less fussy tooling issues. By using X macros or BOOST_FUSION for the struct definitions, you can infer the type definition of message fields and automatically write serializers for JSON, SQL, CSV, Python, or whatever schema format is your favourite. Unfortunately, all that macro/variadic magic is too verbose and fiddly for a hnews comment. The important thing is that the drudgery of schema maintenance and conversion can be eliminated using C++'s (limited) reflection capabilities. Your code is the schema. In WhateverMessages.h: #pragma pack(1) template<typename Message> struct Header { uint32_t type = Message::MessageType; uint16_t size = sizeof(Message); }; struct Msg : Header<Msg> { static constexpr uint32_t MessageType = 'msg1'; // Can constexpr bswap this to appear correctly in GDB and hex dumps uint32_t somedata; char someText[32]; }; #pragma pop(pack) In EventHandler.cpp: void processMessage(const Msg& msg) { // Everything is type safe now, do real stuff } void processMessage(char* data, uint32_t len) { Header* hdr = reinterpret_cast<Header*>(data); // should check the packet length is sufficient, return length processed, process multiple messages, etc switch(hdr->type) { #define DISPATCH(m) case m ::MessageType: assert(hdr->size == sizeof(m) /*or warn, whatever */); \ processMessage(*reinterpret_cast<m*>(hdr)); break; DISPATCH(Msg); DISPATCH(OtherMsgDeprecated); DISPATCH(OtherMsgButForRealThisTime); #undef DISPATCH default: err("Unknown message type: %x (%s)\n", hdr->type, fourCCToString(hdr->type)); break; } } Another important thing is that the raw binary is the first class format that requires no translation of any kind on x86-64 after the initial dynamic type inference in DISPATCH. Python, JSON, SQL, etc are all slow anyways, so we can spend time massaging their serialization format afterwards.
- mcguire 7y ago"Writing code to map back and forth between bits-on-the-wire and program data structures is a very bad use of developer time." I'm going to disagree. Those transitions happen very frequently and can materially affect both latency and throughput. Spending developer time on these things can give a large return. "Among other things, most messages are in JSON..." Ah.
- IshKebab 7y agoI think the proper solution is to use a format that explicitly requires a schema (e.g. Protobuf). If your schema is implicit then people won't bother using it and it'll be a hassle to get them to even write it down.