11 ms·
Differential: Type safe RPC that feels like local functions
- CipherThrowaway 3y agoHow does the idempotency work? I don't understand how functions can be made idempotent without changing or instrumenting their implementation for end-to-end idempotency.
- scarmig 3y agoYou could have some side channel with an identifier for the RPC. If the server receives duplicates of one RPC, it ignores the extras. It would require some client support, though. Alternatively, if the RPCs have the same arguments, the framework could choose to ignore seemingly identical ones, though that has its own issues.
- CipherThrowaway 3y agoNeither method works. If the framework records the occurrence of the call before the effect of the function, you achieve at-most-once semantics. If it records the call after the effect, you get at-least-once. The framework might perform its idempotency bookkeeping within the same transaction boundary as the function's side-effects, but this means the function implementation is no longer a black-box. E.g. you can no longer perform arbitrary side-effects while preserving the idempotency guarantees of the framework.
- jitl 3y agoIt’s at-least-once. From https://docs.differential.dev/advanced/compute-recovery/ https://docs.differential.dev/advanced/compute-recovery/ > If a machine fails to send any heartbeats within an interval (default 90 seconds): > It is marked as unhealthy, and Differential will not send any new requests to it. > The functions in progress are marked as failed, and Differential will retry them on a healthy worker. I would guess that “idempotent” functions in the system also take a lease out on the idempotency key. Perhaps they release the lease on failure since they can observe errors thrown, and they commit the key as consumed after success. ¯\_(ツ)_/¯ the docs are not clear on these semantics!
- Salgat 3y agoThat's how we handle idempotency. It's basically a mutex on the idempotency key with a timeout (in our case using redis with redlock since it's a distributed system). Once the command finishes, the key is marked as handled and the lock is released. At that point any future or queued requests with the same key immediately return.
- CipherThrowaway 3y agoIt sounds to me like there are scenarios where calls to, and certainly effects of, "idempotent" functions can take place twice. I haven't looked at Redlock for a while, but I'm assuming you're all across the historical objections to it from distributed systems researchers: https://martin.kleppmann.com/2016/02/08/how-to-do-distributed-locking.html https://martin.kleppmann.com/2016/02/08/how-to-do-distribute...
- fweimer 3y agoIs there really consensus that practical systems need to have lock lease timeouts?
- jitl 3y agoIt’s a choice between at-least-once and at-most-once right? Without a timeout you get at-most-once since the lease holder may die off and never complete the task, and so the task of a completed 0 times. I’m building a system right now and we’re going to use lease/heartbeats etc etc to elect leaders and whatnot, all this discussion seems very practical to me.
- Salgat 3y agoExactly. It comes down to whether the consequences of these race conditions are acceptable or not. Sometimes (known and understood) race conditions are okay for performance reasons. Shoot, that's the whole philosophy behind eventual consistency. Sometimes they aren't, such as bank transactions. It's all situational.
- reissbaker 3y agoYeah, this is what it looks like in the docs; you choose an "idempotency key": https://docs.differential.dev/advanced/idempotency/ https://docs.differential.dev/advanced/idempotency/
- asplake 3y ago"Just decorate your idempotent operations" – implies to me that the idempotency of your functions is your responsibility.
- CipherThrowaway 3y agoThat's one way of reading it - although I find it's at odds with the heading "Idempotency in one line of code." If it's the case that the idempotency wrapper is only for operations that are already idempotent, it makes me wonder that the docs couldn't do a better job of communicating the purpose and function. Decorate your idempotent operations.. to make them more idempotent?
- lunarcave 3y agoNot quite. The decorator [1] enforces you to provide an idempotency key, if idempotence is needed. The control-plane intercepts that and implements idempotence for you. [1] https://docs.differential.dev/advanced/idempotency/ https://docs.differential.dev/advanced/idempotency/
- hmeh 3y agoYes, quite. Unless my function has no side-effects, idempotency will always be my responsibility. This feature is a lie and an outright hazard. It should not be called idempotency and it should have a big bold warning stating that the function may be called more than once.
- lunarcave 3y agoI'd appreciate it if you could point out how the function can be called more than once. Here's the source code for the orchestrator [1] and a demonstration of it [2]. If there's an error in the source code or the copy, I'll happily retract it. [1] https://github.com/differentialhq/differential/blob/236ffc5367fc0c8d8645fbdf41603592f7a2630d/control-plane/src/modules/jobs/create-job.ts#L45 https://github.com/differentialhq/differential/blob/236ffc53... [2] https://github.com/differentialhq/differential/blob/236ffc5367fc0c8d8645fbdf41603592f7a2630d/ts-core/src/tests/idempotency/idempotency.test.ts#L4 https://github.com/differentialhq/differential/blob/236ffc53...
- lunarcave 3y agoHere are the docs on idempotency: https://docs.differential.dev/advanced/idempotency/ https://docs.differential.dev/advanced/idempotency/
- CipherThrowaway 3y agoYou mentioned elsewhere that you're still finding the right wording for the product and its benefits. Fair enough. Personally, I think you should probably stop referring to this as idempotency. Maybe describe it as a retry policy, or as a choice between at-least-once and at-most-once. I expect most people with knowledge of - and respect for - distributed systems theory will be completely put-off by this wording. In distributed systems, the reason idempotency is such a Big Deal™ is that it can be combined with at-least-once delivery to achieve exactly-once semantics. This isn't that. What you're describing as idempotency has the potential to mislead users, especially since you use examples like credit card processing.
- lunarcave 3y agoThanks for the thoughtful reply. You might be on to something here: > Maybe describe it as a retry policy, or as a choice between at-least-once and at-most-once and we will consider changing it because I agree that conflating the concepts (or the appearance thereof) is not worth diluting the other parts of the offering. I still struggle to find the difference between this approach and the likes of Stripe [1] and Temporal [2] because in practice, it yields the same result. [1] https://docs.stripe.com/api/idempotent_requests https://docs.stripe.com/api/idempotent_requests [2] https://www.restack.io/docs/temporal-knowledge-temporal-io-idempotency https://www.restack.io/docs/temporal-knowledge-temporal-io-i...
- CipherThrowaway 3y agoOne difference is that you are specifically creating distributed systems middleware, where your audience is going to be using your product in the implementation of a system of their own, and is more likely to understand the terminology and wish to know the exact guarantees and trade-offs being provided. I find both those comparisons are apples to oranges: 1. Temporal does not claim, on that page, to be able to automatically make your calls idempotent. Instead, it is explaining how you - as the user - can and should write your activities to be idempotent. Take the payment code snippet: you can't just wrap `ExternalPaymentAPI.Process(paymentId, amount)` in an idempotency provider. You have to implement specific idempotency logic inside the activity. This is in line with the distributed systems theory of idempotency. 2. In these docs, Stripe describes an interface for achieving idempotency as an external caller. Since Stripe control the implementation of their own system, it is certainly possible that they have implemented their operations in a way that is truly idempotent. What your docs describe - the ability to make arbitrary operations idempotent by wrapping them - is simply not possible. I wonder whether you're focusing on the idempotency key as the thing that is wrong here. There is nothing wrong with idempotency keys themselves - they are a great way for APIs to provide idempotency. The problem is that this guarantee can't be provided by a wrapper layer, it requires the operation itself to be written in an idempotent way. This is an example of the age old "end-to-end principle" from networks class.
- jahewson 3y agoOk… but how does it work? Too much magic.
- shoulderfake 3y ago[dead]
- lunarcave 3y agoThis (and the linked docs) might help understanding how it works: https://docs.differential.dev/getting-started/thinking/ https://docs.differential.dev/getting-started/thinking/
- mblode 3y agoHuge congratulations on the launch!
- lunarcave 3y agoThanks Matt :)
- pjmlp 3y agoUsually type safety isn't the issue, rather pretending the network doesn't exist. RPC, even if running on the same machine, is distributed systems land.
- nextaccountic 3y agoIs this just like trpc? https://trpc.io/ https://trpc.io/ Or like rspc (in Rust) https://www.rspc.dev/ https://www.rspc.dev/
- kszyh 3y agoIf I understand the idea correctly, differential is for backend only (microservices architectures). There is nothing about browser support in the docs.
- benatkin 3y agoI think they’re probably paying more attention to gRPC and GraphQL. They have a cloud offering whereas trpc seems to have a smaller scope and prioritize supporting every client and server.
- lunarcave 3y agoClose. Think of it like trpc / rspc with a service mesh that orchestrates calls between services. It _does_ put an extra hop in your network calls, but it allows some the benefits of orchestration. Here's an opinionated take: https://docs.differential.dev/getting-started/thinking/ https://docs.differential.dev/getting-started/thinking/
- eqvinox 3y agoIf it "feels like local functions", how do you handle errors from the RPC layer / network?
- JonChesterfield 3y agoThis really should be addressed in the docs and I can't find it either. There's some stuff about retrying functions that failed sometime later which seems very different to local functions and very like not-local functions.
- IshKebab 3y agoIt's Typescript so presumably it makes the functions async to handle network delays, and errors cause exceptions to be thrown. Doesn't seem like a big issue to me.
- pjmlp 3y agoThat isn't even scratching the surface of what can go wrong on distributed systems.
- lunarcave 3y agoI agree that OP only reduced his statement to a subset of the things that can go wrong. But the error handling mechanism is uniform across any errors that can be caused.
- JonChesterfield 3y agoTherefore the function has its own failure modes and also a set of extra ones from the execution framework, folded into an exception hierarchy?
- IshKebab 3y agoThat is how exceptions always work in these sort of languages. I think it's one of the reasons exceptions suck as an error mechanism (checked exceptions are the "answer" but hardly anyone uses those - so few in C++ that they were removed from the language!), but it isn't a new problem introduced by this framework. In any case Typescript doesn't have checked exceptions so you're pretty much stuck with "any error can happen anywhere" in Typescript whether you use this framework or not.
- EdSchouten 3y agoThe downside of an approach like this is that it limits you to a single programming language on both the frontend and backend. This is exactly the reason RPC frameworks like gRPC use an IDL.
- catern 3y agoWell, that's not fundamental. There are plenty of projects that support generating typed bindings for libraries in one language for use by programs in another language. You could do the same thing here.
- xyzzy_plugh 3y agoFundamentally you go against the grain by generating an IDL from source code. I've never seen this done properly such that it is an improvement over just authoring the IDL in the first place and generating from that. If your IDL isn't a first-class citizen then what do you expect the derivatives to be?
- lunarcave 3y agoI agree, and disagree. I agree that it limits the usage of the tool in a polygot environment. I disagree it's a downside. The absence of an intermediary language does give some benefits to the first class citizen (in this case, Typescript). However, there are some other developments [1] which attempt to make the Typescript type system an IDL to allow for better interop. [1] https://github.com/Aleph-Alpha/ts-rs https://github.com/Aleph-Alpha/ts-rs
- moltar 3y agoHow would this work in a serverless scenario where handler side runs in a Lambda for example? It seems a bit of a chicken and egg in this case. A service needs to register itself with a control plane. But for the service to start it needs to get a request (via lambda invocation).
- lunarcave 3y agoFair question. > But for the service to start it needs to get a request (via lambda invocation). A service can also start by manually `.invoke`-ing the lambda. The control-plane will start the lambda function when there's work. Lambda "asks" for work to do. Once work has finished, lambda function exits. When deploying, we start the lambda function once so it can come out for air and advertise itself to the control-plane. This is an affordance we do only do for lambda, and currently in development with our deployment offering here [1] [1] https://github.com/differentialhq/differential/blob/236ffc5367fc0c8d8645fbdf41603592f7a2630d/control-plane/src/modules/deployment/lambda-provider.ts#L48 https://github.com/differentialhq/differential/blob/236ffc53...
- bugbuddy 3y agoThe key word here is “feels.” I would not trust feelings when it comes to software systems’ reliability.
- lunarcave 3y agoThe "feels" part is around the ergonomics of function calls. It doesn't pertain to system reliability.
- JonChesterfield 3y agoThis might be an interesting project. "Type safe RPC" is not a great description to lead with as remote procedure calls have a large amount of negative associations that have little to nothing to do with static type systems. I think it's a distributed language runtime where you write functions in typescript and the runtime deals with passing data between machines. The website is quite good at presenting the system as a nice thing that does nice things and you should like it and it's all very nice. It does a very poor job of convincing me that it works properly. I'd want to know how the scheduler works and how it handles mutable state in the presence of network partitions. Also just found a caveat which amounts to they've got the implementation wrong: > Arguments must be JSON serializable. This means that you can pass strings, numbers, booleans, arrays, objects, and null. You cannot pass functions, promises, or other non-serializable objects
- jtsiskin 3y agoWhat do you mean by “they’ve got the implementation wrong”?
- winwang 3y agoNot sure, but functions are serializable.
- williamdclt 3y agoIt’s technically correct but really we all know what was meant. They’re serialisable but not really deserialisable, even if it’s a pure function it might not even be serialised in a version of JS that the receiver understands.
- jrumbut 3y ago> serialised in a version of JS that the receiver understands I don't think the use case for this is as general as most RPC systems. The linked site envisions a situation where you had an app running on one server, now you want to break that workload up while making the absolute minimum number of code changes. All the clients and code would still be controlled by the same team.
- ushakov 3y agoHow is this a company?
- infamouscow 3y agoA ZIRP phenomena.
- lunarcave 3y agoIt's really not a company (at least not yet). We've been a bunch of devs hacking together on this for a few months (~6). And we definitely haven't had the blessing (or the curse) of ZIRP.
- valenterry 3y agoRemote procedure calls (or whatever it is called) simply can never really feel like local functions. It has been tried again and again for decades. Maybe call it "removes as much friction as possible compared to local functions" but please don't call it "feels [exactly] like local functions". It's simply impossible and pretending so will just cause trouble.
- DrDeadCrash 3y agoI once wrote a system for distributed data processing (8-10 years ago). There was a base class that when inherited from would, via Interception reroute a call to any virtual method to a server with the same interface loaded (plug-in architecture). The server would then delegate the method to one of its worker clients, which had the (same) required plugin(s) loaded. It worked over TCP on a local intranet. Once the admin-server-worker plugin system was figured out, projecting a virtual method was fairly transparent. There's always going to be some boiler-plate!
- calebpeterson 3y agoI suspect parent was referring more to the runtime characteristics of latency, error handling, retries, load balancing, etc… more than the syntax. Nonetheless, what you built sounds quite interesting: can you say anything more about the language, tech stack, etc…?
- photonthug 3y agoNot OP but I too have built this, basically a few hundred lines of python to implement something like json rpc (plus introspection to answer help requests / publish typed api info) on top of AWS lambda.
- osigurdson 3y agoI had to replace a .NET remoting based system with thousands of individual calls, with all manner of in process like conventions, with an http layer without changing client code at all in a few weeks. The result was good but I still don't like transparent rpc (on large dev teams anyway).
- eqvinox 3y ago"Predictive Retries: With the help of AI Differential detects transient errors (database deadlocks) and retries the operation before the client even notices. See Predictive Retries for more." [https://docs.differential.dev/ https://docs.differential.dev/] "If it's predicted to be transient, Differential will retry the operation on a healthy worker before the client even notices." [https://docs.differential.dev/advanced/predictive-retries/ https://docs.differential.dev/advanced/predictive-retries/] I… er… wat? Does that mean this behavior is nondeterministic, driven by what some extra component cooks¹ up on a per-call basis? … Doing this without developer input is such a horribly bad idea, and then debug it when something (not even in the RPC layer) goes wrong? [Edit:] "You can turn this feature on for your cluster using the Console. It's off by default." – a breath of relief. -- ¹ I really had to resist the temptation to say "hallucinates" there.
- rco8786 3y agoWoof, yea that sounds really, really dangerous.
- lunarcave 3y agoOne of the authors here. So, the implementation _is_ deterministic per observed error, that is because it caches the previously cooked up result of "is this retryable or not" and apply it to next observed error without trying to process it for each function call. Whether it does correctly infer retry-ability - well, that's another question. It's an opt-in feature, and it's not default-on. So, it does do this _with_ developer input. If you prefer to not use this abstraction and have total control over error handling, you definitely can do that. (And it's the default) If I reduce the tradeoffs here, it's between: 1. Explicitly handling all your errors for complete control of retry-ability. 2. Using predictive retries when you want some "good enough" error handling. I agree that we can't really be 100% "correct" in determining whether something is retryable. But I think it's a bounded enough problem for a language processor (ala LLM) that can yield good enough results for vast majority of the use cases. Here's a test case that asserts this behaviour (without any caching): https://github.com/differentialhq/differential/blob/236ffc5367fc0c8d8645fbdf41603592f7a2630d/control-plane/src/modules/jobs.test.ts#L360 https://github.com/differentialhq/differential/blob/236ffc53... So in a way, it is bounded enough to yield this result deterministically in our test suite. But I agree there's a non-zero chance it fails, although we haven't observed it yet after thousands of test runs.
- rco8786 3y agoIt's an interesting/useful concept and most big tech co's I've worked at have some form of this internally. But I think it's a little dangerous to market this as "feels like local functions". It glosses over a lot of really important technical considerations (like, the network and all the dragons that come with it) which will eventually bite people who don't understand that, and will immediately turn people off who do understand that. I'm not a marketing or copy expert by any means, but what might be a neat product is (rightfully) getting criticized in the comments here for the positioning.
- lunarcave 3y agoOne of the authors here. Completely agree. To be honest, we didn't intend to have this many eyes on the product as we're still iterating (both the abstraction and the copy). We've gone through a lot of iterations on the copy itself, and I agree that there might be a better framing here which works.
- kgeist 3y agoThe site lists all those features but it's not clear what problem it's trying to solve. Only in the FAQ at the end of the page it's said it's a service mesh basically with added bells & whistles, from what I understand. But I'm still confused: >Monolithic codebases don't have to result in monolithic services. Why would I want to replace straighforward function calls with network calls in my monolith? >By using a centralised control-plane, Differential transparently handles network faults and machine restarts with retries, all without changing your existing programming paradigm. It is designed to be a drop-in replacement for any function call that you'd like to make distributed and reliable. "Centralization" and "reliability" don't sound like words that often come together. What if the control plane goes down? Unless they have a cluster? It's all self-imposed problems anyway: if we don't have network calls in the monolith, we don't have all those problems in the first place (which need to be solved with retries etc.)
- lunarcave 3y agoThanks for the feedback. One of the authors here. > Why would I want to replace straighforward function calls with network calls in my monolith? I agree that there's nothing inherently virtuous with replacing a functioning local call with a remote procedure call. But if you decide to break your monolith up for any one of the reasons people tend to follow service-oriented architecture, then Differential reduces the friction. Here are our thoughts on this: https://docs.differential.dev/advanced/soa/ https://docs.differential.dev/advanced/soa/ > "Centralization" and "reliability" don't sound like words that often come together. What if the control plane goes down? Unless they have a cluster? Yes, there's a cluster.
- hmeh 3y agoThe problem though is that you're taking the known failure mode interpretation of microservices (what we call a distributed monolith) and making it "easier" by adding significantly more complexity. As many of the other comments here have said, this has been done before. You are repeating the mistakes of DCOM and gussying it up with a modern-looking marketing page. You demonstrate a lack of understanding of idempotence on your marketing page, which is enough to dismiss your service outright. With all due respect, I'm sure you've put a ton of work into this, and you may even manage to sell it to some teams, but this has been tried and failed so many times in software development's history. It won't be any different now. This is very likely to cause long term harm to any team that adopts it. For anyone considering it, please make sure you actually understand idempotence and autonomy. Look into event sourcing. Study the fallacies of distributed computing. Check out the work of Udi Dahan and Scott Bellware. Run, do not walk from anything purporting to do the things that this does.
- danlugo92 3y agoREST is enough
- lunarcave 3y agoCan definitely appreciate this sentiment, and it's one we talk about [1]. [1] https://docs.differential.dev/advanced/comparisons/#comparisons-with-http--rest-apis https://docs.differential.dev/advanced/comparisons/#comparis...
- novoreorx 3y agoIs it suitable to use in the browser-server architecture? I think the documentation lacks scenarios or use cases, which should also be included on the home page.
- lunarcave 3y agoIt is theoretically possible to be used in a browser-server setting, but it's not something we're optimising for at the moment. Thanks for the feedback on the use-cases. Will work that into the website. In the meantime, here's a write up on our value prop: https://docs.differential.dev/getting-started/thinking/ https://docs.differential.dev/getting-started/thinking/
- lunarcave 3y agoOne of the authors here. Ok, this blew up, and I appreciate the engagement. We will attempt to get to all the questions. We are still polishing the initial offering, so if the documentation seems lacking, it's because it very much is so. Honestly, we did not anticipate having this many eyes on the project this early.
- syngrog66 3y agophone rings "Sir, its The Fallacies of Distributed Computing for you, on line 1!"
- lunarcave 3y agoI'm more than happy to engage you on the fallacies [1] that you point to and how they are considered in the context of Differential. But as your earlier comment would show, I don't think you're engaging in a good-faith conversation. So, I'm happy to leave it here and wish you a good day. [1] https://en.wikipedia.org/wiki/Fallacies_of_distributed_computing https://en.wikipedia.org/wiki/Fallacies_of_distributed_compu...
- syngrog66 3y agoThe feature pitch is one long case of loltears. I needed the chuckle "Kids , get off my lawn."
- emilehere 3y agoI'm curious if people's qualms around abstracting the cloud/network also apply to the https://modal.com https://modal.com product. Differential seems like a similar project at its core, just focused on the Typescript + microservices ecosystem.