7 ms·
The Valley of Webhooks
- hungryhobbit 2mo agoDude is not wrong ... but good luck convincing the Internet to switch to a sane system, when everyone already thinks web hooks are a "solved problem".
- lobofta 2mo agoI can't imagine any reasonable engineer thinking webhooks are a solved problem. Imagine how much time and effort those provider must spent to offer high quality webhooks, to then still get customers complaining about missing packets and such. Webhooks are an operational nightmare.
- Terr_ 2mo agoThe end here reminds me of "The Log: Real-time data's unifying abstraction" [0], which has unfortunately had a bit of link-rot since 2013. One complication in this approach involves access-windows: What if my system is only supposed to be seeing stuff that happened during two separate weeks in the year, because those are the spans when it was subscribed or authorized? So the data-host would need to maintain a concept of "connection history" for other services, and also use that to filter/modify its real event stream, inserting artificial "initial state" roll-ups of events that happened in dark periods. [0] https://news.ycombinator.com/item?id=6916557 https://news.ycombinator.com/item?id=6916557
- zbentley 2mo agoThat requirement seems ... uncommon. How many systems/industries have a frequent notion of transient/sliced views of history for their customers? If that is a real requirement, it seems like it'd be easier to meet by giving customers a realtime-stream/log API whose history starts when they were most recently granted access, and providing them older historical events via a separate API of the classic "ask for a report and we'll get back to you within a day or two with an S3 presigned URL" variety, then synthesizing that huge historical report in batch code that's aware of the subtleties of the customer's visibility windows.
- Terr_ 2mo agoPerhaps less a requirement and more a kind of consistency with prior limitations? For example: 1. I have events in a Calendar service. 2. I want to authorize Reminder service to see upcoming events, so that it can send reminders to attendees in a way the Calendar service does not directly support, e.g. SMS/WhatsApp. 3. With the necessary credentials/SSO, the Reminder service subscribes to Calendar and Calendar periodically POSTSs webhook updates. Reminder needs to recognize when an event is cancelled or rescheduled, so that can alter its reminders. Do I want to give Reminder potential access to all events ever, or just ones active across the usage period? Meanwhile, the Reminder guys probably don't want to step through the whole Calendar-wide event stream to reconstruct which events haven't finally happened yet.
- zbentley 2mo agoFair, and thanks for explaining. Other than my original proposal of a simple webhook stream and a complicated offline backfill, all the good ways to do that with current tools probably involve building support for fine-grained read permissions (Reminder service's windowing) onto the producer (Calendar) side and emitting a tailored reminder stream. If that challenge is indeed common, then there might be a market/demand/opportunity to implement ... I don't know what to call it, "Postgres RLS for Kafka" or something--a record-level security model for event streams, basically. Event compaction/deduplication would make this hard, though.
- gnat 2mo agohttps://archive.ph/oIfml https://archive.ph/oIfml swerves the bitrot
- zrail 2mo agoWebhooks are a painful problem. To clarify, Stripe's events API definitely ships a cursor and polling it has been the method preferred by large consumers for a long time.
- weli 2mo agoStripe events API is one of the examples of how to do things properly. And SCROLL is just trying to create a common spec so that everyone offers a stripe-like event polling api.
- seandoe 2mo agoYea I had the same thought: "It sounds like the Stripe events API would fulfill his needs." But then you said > Nobody serves this today. That’s the catch, and it’s also the point. and I was left wondering what was missing.
- zffr 2mo agoWith webhooks, consumers get to asynchronously respond to updates from a provider. If no data has changed, a provider will not send any updates. With SCROLL, consumers are responsible for choosing when to ask a provider for updates. Without a mechanism for knowing when data has changed, consumers will be forced to be pessimistic and poll providers for new data on some cadence. I see two issues with the proposal: (1) SCROLL will lead to an increase in unnecessary network traffic for both the consumer and provider, and (2) because a consumer cannot know when data has changed, the lag between a consumer's local model and the provider's data model will be larger when with Webhooks.
- lxgr 2mo agoAssuming you're not using the proposed streaming option, I suppose you could always send a webhook for that fact alone? In other words, an empty notification, with semantics of "something has probably changed, better poll the SCROLL if you aren't already".
- Multicomp 2mo agoDoesn't doing that just reinvent eTags on hypertext resources from the RESTful wars and XML Web Services days 20 years ago? 1. Long Poll the cursor to pull down the latest events 2. Trigger a long-poll even if in exponential backoff because they shot you a webhook saying 'eTag changed!'
- inigyou 2mo agoIf your API is just wrapping Kafka, it can long-poll
- cobbzilla 2mo agoThe proposed “feed” solution is functionally indistinguishable from incremental reconciliation. The article presents a good framing and is well-written, but doesn’t really propose anything new.
- cadamsdotcom 2mo agoYes, and I think that's good news. If what you need is log replication, having a standard means it's more likely providers offer an endpoint that conforms - and covers the common pitfalls (tombstones, what to do if cursors can't be sequential, etc) - because at scale any missed requirement can be a showstopper and send you back to emulating log replication by listening to out-of-order events.
- stymaar 2mo ago> and is well-written Nah, it's almost pure slop with just the em-dashes edited away.
- qlkzy 2mo agoThis topic always surprises me. I do not understand the sequence of logic that leads people to build synchronisation mechanisms based only on webhooks. Webhooks aren't at-least-once, nor at-most-once, nor are they guaranteed in-order. Some people build systems to make them more reliable, but if you really care about the data you need to think of a webhook delivery as best-effort, a bit like UDP. That's before you get into all the extra complexities around these systems being owned by different people. For example, either or both system might have to roll back their database. Or either side might have a long-term bug in how they process webhooks, and now you have months of broken data. My view is that the only reasonable thing is to start with the process that gets things back into sync if everything is broken. That almost certainly involves a poll or query of at least the upstream side, and maybe both sides. I find that if you put a decent bit of engineering effort into that "disaster recovery" synchronisation, it can often act as the main or only synchronisation process for quite a lot of systems. Stepping up from that, it's often useful to introduce webhooks as notifications only; that is, to provide a signal that some or all of the data is stale. You have to do a bit of consolidation, but this approach is usually enough to get completely reasonable latency for the kind of applications the author is describing. Only if that wasn't enough for speed/scale reasons would I reach for a truly "push-driven" fast path. But you always have to be able to disaster recovery assuming the stream is wildly out of sync. Some bits of the author's idea seem reasonable: certainly, I would love for there to be a standard protocol to request new data since some cursor or since some timestamp, ideally with some webhook notifications to give hints on when to poll. The problem I have with the author's idea is that it is very strongly event-based, but the desired outcome isn't event-based. The desired outcome is almost always "the state over here looks like the state over there". Relying too strongly events ends up at the same kind of problem another level down: the "disaster recovery" script ends up wanting to compare the states anyway to figure out whether the events are broken. Going fully event-sourced can work (although, I think, less often than advertised), but it really relies on everyone collectively agreeing on the same event stream being the source of truth. Once you start doing work across multiple organisations then that coordination is relatively rare. What really surprises me is the variation in maturity on this topic. There seem to be people at all experience levels who are both doing this well and doing it badly. I have worked with people with decades of experience whose whole design just collapses if you ask "but what if X?" for some really banal values of X like "we have an outage for more than five minutes" or "we have to restore the DB to yesterday" or "someone, one time, accidentally merges a bug into master". As an aside, I do find the obvious LLM-ness of the blog post and the proposal a bit disheartening. These are problems that require diligence and precision of thought. LLMs may be able to achieve those things, but that level of quality just isn't expressible in "Claudish".
- tasn 2mo agoWebhooks are simple and ubiquitous, and that's both a weakness and a strength. It's also why they are used for a lot of things, even things they are not great for (state sync). These weaknesses are why we[1] added FIFO endpoints, Polling Endpoints, and what we call "Svix Stream" as ways to do ordered state synchronization (each with its own tradeoffs). This lets people consume the events in the way that best fits their use-case. We are working on more things to make the state sync even easier. I'd love to hear about more challenges people are facing with webhooks, as we want to make these things better. OP: I'd love to hear more about your thoughts there, and will send you an email in a moment. P.S, if you're unfamiliar, please check out Standard Webhooks[2]. It's a spec we created to help with signature verification that has been adopted by OpenAI, Anthropic, Google, and many others. We are chipping at one webhook challenge at a time. :) 1: I'm the founder of Svix (mentioned in the post), we do webhooks infrastructure as a service. 2: https://www.standardwebhooks.com/ https://www.standardwebhooks.com/
- lubujackson 2mo agoNice write-up of the webhook data consistency problem - I've been bitten by some of those Stripe issues myself. I agree that webhooks should come hand-in-hand with a "records updated since" endpoint that allows high-level discrepancy checks and API endpoints to pull full logs for any record. There needs to be some sort of reconciliation process that is both fully verifiable and not a firehose of data.
- alt227 2mo agoI had the exact same thing with the Quickbooks api recently. You cannot trust the responses or webhooks at all. On create a user or invoice for example sometimes it will return an error, yet it actually created the entity. This means you have to check manually after creating everything to know if its created properly. Then you have the issue that sometimes quickbooks takes a while to update, and locks the company file while it does some background magic. This means you cannot immediately do the existence check, and also sometimes the check errors or times out which essentially means you need to keep checking forever until you can properly reconcile your db against theirs. But with hundreds/thousands of transactions per minute this state is never reached. You perpetually live in a state of trying to catch up but never managing it. When I brought it up with Quickbooks dev support their response was literally "Its your job to make sure things are created properly in our system". How did we get to this place where we started putting up with systems that cannot ever be trusted?
- hyperhello 2mo agoThat is the way enterprise software works as a system. It demands to be the central focus of everything. Workers want to route around these turbo productivity theater nonsense that could be replaced by a few K script that gates access to a text file and checks validity of appends. That can’t be allowed, so you need what is essentially whole poorly documented OSs to enable an economy of brokers to it, or the whole con would collapse.
- lelanthran 2mo ago> You cannot trust the responses or webhooks at all. Well... yeah. I mean it's pretty obvious, no? Here's some things that could go wrong regardless of what care the software tries to provide: - The transaction completed on the backend cluster but the app instance died before if could create the response and after it committed the transaction. - The transaction completed, the app instance transmitted a response, but the load-balancer/reverse-proxy in-between died before it could relay that response. - Everything went well, but the ISP dropped some packets before it could get to you. - Everything completed and the ISP stayed up, but on your end the response was flagged as malicious, or never made it through your load-balancer. So, yeah. in general when you make an API request and get an error you have to check if the state was changed anyway, and if you aren't doing that you're doing it wrong anyway and cannot blame the system on the other side for returning errors.
- ninju 2mo ago> Ask with a cursor and you resume where you left off. How does the provider know what event the cursor you provided refers to? Sounds like external state that needs to be managed ("cursor" -> timestamp)
- weli 2mo agoThat is completely fine. In the spec cursors just need to be lexicographically comparable. They could easily be timestamps if the provider chooses to. In many of my example they are just alphanumeric dictionaries.
- sandeepkd 2mo agoThis is really a great article from implementation perspective and the challenges associated with it. The OP already captured the reasons. Even though the reason 1 has been mentioned it did not go deep into it and somehow focussed a lot more on 2 > 1. Trigger a side effect: send the receipt, start the build, ping the channel. > 2. Keep a copy of the provider’s data correct: On high level 1. System can either be PUSH or PULL, webhooks are essentially push and towards the end the OP is exploring the possibility with PULL. The caveat is that OP already iterated the PUSH mechanisms thrice and is aware of all the hardships and is somehow hoping that PULL would solve them. Unfortunately the grass is same on other side too. a) The availability of the server can always be questionable in PULL mechanisms and its a lot of load on servers to support this kind of data at scale in bulk to multiple customers. You are essentially getting into database table scans. Its becomes a lot more costly with NOSQL databases. b) The customer would end up making way too many calls to server even if data is not available or there would be additional latency when data was updated and when it was queried. This is one of the reasons why servers prefer to push instead of pull if they can find a listener available on other side. c) CRLs (Certificate revocation lists) are good example which are available for PULL, same for all clients and yet rarely anyone does it correctly or does it at all even though its in security domain. In fact they are simple files on webservers in most of the implementations. 2. The primary use case for Webhook is for triggering the side effect and allowing the customers to choose if they want to subscribe for that event. A customer subscribing for everything even if its non-actionable should just treat it as logging data. 3. Logging data can and always have gaps, it should never be treated as source of truth. I might question the need for deduplication, usually there is a unique identifier and almost all databases support insert ignore kind of clause. Logging the event data just provides you with better availability and latency, the source of truth is still with the provider if a next step needs to happen. 4. If user cancelled the subscription in stripe and the event never arrived then its a system design issue or system availability issue on the client side. The complete data checksum or bulk imports at night are attempt to fix the problem in a hammerhead way . I understand it exists in lot of places, however it defeats the whole purpose.
- Elucalidavah 2mo ago> its a system design issue or system availability issue on the client side If all events need to arrive, then the problem is not "notification" (which would be solved by webhooks) but "database replication": subscribe to new events, fetch the full snapshot, fetch the updates in range, have the monotonic value to establish the "range" in the first place. Reach the eventual consistency. The proposed SCROLL handles half of these, which limits its use-cases.
- __MatrixMan__ 2mo agoIf you have two or more services that need to agree about state, and you have some set of rules that govern what state changes are valid, and you don't want to mess around with any of this "what do we do when we miss and update vs when we get two of the same update" nonsense, and the services aren't in a position to query the same database, then you should really consider a permissioned blockchain. Consensus hard, but it's harder if you're not using tools that understand that what they're going for is consensus.
- Xirdus 2mo ago99% of the time (and all 3 times in the blog post), there is only one source of truth for any piece of data, and state transitions are completely arbitrary. Blockchain is almost always the wrong solution.
- Terr_ 2mo ago> Blockchain is almost always the wrong solution. Especially since in most of the cases where it's not-totally-insane to use, the right solution is still the classic distributed database which already existed. In those, the ledger is kept among a predefined/controlled node-membership... as opposed to a bloated mass of workarounds and limitations to make it barely survive being ungovernable. I've seen some boosters pivot to saying "private blockchain is good", but that's contradictory buzzword nonsense. It's like selling a blog as "single-user Twitter" or advertising a regular car as "user-controlled autonomous vehicle."
- __MatrixMan__ 2mo agoI'm not sure what a private blockchain is, but a permissioned blockchain is one where only certain parties have keys that allow them to write blocks. You end up with a message queue optimized to eliminate anything that would lead to the inconsistency nonsense that this article is talking about as soon as it is detected. It then becomes the writer's problem to retransmit in a way that doesn't cause a problem next time, rather than the reader's problem to recover. The craziness comes in when you'll accept blocks from anybody willing to burn enough electricity to do so, or gamble enough tokens to do so, or whatever other artificial scarcity game people like to play. But if you're only planning to consume data from the eight other companies you do business with then there's no reason to bother with any of that, you can just hard code their public keys into your consensus protocol and you've sidestepped the nonsense.
- thingification 2mo agoI wonder if this is a CS problem somebody solved in 1954. Does somebody have the link to that paper? (I'm not serious about 1954 in particular, I am about hoping somebody here knows the CS literature better than me)
- rawgabbit 2mo agoThe old timey systems modeled this problem using accounting 101. When things change, you don’t immediately update the balance. Instead it is written to a transaction journal aka a log. The thing is this log is the source of truth. State or the balance is derived from the log. You don’t send a continuous stream of logs. Instead it is batched and sent asynchronously. It is also applied asynchronously. It also records if the batch was successful or not. If you have multiple systems sending their logs to a central server. No problem. The central server orders them all before applying the batches. Every so often. The books are “closed”. Meaning the central server won’t accept any more journal entries for things that happened older than X dates.
- zrail 2mo agoI dunno about "old timey", I wrote a system that does exactly this like six months ago. Granted, it's a money tracking system. It's always surprising to me how much of the real world runs on CSV and EDI files sent back and forth over SFTP.
- gregw2 2mo agoI'm not an academic but I've worked through practical problems in this space for longer than I'd like to admit, often in ignorance of the literature. The most CS-y fundamental paper would probably be Lamport(1978) below, but the database papers are pretty fundamental on this topic in their own right. Here are some pointers: * Astrahan et al., "System R: Relational Approach to Database Management" (1976), oldest paper talking about logs in a database which is kinda what at its root this is recapitulating * Jim Gray, "Notes on Data Base Operating Systems" (1978), covers logs with the goal of transaction management, undo, redo, the famous "two phase commmit" process for aligning state across a network boundary in two systems * Leslie Lamport, "Time, Clocks, and the Ordering of Events in a Distributed System" (1978), not a database paper, more about state synchronization in general * Bruce Lindsey, "Notes on Distributed Databases" (1979), standards for replicating data across multiple identical database nodes * Jim Grey et al, "The Recovery Manager of the System R Data Manager" (1981), describes "write to the log first, then the database (WAL)" pattern * Silberschatz, Stonebraker, Ullman: "Database Systems: Achievements and Opportunities"/"A Architecture for Heterogeneous Database Replication" (1980), discusses tricky bits about replaying logs in client database systems that work differently from the producer * Leslie Lamport, "The Byzantine Generals Problem" (1982), on the math and needed consensus error handling when distributing state For other later topics on state synchronization, read up on Paxos (Lamport, 1989/98), RAFT (2013), CRDTs (2011), etc. The above approaches of "write-to-the-log-then-update-state" were applied to UNIX filesystem first in AIX 3.1 in 1990 and then adopted by other UNIX vendors and then in Linux ReiserFS/ext3/XFS(SGI) in 2001ish. Lotus Notes took a different path in the 1990s to synchronizing state within documents which was not ACID-oriented like the above, but was more like the modern append-only-log with optimistic eventual consistency. Not sure about the best paper on this. Then Martin Fowler popularized event sourcing in 2005 with Enterprise Application Architecture, and later described CQRS in 2011 and microservices. So "we" decoupled everything with Webhooks (2007) and Kafka (2011) and reinvented this problem. Oh, and did I mention Blockchain (2008), etc. Oh, did I mention Bitkeeper dvcs (1998) and git (2005) handling of distributed state? Along the way databases (Snowflake, BigQuery, Iceberg, etc) started exploiting logs to show "older" state via features like "Time Travel" queries. Which is actually what Stonebreaker tried to do in the earliest versions of PostgreSQL but the first implementation in the 80s didn't perform well without good compaction support and on much more constrained resources. Ask your local LLM for a good survey paper and it might be a bit easier to digest, but the above gives you some color and keywords.
- russellbeattie 2mo ago> Then the dedup table, because deliveries arrive twice and the docs cheerfully call this “at-least-once.” Documenting unexpected or intermittent behavior: The easiest bug fixes of all.
- WorldMaker 2mo agoThe proposed feed solution looks a lot like the CouchDB replication protocol, as another object in a convergent evolution space to consider.
- toomim 2mo agoThis is a nice writeup of the problems in using Webhooks for State Synchronization. I also noticed that the proposed solution is a pseudo IETF-style draft protocol called SCROLL... that happens to be remarkably similar to an actual IETF draft I am bringing to IETF 127 this November called "Braid-HTTP Subscriptions." Both drafts request a subscription with a GET plus a header: Scroll Request: GET /scroll/feed/customers Prefer: stream Braid Request: GET /customers Subscribe: In both systems, the GET leaves its response open to stream events. SCROLL responds with application/x-ndjson. Braid subscriptions are a 209 Multiresponse, with content-type application/http-history. This lets them support more than just JSON. You can send updates to the state of CSV, or PNGs, XML, HTML, plain text, or any media type. The author noted that it's hard to get adoption. Well, the reason that Webhooks are so common is that they are bog-standard HTTP. For this to get adopted, we need to put it into bog-standard HTTP. So we need to go to the IETF, and and extend HTTP in a general way to support state synchronization. It should just work for any existing HTTP media type (not just JSON), and any resource/URL (not just special /scroll/* URLs), and any way of marking timestamps (not just the ordered strings proposed in SCROLL). Then we can bake this stuff into HTTP, and thus into all our bog-standard libraries, utilities, and code, and you won't have to reimplement the same sync-logic-over-webhooks again, and again, and again. Reach out if you're interested!
- weli 2mo agoI reached out through email :) Just one correction. My spec doesn't force /scroll/ URL's, just proposes it as a convention.
- bobbiechen 2mo agoIs it accurate to say this is something like long polling except you continue to hold the connection open for subsequent updates? Does this mean a server potentially needs to hold open a very large number of connections (one per client) even if there are no updates? And why formalize on HTTP rather than on a similar protocol over websockets?
- anamexis 2mo ago
- shreygupta 2mo agoGerard mentions it super quickly, but another massive issue with webhooks generally is local development. Yes, you can use a tunnel, but that requires all engineers on a team to add their own tunnel urls. This causes even more issues when you use the platform as a source of truth, like for auth or payments. With WorkOS specifically, your whole team develops with one shared development sandbox. You run into issues when your local dev auth (in postgres) is not synced with the shared dev sandbox that WorkOS has since not all team members have their dev environments running at once. So yeah, then you use events API. But WorkOS only preserves the events API data for 90 days (and u have make 3 calls since its a max of 30 days per call). So then you load all the data with the state API first, then you start running the events API. It's a mess. Tried to talk about this on X until the CEO of WorkOS wanted to bring it in private, then proceeded not to help at all. https://x.com/grinich/status/1913035839866835297?s=20 https://x.com/grinich/status/1913035839866835297?s=20
- pphysch 2mo agoThe core tradeoff here is "Push vs. Pull". Webhooks are a model for pushing data to subscribers. A traditional API allows clients to request and pull data. The data flow is in the same direction, but control flow is opposite. Pull-oriented models are much easier to reason about and should be preferred where possible (cybernetically they are a closed loop, vs. push models which could literally just be a barrage of UDP packets). But they do have a little bit of overhead which makes them the wrong tool for some cases, like live-streamed entertainment or massive telemetry flows which value performance (latency, throughput) over missing a few packets.
- foresterre 2mo agoI had this almost exact discussion today. Adyen (the payment provider) provides merchants with webhooks so they can update a local modal of payment and payment modification data (checkout). But there is no way (for a merchant) to get the latest 'true' state as held by Adyen. So you better hope your data is exactly in sync with the notifications you got from the webhook (which it never exactly is, because there are so so many points of failures, and unlike what this author says, the docs aren't thát well presented to hold the same model as the PSP does. It is often close enough though, but you are constantly gardening your implementation, because the model also changes on their end with little information in the changelogs). The "latest state" data exists though! If you open the customer portal it is presented to you without problem.
- bytesandbots 2mo agoWith the proposed solution every consumer will have a persistent connection to the server irrespective of the frequency of events. This setup seems inefficient unless you have a very high volume of events coming in. Many CDN networks have a limit on how long a connection you can open. And data providers will not prefer serving persistent requests. Problems listed are signatures, dedup, buffering, bootstrap, cron. Everything other than signatures and bootstrap, can be solved by having a counter in every webhook payload. It will increment each time. When you receive a webhook and the counter does not match, the consumer can fetch the missing data from the events API. I agree with the author that providers simply saying "at least once delivery" is insufficient. they should have solutions that does not require an architecture diagram. Bootstrap is better served with a bulk events API so you don't make one call per request. It can have an after/cursor pagination. Solutions that work for our internal Kafka might not be suited to work across services, over the internet.
- oasisbob 2mo agoOpen TCP connections can also be wildly cheap and efficient - Apple Push Notifications (APNS) and Android's push systems maintain open TCP connections to just about every mobile device on this planet. An open connection is just a bit of state on either end. The C10K problem has been solved for ages. Anyone remember consuming Twitter hoses back in the day? Those were also long-lived persistent connections for efficiency reasons.
- weli 2mo agoNot only this but the spec explictly allows provider-terminated connections followed by a 429 with a respected "Retry-After" if you want to definitely kick "slow" feeds. This combined with HTTP2/3 multiplexing makes holding feeds open pretty cheap.
- zbentley 2mo agoThis article resonates with similar struggles I've had at multiple companies. Its content is valuable. I wish the writing and the article itself had less of a slop odor. That aside, what I don't understand (especially having worked on the side of the webhook sender, which is itself really tricky to get correct/performant/cheap) is why more companies which broadcast webhooks don't, say, provide direct access to Kafka topics, S3 buckets with ordered data objects landing, SQS queues, or any of the alternatives to those things. "But it's irresponsible to expose an internal-use-only datastore directly to clients" goes one objection. But plenty of log-store systems have the notion of sharing a subpart of the log with a less-than-trusted external peer, so while exposing Kafka directly might be asking for the same kind of trouble as exposing your customer's SQL database for authenticated connection over the open internet (e.g. "we said you could issue reads, not that you could open/close TCP connections a million times a second! You just took out our message broker!"), exposing, say, an S3 bucket or Kinesis stream is much less risky because those systems have put some thought towards semi-trusted sharing. "But everyone is used to getting HTTP webhooks and doesn't have the expertise to connect to something else"--that'd be true if, say, reading from a websocket or Postgres NOTIFY stream or Kafka topic or S3-change-notification stream were advanced techniques, but libraries around those things are so good nowadays that even the most web-tech-only low-skill developer can probably integrate with them with minimal hassle. Maybe it's just that a lot of shops literally only know how to run their code in a webserver, and have never deployed any other kind of application service/cronjob/queue worker? That seems unlikely to me, but I might be surprised. "If we do something weird our competition will beat us on ease-of-use" goes another objection. But is it that hard given the libraries available? And can't you hedge back on the ease-of-use sell with "our data is fresher and more provably ordered and correct"? I'm glad that SCROLL exists as a possible solution here. I'm just puzzled why more people haven't been using existing technologies to achieve this property. Do most webhook senders literally not have a log store? Are they just firing webhooks in the middle of business event handlers and giving up synchronously if they can't be delievered? Because if that's not the case (and I don't think it's the case), then it seems like the SCROLL API is ... basically just the Kafka consumer API. Or Kinesis. Or SQS. And so on.
- tasn 2mo agoWe (Svix, webhooks infra) actually help our customers directly write to Kafka topics, S3 buckets, SQS, etc. and have for a few years now. There are definitely people that adopt that, but receivers as well prefer the simplicity of webhooks.
- delusional 2mo agoYou don't really need a new protocol. If you trust your consumers, you can just give them a paginated "/events" endpoint and have them poll that. The delay will be variable, but tunable, based on the required timeliness. If you want more prompt responses you do long-polling or a websocket. The key, and only thing that matters, is that the cursor rides in your database, and is therefore transactionally consistent with the event. That's the whole magic trick. We've done event streams like this at the bank I work at for years.
- zie 2mo agoSeems like we are trying to solve 2 problems at the same time with SCROLL: * Database sync of log events. * "Real" time updates. Webhooks can already mostly handle the "real-time" event portion. Of course there are problems. as the article expounds on, but a lot of those won't magically get solved with other solutions either. Distributed real-time communication is hard. Webhooks are good enough for this purpose. For the database of log events, personally I'd rather just have a SQLite DB I can yank whenever. Don't give me CSV or JSON or whatever I have to parse and manage, just give me a SQLite DB ready to go. I'd love you for it. I'll just take a whole fresh copy with everything thanks. Maybe you limit it to to the last X events, say 90 days or 365 days or whatever, depending on sizing of events, but just send it all every time I fetch and I'm happy enough. If I need to generate a delta to keep some other DB in sync, well that's my problem. Just give every row a stable identifier.
- cyberax 2mo agoJust treat webhooks as a hint. Write your code to work as a "reconciler" that checks the state of the remote system and reconciles it with the local view of that system. Run reconciliation for the full state periodically using a scheduler, and then treat webhooks as a hint to run the reconciler immediately. This way, you will have a robust system that can survive logical bugs and outages because you don't store the synchronization state per se. In the case of Stripe, for example, have a process that polls every open checkout session every couple of minutes. A webhook then just triggers the run earlier. If you're worried about DDoS, have an exponential backoff for the poll period. Theoretically polling doesn't scale, but in practice it works just fine.
- tlonny 2mo agoI much prefer cursor paginated API requests vs. webhooks. The obvious downside being that in order to not get 429'd you need a respectable poll frequency - meaning you lose reactivity to new events. Thus I think webhooks still have a place - but as a simple "poke" that can be sent to the client to tell them something has changed - supplementing a default low frequency polling interval. This gives us the best of both worlds: 1. No need to bother de-duping/retrying pokes - if you miss a webhook you will shortly recover anyway when you next poll. 2. No need for any local-specific tunnelling/tooling - the local app will work just fine with the default poll interval. 3. No need to keep a connection live for each client. 4. All the good stuff OP mentioned in his blog post.
- throwaway7783 2mo agoAbsolutely. I have been so livid at so many applications for not providing a decent CDC API (and dont forget deletes). Salesforce perhaps is the best out there. Imagine if every application exposed a standard CDC API, the world of integrations would be so much better. Webhooks are fine, but a pollable API is a must have. The amount of hacks I had to do at work to workaround shitty APIs gives me nightmares.
- AgentME 2mo agoThe Gmail API works nicely like this. There's a history.list endpoint where you can see the recent history of message additions and removals, you can query for just the history that's taken place since a specific event's historyId, and you can subscribe to push notifications (that can be delivered by webhook) which just tell you when there are new history events, and you're expected to hit the history.list endpoint to see what's new. Some dropped push notifications aren't a big deal.
- evolve-maz 2mo agoYep, the poke pattern is additionally nice because it means my state reconciliation function is the same when running on cron interval vs event driven. On the flip side, it helps to have endpoints which have a query param linking to some sort of resource update time stamp. That way you can query to only get those items changed since last poll.
- stymaar 2mo agoThe text ought to have been interesting, as the insight is real, but for god sake how can someone decently read the slop their LLM outputted as reply to their prompt “write a blog post about X” and just copy-paste as it is?!
- sicromoft 2mo agoCan we ban AI slop posts?
- kzmttkc 2mo ago[flagged]
- bobtheborg 2mo ago"Nobody serves this today" except for every real estate listing service (https://www.reso.org/reso-web-api/ https://www.reso.org/reso-web-api/)
- aurumflux20 2mo ago[flagged]
- ZenithBar 2mo ago[flagged]
- nektro 2mo agoi think a nice middle ground would be the "flip the arrow" but the webhook tells you when the cursor has moved and that's it so you don't hammer the origin server will polls or have N-time stale data.
- jmaw 2mo agoI agree that this could be a good middle ground. This causes the webhook to be a push notification with the eventId/cursor, that then causes the consumer to pull the actual data (along with other recent events if desired).
- abrookewood 2mo agoCould be completely off base, but isn't this the equivalent of polling a Provider's commit log or distributed log? So accessing a third party's Apache Kafka?
- mrkeen 2mo agoYep! And since your partner will be unlikely to build a new architecture on the request of one customer, you should 'do the Kafka thing' on your end, capture everything (double-sends and all), then smartly query your events, rather than try to program correct double-send-friendly db updates.
- ksymph 2mo agoThis is kind of beside the point but... Shouldn't the the evolutionary metaphor be the opposite of how the author's using it? They go on about how we've settled in the trough, but the whole local optimum problem is about settling in an early peak, and being unwilling to cross a trough to get to a higher peak. The workarounds would be propping up the first peak, not evidence that we've settled in the trough. Either way, the metaphor feels kinda forced IMO.
- marklar423 2mo agoThis is a great article and exactly mirrors my experience consuming webhooks. My team built almost the same workarounds as the author. The solution my team eventually settled on to the dedup, buffering, and race problems was: * When a webhook came in, store a single copy of the payload (usually json) in some temporary storage and enqueue the id. * Any subsequent updates (while the id was enqueued) overwrote that payload and skipped the queue. So, we were able to ensure we only updated the object once and with the latest state (since the webhook payload always held the entire object state). We didn't have a great solution for the bootstrap problem though. I do like the SCROLL proposal, but I wonder about the cost of keeping around log-structured data forever - every log structured DB I know about does compactions for this reason.
- weli 2mo agoCompactions are supported and heavily recommended! https://github.com/welidev/scroll/issues/1 https://github.com/welidev/scroll/issues/1
- mrkeen 2mo agoI usually preach that a big blob of mutable state in the middle of your system (the db) is the problem, and events are the solution, and this seems like its most obvious case. All your problems start when you try to modify your current state representation in response to hearing the USER_ADDED or USER_BLOCKED events. Just don't. Leave them as events. Any time you have a new stupid edge case (double send, out-of-order, add-then-delete-then-add), this becomes one new unit test, where you can soberly decide what it means, and update your read path to understand it. If you bake the nonsense into the current db state every time a new event arrives, your first step in any debug or reasoning scenario is to unbake it: what events led to this mess? If you don't throw away the events, your debugging is done for you.
- emmanueltsakpo 2mo ago[flagged]
- NortySpock 2mo agoYeah, been fighting variants of this problem as a data engineer where the upstream database would only let you see the current state of the table and did not provide you a changelog / cdc So you have to (a) repeatedly query incremental date ranges, with some overlap (b) deal with late arriving rows (c) deal with occasional bugs where a bug in a sproc caused the timestamp not to update (d) query each day for all primary keys and remove rows that had been deleted (but only if you have newer data, otherwise you delete old values but don't have the updates) (e) Sweat about if this goes wrong on a large table, how can you rapidly determine what state is out of sync? (Idea: how do I "hash" this in a way that hints at where the problem is?) (f) Deduplicate, as mentioned Martin Kleppmann's "Designing Data-Intensive Applications" book has some detailed discussions of some of the challenges...