13 ms·
Sync Engines Are the Future
- paduc 2y agoBefore I write anything to the DB, I validate with business logic. Should I write this logic in the DB itself ? Seems impractical.
- scotty79 2y agoI think that's the main issue. It's not enough to have a database that can automatically sync between frontend and backend. It would also need to be complex enough to keep some logic just on the backend (because you don't want to reveal it and entrust adherence to the client) and reject some changes done on frontend if they are invalid. Database would become the app itself.
- acac10 2y agoWhich many DBs allow: - stored procedures - Oracle PL/SQL I used to work for Oracle but never liked that approach.
- Sammi 2y agoThe issue with stored procedures is testing and code maintenance. How do I run unit tests? How do I version control and code review?
- TeMPOraL 2y agoIt's the same issue that killed the image-based programming in favor of edit-compile-run cycle we're all doing. "How do I test? How do I do version control? How do I migrate?". These are valid concerns, but $deity I wish we focused on finding solutions for them, because the current paradigm of edit/compile/run + plaintext single source of truth codebase, is already severely limiting our ability to build and maintain complex software.
- brulard 2y agoWhile I don't like the idea of putting logic to the DBRMS (if not for a really good reason), you can do unit tests and code reviews. In a serious project you already should have a way to make migrations and versioning of the DB itself (for example using prisma, drizzle, etc.). Procedures would be just another entry in the migrations and unit tests can create testing temporary DB, run the procedures and compare the results. I agree tooling is (AFAIK) not good and there will be much more work around that, but it is possible.
- x0x0 2y agoThe other issue, from experience, is needing to reimplement logic as well -- you end up with stored procedures that duplicate logic that also must be run either in your server or on your client. eg given the state of the system, is this mutation valid. Then those multiple implementations inevitably suffer different bugs and drift, leading to really ugly bugs.
- scotty79 2y agoI don't think a stored procedure that operates only on master copy of the database can reject update comming from a second copy and nicely comminicate thus happened so that the other copy can infrom the user through some ui.
- Terr_ 2y ago> logic in the DB Something similar but in the opposite direction of lessening DB-responsibilities in favor of logic-layer ones: Driving everything from an event log. (Related to CQRS, Event-Sourcing.) It means a bit less focus on "how do I ensure this data-situation never ever ever happens" logic, and a bit more "how shall I model escalation and intervention when weird stuff happens anyway." This isn't as bad as it sounds, because any sufficiently old/large software tends to accrue a bunch of informal tinkering processes anyway. It's what drives the unfortunate popularity of DB rows with a soft-deleted mark (that often require manual tinkering to selectively restore) because somebody always wants a special undo which is never really just one-time-only.
- TeMPOraL 2y ago> Should I write this logic in the DB itself ? Yes? If it sounds impractical, it's because the whole industry got used to not learning databases beyond most basic SQL, and doing everything by hand in application code itself. But given how much of code in most applications is just ad-hoc reimplementation of databases, and then how much of the business logic is tied to data and not application-specific things, I can't help but wonder - maybe a better way would be to treat RDBMS as an application framework and have application itself be a thin UI layer on top? On paper it definitely sounds like grouping concerns better.
- lloeki 2y ago> treat RDBMS as an application framework and have application itself be a thin UI layer on top? Stored procedures have been a thing. I've seen countless apps that had a thin VB UI and a MSSQL backend where most of the logic is implemented. Or, y'know, Access. Or spreadsheets even! And before that AS/400&al. But ORMs came in and the impedance mismatch is then too great. Splitting data wrangling across two completely differing points of views makes it extremely hard to reason about.
- brulard 2y agoWhile stored procedures/triggers etc. can be powerful, it has been taught for decades now that it is an antipattern to put business logic to the RDBMS (for more or less valid reasons). Some concerns I would have would be vendor lock-in and limits of the provided language.
- Tobani 2y agoIn very simple systems that makes sense. But as soon as your validation requires talking to a third party, or you have side effects like sending emails you have to suddenly move all that logic back out. You end up with system that isn't very easy to iterate on.
- Nextgrid 2y agoYou can model external system interactions with tables representing "mailboxes" - so for example if a DB stored procedure needs to call a third-party API to create a resource, it writes a row in the "outbox" table for that API, then application-level code picks that up, makes the API call, parses the response (extracts the required fields) and stores it in an "inbox" table so now the database has access to the response (and a trigger can run the remainder of the business process upon insertion of that row).
- tonsky 2y agoIf you think of an existing database, like Postgres, sure. It’s not very convenient. What I am saying is, in a perfect world, database and server will be the one and run code _and_ data at the same time. There’s really no good reason why they are separated, and it causes a lot of inconveniences right now.
- Tobani 2y agoSure in an ideal world we don't need to worry about resources and everything is easy. There are very good reason why they are separated now. There have been systems like 4th dimension and K that combine them for decades. They're great for systems of a certain size. They do struggle once their workload is heavy enough, and seem to struggle to scale out. Being able to update my application without updating the storage engine reduces the risk. Having standardized backup solutions for my RDBMS means is a whole level of effort I don't have to worry about. Data storage can even be optimized without my application having to be updated.
- onion2k 2y agoIsn't this what CouchDB/PouchDB solves in quite a nice way?
- paul_h 2y agoI always found the documentation lacking and it not 100% clear what was in couchbase (commercial & OSS) vs couchdb and which I really wanted
- fridder 2y agoThat was my first thought! https://couchdb.apache.org/ https://couchdb.apache.org/ is pretty good though is it still the incremental views with JS?
- theamk 2y agoTL/DR: > If your database is smart enough and capable enough, why would you even need a server? Hosted database saves you from the horrors of hosting and lets your data flow freely to the frontend. (this is a blog of one such hosted database provider)
- Sytten 2y agoThat quote is why security people will always be employed. Jokes aside firebase access control is a nightmare and all those database as an APi thing have the same problem.
- SuperNinKenDo 2y agoApropos of the other reply to you about security. Maybe some security people could let me know their thoughts on this. It seems like generally, the best way to expose your database to the internet is considered to be not doing so in the first place, i.e., have your webserver query and cache a hosted database that isn't directly exposed. Is my understanding correct? It seems that almost all data breaches we hear about are directly exposed databases or their cloud equivalents. Is doing this in the era of "cloud" being made impossible?
- worthless-trash 2y agoI don't think most largescale breaches are directly exposed databases, they are just the ones that summon the largest face palms.
- TeMPOraL 2y agoThat's in some sense a "Swiss cheese security model". It's not that databases should, in principle, never be directly exposed. It's that they rarely are designed for it security-wise[0]; meanwhile, adding whatever complex assembly of containers and applications written in random languages and frameworks, to sit between users and the database, introduces a swamp of better-secured systems that attackers also needs to get through. The more cruft you pile on, the more annoying it gets for attackers and users alike. In fact, there are many benefits of directly exposed databases - many of which would remove the need for applications normally sitting on top of those databases, which are strictly inferior and less ergonomic and overall more shitty than a generic database browsing interface. But that's another reason for why things are the way they are: people wanna make money, and having your application be a toll booth between useful data you own and the rest of the world, is tried and true way of making money. -- [0] - Because they're not normally exposed, because they're not designed for it, because... it's a self-reinforcing loop.
- tbrownaw 2y ago> decoupled from the horrors of an unreliable network The first rule of network transparency is: the network is not transparent. > Or: I’ve yet to see a code base that has maintained a separate in-memory index for data they are querying Is boost::multi_index_container no longer a thing? Also there's SQLite with the :memory: database. And this ancient 4gl we use at work has in-memory tables (as in database tables, with typed columns and any number of unique or not indexes) as a basic language feature.
- anonyfox 2y agoIn Elixir/Erlang thats quite common I think, at least I do this for when performance matters. Put the specific subset of commonly used data into a ETS table (= in memory cache, allowing concurrent reads) and have a GenServer (who owns that table) listen to certain database change events to update the data in the table as needed. Helps a lot with high read situations and takes considerable load off the database with probably 1 hour of coding effort if you know what you're doing.
- TeMPOraL 2y ago> Is boost::multi_index_container no longer a thing? Depends on the shop. I haven't seen one in production so far, but I don't doubt some people use it. > Also there's SQLite with the :memory: database. Ah, now that's cheating. I know, because I did that too. I did that because of the realization that half the members I'm stuffing into classes to store my game state are effectively poor man's hand-rolled tables, indices and spatial indices, so why not just use a proper database for this?. > And this ancient 4gl we use at work has in-memory tables (as in database tables, with typed columns and any number of unique or not indexes) as a basic language feature. Which one is this? I've argued in the past that this is a basic feature missing from 4GL languages, and a lot of work in every project is wasted on hand-rolling in-memory databases left and right, without realizing it. It would seem I've missed a language that recognized this fact? (But then, so did most of the industry.)
- phyrex 2y agoABAP, the SAP language has that, if i remember correctly
- ximm 2y ago> have a theory that every major technology shift happened when one part of the stack collapsed with another. If that was true, we would ultimately end up with a single layer. Instead I would say that major shifts happen when we move the boundaries between layers. The author here proposes to replace servers by synced client-side data stores. That is certainly a good idea for some applications, but it also comes with drawbacks. For example, it would be easier to avoid stale data, but it would be harder to enforce permissions.
- szundi 2y ago[dead]
- worthless-trash 2y agoI feel like this is the "serverless" discussion all over again. There was still a server, its just not YOUR server. In this case, there will still be servers, just maybe not something that you need to manage state on. This misnaming creates endless conflict when trying to communicate this with hyper excited management who want to get on the latest trend. Cant wait to be on the meeting and hearing: "We dont need servers when we migrate to client side data stores".
- TeMPOraL 2y agoI think the management isn't hyper excited about naming - in fact, they couldn't care less for what the name means (it's just a buzzword). What they're excited about is what the thing does - which is, turn more capex into opex. With "cloud", we can subscribe to servers instead of owning them. With "serverless", we can subscribe directly to what servers do, without managing servers themselves. Etc.
- Diederich 2y agoRecently, something quite rare happened. I needed to Xerox some paper documents. Well, such actions are rare today, but years ago, it was quite common to Xerox things. Over time, the meaning of the word 'Xerox' changed. More specifically, it gained a new meaning. For a long time, Xerox only referred to a company named in 1961. Some time in the late 60s, it started to be used as a verb, and as I was growing up in the 70s and 80s, the word 'Xerox' was overwhelmingly used in its verb form. Our society decided as a whole that it was ok for the noun Xerox to be used a verb. That's a normal and natural part of language development. As others have noted, management doesn't care whether the serverless thing you want to use is running on servers or not. They care that they don't have to maintain servers themselves. CapEx vs OpEx and all that. I agree that there could be some small hazard with the idea that, if I run my important thing in a 'serverless' fashion, then I don't have to associate all of the problems/challenges/concerns I have with 'servers' to my important thing. It's an abstraction, and all abstractions are leaky. If we're lucky, this abstraction will, on average, leak very little.
- slifin 2y agoI'm surprised to see Tonsky here Mostly because I consider the state of the art on this to be Clojure Electric and he presumably is aware of it at least to some degree but does not mention it
- profstasiak 2y agothank you for mentioning! I have been reading a lot about sync engines and never saw Clojure Electric being mentioned here on HN!
- tonsky 2y agoClojure Electric is different. It’s not really a sync, it’s more of a thin client. It relies of having fast connection to server at all times, and re-fetches everything all the time. They innovation is that they found a really, really ergonomic way to do it
- quotemstr 2y agoClojure Electric is proprietary software, which disqualifies it immediately no matter its other purported benefits
- dustingetz 2y agoElectric’s network state distribution is fully incremental, i’m not sure what you mean by “re-fetches everything all the time” but that is not how i would describe it. If you are referring to virtual scroll over large collections - yes, we use the persistent connection to stream the window of visible records from the server in realtime as the user scrolls, affording approximately realtime virtual scroll over arbitrarily large views (we target collections of size 500-50,000 records and test at 100ms artificial RT latency, my actual prod latency to the Fly edge network is 6ms RT ping), and the Electric client retains in memory precisely the state needed to materialize the current DOM state, no more no less. Which means the client process performance is decoupled from the size of the dataset - which is NOT the case for sync engines, which put high memory and compute pressure on the end user device for enterprise scale datasets. It also inherits the traditional backend-for-frontend security model, which all enterprise apps require, including consumer apps like Notion that make the bulk of their revenue from enterprise citizen devs and therefore are exposed to enterprise data security compliance. And this is in an AI-focused world where companies want to defend against AI scrapers so they can sell their data assets to foundation model providers for use in training! Which IMO is the real problem with sync engines: they are not a good match for enterprise applications, nor are they a good match for hyper scale consumer saas that aspire to sell into enterprise. So what market are they for exactly?
- ForTheKidz 2y ago> You’ll get your data synced for you How does this happen without an interface for conflict resolution? That's the hard part.
- Sammi 2y agoAll this recent hype about sync engines and local first applications completely disregards conflict resolution. It's the reason why syncing isn't mainstream already and it isn't solved and arguably cannot be. Imagine if git just on its own picked what to keep and what to throw away when there's a conflict. You fundamentally need the user to make the choice.
- porridgeraisin 2y agoPrecisely. The hype articles write all about the journey to The Wall, and then leave out the bit where you smash headfirst into it.
- lifty 2y agoVery good point. The local-sync ecosystem is still in a young phase, and conflict resolution hasn't been tackled or solved yet. Most systems have a |last write wins" approach.
- sgt 2y ago> All this recent hype about sync engines and local first applications completely disregards conflict resolution. Not really true though. I've used a couple of local sync engines, one internally built and another one which is both commercial and now open source called PowerSync[1]. Conflict resolution is definitely on the agenda, and a developer is definitely going to be mindful of conflicts when designing the application. [1] https://www.powersync.com/ https://www.powersync.com/
- Sammi 2y agoMy unfortunate point is that the dev cannot know what the user is doing, and so cannot in principle know what choice to make on behalf of the user in case of a conflict. This is not a code problem. It cannot be solved with code.
- DeathArrow 2y agoI've solved data sync in distributed apps long time ago. I send outgoing data to /dev/null and receive incoming data from /dev/zero. This way data is always consistent. That also helps with availability and partion tolerance.
- irisman 2y ago[flagged]
- zx8080 2y ago> decoupled from the horrors of an unreliable network There's no such thing as reliable network in the world. The world is network connected, there's almost no local-only systems anymore (for a long long time now). Some engineers dream that there's some cases when network is reliable, like when a system fully lives in the same region and single AZ. But even then it's actually not reliable and can have some glitches quite frequently (like once per month or so, depending on some luck).
- 01HNNWZ0MV43FF 2y agoTrue. Even the network between the CPU and an SD card or USB drive is not reliable
- tonsky 2y ago> There's no such thing as reliable network in the world I’m not saying there is
- jimbokun 2y agoI believe the point is that given an unreliable network, it's nice to have access to all the data available locally up to the point when you had a network issue. And then when the network is working again, your data comes up to date with no extra work on the application developer's part.
- Keyz56 2y ago[dead]
- _1tem 2y agoLocally synced databases seem to be a new trend. Another example is Turso, which works by maintaining a sort of SQLite-DB-per-tenant architecture. Couple that with WASM and we’ve basically come full circle back to old school desktop apps (albeit with sync-on-load). Fat client thin client blah blah.
- arkh 2y agoThe future of webapps: wasm in the browser, direct SQL for the API. Main problem? No result caching but that's "just" a middleware to implement.
- TeMPOraL 2y agoAlso the past of webapps. We don't have that because doing this properly, in a way that's maximally useful and ergonomic for the users, pretty much kills the entire business of the web. If you give direct SQL access to the underlying data, you can no longer seek rent by putting a bloated, barely-functional app in front of the database, nor can you use it to funnel users or upsell them stuff. Most of the money in this industry is made from rent-seeking.
- hyperbolablabla 2y agoHow does this compare to supabase?
- jFriedensreich 2y agothat question comes up all the time for some reason but supabase does not support offline or sync, only some form of subscription updating, but this has nothing to do with having sync or local data.
- zareith 2y agoI think an underappreciated library in this space is Logux [1] It requires deeper (and more) integration work compared to solutions that sync your state for you, but is a lot more flexible wrt. the backend technology choices. At its core, it is an action synchronizer. You manage both your local state and remote state through redux-style actions, and the library takes care of syncing and resequencing them (if needed) so that all clients converge at the same state. [1] https://logux.org/ https://logux.org/
- mike_hearn 2y agoI recently took a part time role at Oracle Labs and have been learning PL/SQL as part of a project. Seeing as Niki is shilling for his employer, perhaps it's OK for me to do the same here :) [1]. HN discourse could use a bit of a shakeup when it comes to databases anyway. This may be of only casual interest to most readers, but some HN readers work at places with Oracle licenses and others might be surprised to discover it can be cheaper than an AWS managed Postgres [2]. It has a couple of features relevant to this blog post. The first: Niki points out that in standard SQL producing JSON documents from relational tables is awkward and the syntax is terrible. This is true, so there's a better syntax: CREATE JSON RELATIONAL DUALITY VIEW dept_w_employees_dv AS SELECT JSON {'_id' : d.deptno, 'departmentName' : d.dname, 'location' : d.loc, 'employees' : [ SELECT JSON {'employeeNumber' :e.empno, 'name' : e.ename} FROM employee e WHERE e.deptno = d.deptno ] } FROM department d WITH UPDATE INSERT DELETE; It makes compound JSON documents from data stored relationally. This has three advantages: (1) JSON documents get materialized on demand by the database instead of requiring frontend code to do it, (2) the ORDS proxy server can serve these over HTTP via generic authenticated endpoints (e.g. using OAuth or cookie based auth) so you may not need to write any code beyond SQL to get data to the browser, and (3) the JSON documents produced can be written to, not only read. The second feature is query change notifications. You can issue a command on a connection that starts recording the queries issued on it and then get a callback or a message posted to an MQ when the results change (without polling). The message contains some info about what changed. So by wiring this up to a web socket, which is quite easy, the work of an hour or two in most web frameworks, then you can stream changes to the client directly from the database without needing much logic or third party integrations. You either use the notification to trigger a full requery and send the entire result json back to the browser, or you can get fancier and transform the deltas to json subsets. It'd be neat if there was a way to join these two features together out of the box, but AFAIK if you want full streaming of document deltas to the browser and reconstituting them there, it would need a bit more on top. Again, you may feel this is irrelevant because doesn't every self-respecting HN reader use Postgres for everything, but it's worth knowing what's out there. Especially as the moment you decide to paying a cloud for hosting your DB you have crossed the Rubicon anyway (all the hosted DBs are proprietary forks of Postgres), so you might as well price out alternatives. [1] and you know the drill, views are my own and nobody has reviewed this post. [2] https://news.ycombinator.com/item?id=42855546 https://news.ycombinator.com/item?id=42855546
- avodonosov 2y agoWhy he haven't implemented a full Datomic Peer for his DataScript I never understood. Having a datalog query engine, supplying it with data from Datomic indexes - b-tree like collections storing entity-attribute-value records - seems simple. Updating the local index cache from log is also simple. And that gets you a db in browser.
- tonsky 2y agoIt’s not as simple as you make it sound: - Reliable communication is hard - Optimistic writes should on client are hard - Tracking subsets of data is hard (you don't want the entirety of Datomic on the client, do you?) - Permissions are hard in this model Why didn't I implement it? Mostly comes down to free time. It's a hobby project and it's hard to find time for it. I also stopped writing web apps so immediate pressure for this went away.
- avodonosov 2y ago> - Optimistic writes should on client are hard This is out of scope - I don't mean a functional equivalent of instantdb. Just a database in browser. > - Reliable communication is hard The same, no special requirements. Just send a request, maybe retry several times (with increasing delays), and give up throwing an error. > Tracking subsets of data is hard (you don't want the entirety of Datomic on the client, do you?) That's the only thing really missing. And it doesn't seem hard. I think Datomic Peer just keeps a fixed number of index pages in cache. Pages missing in the cache are just retrieved from storage. In result the cache keeps the working subset - the elements related to the entities needed for queries and entity api requests made by the application. Especially since the indexes are ordered (EAVT, AEVT, AVET, VAET), much of the data in the cache will be relevant to the application. > - Permissions are hard in this model Permission is a question, but there are useful applications where permission control is not needed. Similar to what you say in another comment about conflict resolution: "Another misconception is that conflict resolution needs to be “solved” perfectly before any progress can be made. That is not true as well. You might have unhandled conflicts in your system and still have a working, useful, successful product." Back in the day when DataScript first appeared and I was eager to see it working with larger-than-memory datasets (and maybe even reading data saved by Datomic by understanding its format), I wanted that to enable public to run queries (read-only) on a large database I was assembling, that didn't fit into memory. In some applications all users may have equal write access to the document / data. Server-side usage of DataScript could be another case that does not require permissions support in the DB. That's how Datomic itself is used. I am not complaining, and understand there are limits on what people can do in free time. You did huge work on your open source projects. But I regretted that seemingly a small step to open DataScript to out-of-memory data, which I though would greatly expand its applicability, was missing. Good luck with instantdb. Hopefully commercial success will allow to continue putting work in it and improve the tech landscape.
- joeeverjk 2y agoIf sync really is the future, do you think devs will finally stop pretending local-first apps are some niche thing and start building around sync as the core instead of the afterthought? Or are we doomed to another decade of shitty conflict resolution hacks?
- Zanfa 2y ago> Or are we doomed to another decade of shitty conflict resolution hacks? Conflict resolution is never going away. It's important to distinguish between syntactical and semantical conflicts though, the first of which can be solved, but the other will always require manual intervention.
- deleted 2y ago[deleted]
- Tobani 2y agoI think this makes sense for applications applications that are just managing data maybe? But if your application needs to do things when you change that data (like call to a third party system)... Syncing is maybe not the solution. What happens when the total dataset is large, do you need to download 6gb of data every time you log in? Now you've blown up the quota on local storage. How do you make sure the appropriate data is downloaded or enough data? How do you prioritize the data you need NOW instead of waiting for that last byte of the 6gb to download? It is like a useful tool, but not the only future.
- keizo 2y agodidn't know that about roam research. I was a user, but also that app convinced me that front-end went in the wrong direction for a decade... Rocicorp Zero Sync, instantdb, linear app like trend is great -- sync will be big. I hope a lot of the spa slop gets fixed!
- mentalgear 2y agoHonourable mentions of some more excellent fully open-source sync engines: - Zero Sync: https://github.com/rocicorp/mono https://github.com/rocicorp/mono - Triplit: https://github.com/aspen-cloud/triplit https://github.com/aspen-cloud/triplit
- guappa 2y ago> - Zero Sync: https://github.com/rocicorp/mono https://github.com/rocicorp/mono Doesn't even have a readme :D Raise the bar a bit maybe.
- thruflo 2y agohttps://zero.rocicorp.dev/docs/introduction https://zero.rocicorp.dev/docs/introduction Hard to raise the bar on Zero. It’s a brilliant system.
- profstasiak 2y agocan you share how are you using this? Production / side projects? Would you recommend it for side projects?
- jakelazaroff 2y agoGP is probably not himself using Zero because he’s the CEO of Electric, which also makes a sync engine: https://electric-sql.com https://electric-sql.com
- thunderbong 2y ago"Website and Docs" is the second line I see
- daveguy 2y agoIt does have a readme. Click the "View all files" button. But you don't have to. GitHub shows the readme just below the partial file list. That's what all the same-page docs on GitHub/GitLab repositories are. Full docs are linked from the readme.
- mackopes 2y agoI'm not convinced that there is one generalised solution to sync engines. To make them truly performant at large scale, engineers need to have deep understanding of the underlying technology, their query performance, database, networking, and build a custom sync engine around their product and their data. Abstracting all of this complexity away in one general tool/library and pretending that it will always work is snake oil. There are no shortcuts to building truly high quality product at a large scale.
- tonsky 2y ago- You can have many sync engines - Sync engines might only solve small and medium scale, that would be a huge win even without large scale
- wim 2y agoWe've built a sync engine from scratch. Our app is a multiplayer "IDE" but for tasks/notes [1], so it's important to have a fast local first/office experience like other editors, and have changes sync in the background. I definitely believe sync engines are the future as they make it so much easier to enable things like no-spinners browsing your data, optimistic rendering, offline use, real-time collaboration and so on. I'm also not entirely convinced yet though that it's possible to get away with something that's not custom-built, or at least large parts of it. There were so many micro decisions and trade-offs going into the engine: what is the granularity of updates (characters, rows?) that we need and how does that affect the performance. Do we need a central server for things like permissions and real-time collaboration? If so do we want just deltas or also state snapshots for speedup. How much versioning do we need, what are implications of that? Is there end-to-end-encryption, how does that affect what the server can do. What kind of data structure is being synced, a simple list/map, or a graph with potential cycles? What kind of conflict resolution business logic do we need, where does that live? It would be cool to have something general purpose so you don’t need to build any of this, but I wonder how much time it will save in practice. Maybe the answer really is to have all kinds of different sync engines to pick from and then you can decide whether it's worth the trade-off not having everything custom-built. [1] https://thymer.com https://thymer.com
- Phelinofist 2y agoThe largest feature my team develops is a sync engine. We have a distributed speech assistant app (multiple embeddeds [think car and smartphone] & cloud) that utilizes the Blackboard pattern. The sync engine keeps the blackboards on all instances in sync. It is based on gRPC and uses a state machine on all instances that transitions through different states for connection setup, "bulk sync", "live sync" and connection wind down. Bulk sync is the state that is used when an instance comes online and needs to catch up on any missed changes. It is also the self-heal mechanism if something goes wrong. Unfortunately some embedded instances have super unreliable clocks that drift quite a bit (in both directions). We consider switching to a logical clock. We have quite a bit of code that deals with conflicts. I inherited this from my predecessor. Nowadays I would probably not implement something like this again, as it is quite complex.
- exceptione 2y agoI believe the idea of a Blackboard is that there is a single blackboard for all processes to asynchronously scribble and read from. Syncing blackboards sounds like going straight against the spirit of that design pattern.
- sreekanth850 2y agoWe use indexedDB and signalr for real time sync. What is new about this?
- rockmeamedee 2y agoIdk man. It's a nice idea, but it has to be 10x better than what we currently have to overcome the ecosystem advantages of the existing tech. In practice, people in the frontend world already use Apollo/Relay/Tanstack Query to do data caching and querying, and don't worry too much about the occasional overfetching/unoptimized-ness of the setup. If they need to do a complex join they write a custom API endpoint for it. It works fine. Everyone here is very wary of a "magic data access layer" that will fix all of our problems. Serverless turned out to be a nightmare because it only partially solves the problem. At the same time, I had a great time developing on Meteorjs a decade ago, which used Mongo on the backend and then synced the DB to the frontend for you. It was really fluid. So I look forward to things like this being tried. In the end though, Meteor is essentially dead today, and there's nothing to replace it. I'd be wary of depending so fully on something so important. Recently Faunadb (a "serverless database") went bankrupt and is closing down after only a few years. I see the product being sold is pitched as a "relational version of firebase", which I think good idea. It's a good idea for starter projects/demos all the way up to medium-sized apps, (and might even scale further than firebase by being relational), but it's not "The Future" of all app development. Also, I hate to be that guy but the SQL in example could be simpler, when aggregating into JSON it's nice to use a LATERAL join which essentially turns the join into a for loop and synthesises rows "on demand": SELECT g.*, COALESCE(t.todos, '[]'::json) as todos FROM goals g LEFT JOIN LATERAL ( SELECT json_agg(t.*) as todos FROM todos t WHERE t.goal_id = g.id ) t ON true That still proves the author's point that SQL is a very complicated tool, but I will say the query itself looks simpler (only 1 join vs 2 joins and a group by) if you know what you're doing.
- timita 2y ago> Meteor is essentially dead today Care to explain what you mean by "dead"? Just today v3.2 came out, and the company, the community, and their paid-for hosting service seem pretty alive to me.
- asdffdasy 2y ago> Such a library would be called a database. bold of them to assume a database can manage even the most trivial of conflicts. There's a reason you bombard all your writes to a "main/master/etc"
- profstasiak 2y agoso... what do people that want to have sync engines do? I want to try it for hobby project and I think I will go the route of just one way sync (from database to clients) using electric sql and I will have writes done in a traditional way (POST requests). I like the idea of having server db and local db in sync, but what happens with writes? I know people say CRDT etc... but they are solving conflicts in unintuitive ways... I know I probably sound uneducated, but I think the biggest part of this is still solving conflicts in a good way, and I don't really see how you can solve those in a way that works for all different domains and have it "collapsed" as the author says
- codeulike 2y agoI've been thinking about this a lot - nearly every problem these days is a synchronisation problem. You're regularly downloading something from an API? Thats a sync. You've got a distributed database? Sync problem. Cache Invalidation? Basically a sync problem. You want online and offline functionality? sync problem. Collaborative editing? sync problem. And 'synchronisation' as a practice gets very little attention or discussion. People just start with naive approaches like 'download whats marked as changed' and then get stuck in the quagmire of known problems and known edge cases (handling deletions, handling transport errors, handling changes that didn't get marked with a timestamp, how to repair after a bad sync, dealing with conflicting updates etc). The one piece of discussion or attempt at a systematic approach I've seen to 'synchronisation' recently is to do with Conflict-free Replicated Data Types https://crdt.tech https://crdt.tech which is essentially restricting your data and the rules for dealing with conflicts to situations that are known to be resolvable and then packaging it all up into an object.
- mrkeen 2y agoI've looked at CRDTs, and the concept really appeals to me in the general case, but in the specific cases, my design always ends up being "keep-all-the-facts" about a particular item. But then you defer the problem of 'which facts can I throw away?'. It's like inventing a domain-specific GC. I'd love to hear about any success cases people have had with CRDTs.
- yccs27 2y agoFor me the main issue with CRDTs is that they have a fixed merge algorithm baked in - if you want to change how conflicts get resolved, you have to change the whole data structure.
- WorldMaker 2y agoI feel like the state-of-the-art here is slowly starting to change. I think CRDTs for too many years got too caught up in "conflict-free" as a "manifest destiny" sort of thing more than "hope and prayer" and thought they'd keep finding the right fixed merged algorithm for every situation. I started watching CRDTs from the perspective of source control and having a strong inkling that "data is always messy" and "conflicts are human" (conflicts are kind of inevitable in any structure trying to encode data made by people). I've been thinking for a bit that it is probably about time the industry renamed that first C to something other than "conflict-free". There is no freedom from conflicts. There's conflict resistance, sure and CRDTs can provide in their various data structures a lot of conflict resistance. But at the end of the day if the data structure is meant to encode an application for humans, it needs every merge tool and review tool and audit tool it can offer to deal with those. I think we're finally starting to see some of the light in the tunnel in the major CRDT efforts and we're finally leaving the detour of "no it must be conflict-free, we named it that so it must be true". I don't think any one library is yet delivering it at a good high level, but I have that feeling that "one of the next libraries" is maybe going to start getting the ergonomics of conflict handling right.
- iansinnott 2y agoHave been using Instant for a few side projects recently and it has been a phenomenal experience. 10/10, would build with it again. I suspect this is also at least partially true of client-server sync engines in general.
- kenrick95 2y agoI concur with this. Been using it on my side project that only have a front-end. The "back-end" is 100% InstantDB. Although for me, I found that the permissions part a bit hard to understand, especially when it involves linking to other namespace. Haven't checked them for a while, maybe they've improved on this...
- skybrian 2y agoThis is also a tricky UI problem. Live updates, where web pages move around on you while you’re reading them, aren’t always desirable. When you’re collaborating with someone you know on the same document, you want to see edits immediately, but what about a web forum? Do you really need to see the newest responses, or is this a distraction? You might want a simple indicator that a reload will show a change, though. A white paper showing how Instant solves synchronization problems might be nice.
- qudat 2y agoThe problem with sync engines is needing full-stack buy-in in order for it to work properly. Having a separate backend-for-frontend service defeats the purpose in my mind. So what do you do when a company already has an API and other clients beyond a web app? The web app has to accommodate. I see this as the major downside with sync engines. I've been using `starfx` which is able to "sync" with APIs using structured concurrency: https://github.com/neurosnap/starfx https://github.com/neurosnap/starfx
- wslh 2y agoSync, in general, is a very complex topic. There are past examples, such as just trying to sync contacts across different platforms where no definitive solution emerged. One fundamental challenge is that you can’t assume all endpoints behave fairly or consistently, so error propagation becomes a core issue to address. Returning to the contacts example, Google Contacts attempts to mitigate error propagation by introducing a review stage, where users can decide how to handle duplicates (e.g., merge contacts that contain different information). In the broader context of sync, this highlights the need for policies to handle situations where syncing is simply not possible beyond all the smart logic we may implement.
- zelon88 2y agoHere's an idea.... Stop putting your critical business data on disparate third party systems that you don't have access to. Problem solved!
- voidpointer 2y agoProbably a silly question, but if you take this all the way and treat everything as a DB that is synchronized in the background, how do you manage access control where not every user/client is supposed to have access to every object represented in the DB? Where does that logic go? If you do it on the document level like figma or canvas, every document is a DB and you sync the changes that happen to the document but first you need access to the document/DB. But doesn't this whole idea break apart if you need to do access control on individual parts of what you treat as the DB because you would need to have that logic on the client which could never be secure...
- PaulHoule 2y agoLotus Notes was a product far ahead of its time (nearly forgotten today) which was an object database with synchronization semantics. They made a lot of decisions that seem really strange today, like building an email system around it, but that empowered it for long-running business workflows. It's something everybody in the low-code/no-code space really needs to think about.
- ddrdrck_ 2y agoNo one that has ever had to work with Lotus Notes could forget it. It was atrocious. Maybe the sync engine was great but I really do not know what it was used for ...
- spankalee 2y agoThe problem I have with "moving the database to the client" is the same one I have in practice with CRDTs: In my apps, I need to preserve the history of changes to documents, and I need to validate and authenticate based on high-level change descriptions, not low-level DB access. This always leads me back to operational transforms. Operations being reified changes function as undo records; a log of changes; and a narrower, semantically-meaningful API, amenable to validation and authz. For the Roam Firebase example: this only works if you can either trust the client to always perform valid actions, or you can fully validate with Firebase's security rules. OT has critiques, but almost all of the fall away in my experience when you have a star topology with a central service that mediates everything - defining the canonical order of operations, performs validation & auth, and records the operation log.
- jimbokun 2y ago> This always leads me back to operational transforms. Operations being reified changes function as undo records; a log of changes; and a narrower, semantically-meaningful API, amenable to validation and authz. Sounds like another kind of synchronization database.
- spankalee 2y agoI think it's only a database if you come down on the "logs are the source of truth, not tables" side of the logs vs tables debate. And if you do, any log is a database, I guess...
- VikingCoder 2y agoThere are two hard problems: 1. Naming things 2. Caching 3. Off-by-one errors
- curtisblaine 2y agoRelated: - https://news.ycombinator.com/item?id=43436645 https://news.ycombinator.com/item?id=43436645 - https://greenvitriol.com/posts/sync-engine-for-everyone https://greenvitriol.com/posts/sync-engine-for-everyone
- loquisgon 2y agoThe local first people (https://localfirstweb.dev/ https://localfirstweb.dev/) have some cool ideas about how to solve the data synch problem. Check it out.
- shikhar 2y agoWe have had interest in using our serverless stream API (https://s2.dev/ https://s2.dev/) to power sync engines. Very excited about these kinds of use cases, email in profile if anyone wants to chat.
- finolex 2y agoIf anyone could be kind to give feedback on the local-first x data ownership db we're building, would really appreciate it! https://docs.basic.tech/ https://docs.basic.tech/ Will do my best to take action on any feedback I receive here
- Nelkins 2y agoDiscussion of sync engines typically goes hand in hand with local-first software. But it seems to be limited to use cases when the amount of data is on the smaller side. For example, can anyone imagine how there might be a local-first version of a recommendation algorithm (I'm thinking something TikTok-esque)? This would be a case where the determination of the recommendation relies on a large amount of data. Or think about any kind of large-ish scale enterprise SaaS. One of the clients I'm working with currently sells a Transportation Management Software system (think logistics, truck loads, etc). There are very small portions of the app that I can imagine relying on a sync engine, but being able to search over hundreds of thousands of truck loads, their contents, drivers, etc seems like it would be infeasible to do via a sync engine. I mention this because it seems that sync engines get a lot of hype and interest these days, but they apply to a relatively small subset of applications. Which may still be a lot, but it's a bit much to say they're the future (I'm inferring "of application development"--which is what I'm getting from this article).
- ochiba 2y agoI think that is where sync engines come in that allow doing arbitrary hybrid queries (across local and remote data) and then keeping the results of those hybrid queries in sync on the client. This is one of the ideas that appears to be central to the genesis of Zero [1] ElectricSQL allows for a similar pattern and PowerSync is also working on this [2] [1] https://www.youtube.com/watch?v=rqOUgqsWvbw https://www.youtube.com/watch?v=rqOUgqsWvbw [2] https://www.powersync.com/blog/powersync-2025-roadmap-sqlite-web-speed-and-versatility#1-on-demand-syncing-of-data-in-addition-to-pre-syncing-data https://www.powersync.com/blog/powersync-2025-roadmap-sqlite...
- Nelkins 2y agoInteresting! I'll give these a look. Edit: I watched the presentation (which I really enjoyed) and also read the blog post. For anyone with less time, the answer is essentially: don't sync everything, treat the local data like a cache. Sync as much as you can into that cache, and then reach out to the server for other things.
- deleted 2y ago
- beders 2y agoI found it quite disappointing to find a marketing piece from Nikki. It is full of general statements that are only true for a subset of solutions. Enterprise solutions in particular are vastly more complex and can't be magically made simple by a syncing database. (no solution comes even close to "99% business code". Not unless you re-define what business code is) It is astounding how many senior software engineers or architects don't understand that their stack contains multiple data models and even in a greenfield project you'll end up with 3 or more. Reducing this to one is possible for simple cases - it won't scale up. (Rama's attempt is interesting and I hope it proves me wrong) From: "yeah, now you don't need to think about the network too much" to "humbug, who even needs SQL" I've seen much bigger projects fail because they fell for one or both of these ideas. While I appreciate some magic on the front-end/back-end gap, being explicit (calling endpoints, receiving server-side-events) is much easier to reason about. If we have calls failing, we know exactly where and why. Sprinkle enough magic over this gap and you'll end up in debugging hell. Make this a laser focused library and I might still be interested because it might remove actual boilerplate. Turn it into a full-stack and your addressable market will be tiny.
- hamilyon2 2y agoI am feeling a bit confused. Is not the stated problem solved 99.9% with decades-old battle-proven optimistic locking and some careful retries?
- delusional 2y ago> I’ve yet to see a code base that has maintained a separate in-memory index for data they are querying Define "separate" but my old X11 compositor project neocomp I did something like that with a series of AOS arrays and bitfields that combined to make a sort of entity manager. Each index in the arrays was an entity, and each array held a data associated with a "type" of entity. An entity could hold multiple types that would combine to specify behavior. The bitfield existed to make it quick to query. It waaay too complicated for what it was, but it was fun to code and worked well enough. I called it a "swiss" (because it was full of holes). It's still online on github (https://github.com/DelusionalLogic/NeoComp/blob/master/src/swiss.h https://github.com/DelusionalLogic/NeoComp/blob/master/src/s...) even though I don't use it much anymore.
- Pamar 2y agoMaybe I am just dumb but I really cannot see how data synch could solve what (in my kind of business) is a real problem. Example: you develop a web app to book for flights online. My browser points to it and I login. Should synchronization start right now? Before I even input my departure point and date? Ok, no. I write NYC -> BER, and a dep date. Should I start synching now? Let's say I do. Is this really more efficient than querying a webservice? Ok, now all data are synched. Even potentially the ones for business class, even if I just need economy. You kniw, I could always change my mind later. Or find out that on the day I need to travel no economy seats are available anymore. Whatever. I have all the inventory data that I need. Raw. Guess what? As a LH frequent flyer I get special treatment in terms of price. Not just for LH, but most Business Alliance airlines. This logic is usually on the server, because airlines want maximum creativity and flexibility in handling inventory. Should we just synch data and make the offer selection algorithm run on the webserver instead? Let's say it does not matter... I have somehow in front of me all the options for my trip. So I call my wife to confirm she agrees with my choice. I explain her the alternatives... this takes 5 minutes. In this period, 367 other people are buying/cancelling trips to Europe. So I either see my selection constantly change (yay! Synchronization!!!) or I press confirm, and if my choice is gine I get a warning message and I repeat my query. Now add two elements: - airlines prefer not to show real numbers of available seats - they will usually send you a single digit from 1 to 9 or a "*" to mean "10 or more". So just symching raw data and let the combinatorial engine work in the browser is not a very good idea. Also, I see the pontential to easily mount DDOS attacks if every client is constantly being synchronized by copying high contention tables in RT. What am I missing here?
- earthnail 2y agoYour use case doesn’t benefit from your own data. There’s nothing you can do that doesn’t require a direct interaction from the server. I write an audio recording app, and in my app, users have most to gain from their own data. For most people, syncing is basically an afterthought. In this use case, the ability of having your recordings in your phone is the most important thing. The difference here lies that in my app, the user generates all the valuable data themselves. In your app, nothing valuable can happen without communication with the airline.
- quantadev 2y agoIPFS is a technology very helpful for syncing. One way it's being used in a modern context (although only sub-parts of IPFS stack) is how BlueSky engineers, during their design process a few years ago, accepted my proposal that for a new Social Media protocol, each user should have his own "Repository" (Basically a Merkel Tree) of everything he's ever posted. Then there's just a "Sync" up to some master service provider node (decentralized set of nodes/servers) for the rest of the world to consume. Merkel-Tree based synching is as performant as you can possibly get (used by Git protocol too I believe) because you can tell of a root of a tree-structure is identical to some other remote tree structure just by comparing the Hash Strings. And this can be recursively applied down any "changed branches" of a tree to implement very fast syncing mechanisms. I think we need a NEW INTERNET (i.e. Web3, and dare I say Semantic Web built in) where everyone's basically got their own personal "Tree of Stuff" they can publish to the world, all naively built into some new kind of tree structure-based killer app. Like imagine having Jupyter Notebooks in Tree form, where everything on it (that you want to be) is published to the web.
- ativzzz 2y agoI've always wondered, how do applications with more stringent security requirements handle this? Assume that permissions to any row in the DB can be removed at any time. If we store the data offline, this security measure is already violated. If you don't care about a user potentially storing data they no longer have access to, when they come online, any operations they make are invalid and that's fine But, if security access is part of your business logic, and is complex enough to the point where it lives in your app and not in your DB (other than using DB tools like RLS), how do you verify that the user still has access to all cached data? Wouldn't you need to re-query every row every time? I'm still uncertain how these sync engines can be secured properly
- fxnn 2y agoThe author would be excited to learn that CouchDB solves this problem since 20 years. The use case the article describes is exactly the idea behind CouchDB: a database that is at the same time the server, and that's made to be synced with the client. You can even put your frontend code into it and it will happily serve it (aka CouchApp). https://couchdb.apache.org https://couchdb.apache.org
- ltbarcly3 2y agoThis has been solved every 5 years or so, and along the way people learn why this solution doesn't actually work.
- erichocean 2y agoI designed the sync engine for Things Cloud [0] over a decade ago. It seems to have worked out pretty well for them. (The linked page has some details about what it can do.) When sync Just Works™, it's a magical thing. One of the reason's my design has been reliable from its very first release, even across multiple refactors/rewrites (I believe it's currently on its third, this time to Swift) is that it uses a Git-like model internally with pervasive hashing. It's almost impossible for sync to work incorrectly (if it works at all). [0] https://culturedcode.com/things/cloud/ https://culturedcode.com/things/cloud/
- philsnow 2y agoI started using Things about 4-5 years ago and it's been amazing, partly because of the first-class support for syncing between devices and their cloud. Thanks for making this great! I would be interested to read any articles you've written about Things's sync pattern, if any.
- erichocean 2y agoHashing + custom server merge is the main logical device, the rest is just encoding tricks to make everything fast. Push becomes like git push to a tmp branch, then you do a server-side merge, then a pull on remaining clients. Push/pull is fast just like git, and due to content hashing, there are never any problems overwriting another device's work on the server. I think I had some clever tricks with text syncing (IIRC I implemented a Merkle tree-inspired approach), but that's the general concept. (I don't remember anything else, it was so long ago. :-)
- stopachka 2y ago(Instant team member here) Wanted to drop in and say I _love_ Things and have used it now for about 7 years! Thank you for your work
- jiggawatts 2y ago> Such a library would be called a database. But we’re used to thinking of a database as something server-related, a big box that runs in a data center. It doesn’t have to be like that! Databases have two parts: a place where data is stored and a place where data is delivered. That second part is usually missing. Yes! A thousand times this! Databases can't just "live on a server somewhere", their code should extend into the clients. The client isn't just a network protocol parser / serialiser, it should implement what is essentially an untrusted, read-only replica. For writes, it should implement what is essentially a local write-ahead log (WAL) either in-memory and optionally fsync-d to local storage. All of this should use the same codebase as the database engine, or machine-generated in multiple languages from some sort of formal specification.
- theanirudh 2y agoHow do sync engines address issues where we need something to be more dynamic? Currently I'm building a language learning app and we need to display your "learning path" - what lessons you have finished and what are your next lessons. The next lessons aren't fixed/same for everyone. It will change depending on how the score of completed lessons. Is any query language dynamic enough to support use cases like this? Or is it expected to recalculate the next lessons whenever the user completes a lesson and write it out to a table which can then be queried easily?
- theanirudh 2y agoSeems like a lot of extra work in cases where we change the scoring mechanism, we will then have to invalidate the existing entries, recalculate and write it out again compared to just having an endpoint that will take all previous lessons and generate the next lessons on demand.