3 ms·
Author here, sorry if this was not clear: that specific point was not supposed to be an indictment of all CRDTs, it was supposed to be much more narrow. Specifi
by antics 7mo ago
Author here, sorry if this was not clear: that specific point was not supposed to be an indictment of all CRDTs, it was supposed to be much more narrow. Specifically, the Yjs authors clearly state that they purposefully designed its interface to ProseMirror to delete and recreate the entire document on every collab keystroke, and the fact that it stayed open for 6 YEARS before they started to try to fix it, does in my opinion indicate a fundamental misunderstanding of what modern text editors need to behave well in any situation. Not even a collaborative one. Just any situation at all.
I think it's defensible to say that this point in particular is not indicting CRDTs in general because I do say the authors are trying to fix it, and then I link to the (unpublicized) first PR in that chain of work (which very few people know about!), and I specifically spend a whole paragraph saying I hope that I a forced to write an article in a year about how they figured it all out! If I was trying to be disingenuous, why do any of that?
- cowboy_henk 7mo ago> sorry if this was not clear It's easy to make that mistake reading your post because of sentences like > I want to convince you that all of these things (except true master-less p2p architecture) are easily doable without CRDTs > But what if you’re using CRDTs? Well, all these problems are 100x harder, and none of these mitigations are available to you. It sure sounds a lot like you're calling CRDTs in general needlessly complex, not just the yjs-prosemirror integration.
- antics 7mo agoTo be clear, we ARE arguing CRDTs needlessly complex for the centralized server use case. What I am describing in the "delete and replace all on every keystroke" problem is the point at which it became clear to me that the project did not understand what modern text editors need to perform well in any circumstance, let alone a collab one. I think this is still reasonable to say because the final paragraph in that section is 100% about how they might fix the delete-all problem, and I hope they do, so that I can write about that, too. But also, that the rest of the article is going to be about how you have to swim upstream against their architecture to accomplish things that are either table stakes or trivial in other solutions.
- josephg 7mo ago> To be clear, we ARE arguing CRDTs needlessly complex for the centralized server use case. I've been working in the OT / CRDT space for ~15 years or so at this point. I go back and forth on this. I don't think its as clear cut as you're making it out to be. - I agree that OT based systems are simpler to program and usually simpler to reason about. - Naive OT algorithms perform better "out of the box". CRDTs often need more optimisation work to achieve the same performance. - But with some optimisation work, CRDTs perform better than OT based systems. - CRDTs can be used in a client/server model or p2p. OT based systems generally only work well in a centralised context. Because of this, CRDTs let you scale your backend. OT (usually) requires server affinity. CRDT based systems are way more flexible. Personally I'd rather complex code and simpler networking than the other way around. - Operation based CRDTs can do a lot more with timelines - eg, replaying time, rebasing, merging, conflicts, branches, etc. OT is much more limited. As a result, CRDT based systems can be used for both realtime editing or for offline asyncronous editing. OT only really works for online (realtime) editing. (For anyone who's read the papers, I'm conflating OT == the old Jupitor based OT algorithm that's popular in google docs and others.) CRDTs are more complex but more capable. They can be used everywhere, and they can do everything OT based systems can do - at a cost of more code. You can also combine them. Use a CRDT between servers and use OT client-to-server. I made a prototype of this. It works great. But given you can make even the most complex text based CRDT in a few hundred lines anyway[1], I don't think there's any point. [1] https://github.com/josephg/egwalker-from-scratch https://github.com/josephg/egwalker-from-scratch
- GermanJablo 7mo ago> But with some optimisation work, CRDTs perform better than OT based systems. I read your paper and I think this is a mistake. You assume that OT has quadratic complexity because you're considering classic operation-based OT. But OT can be id-based, in which case operations are transformed directly on the document, not on other operations. This is essentially CRDT without the problems of supporting P2P, and therefore the best CRDT will never perform better than the best OT. > CRDTs let you scale your backend. OT (usually) requires server affinity. CRDT based systems are way more flexible. Personally I'd rather complex code and simpler networking than the other way around. All productivity apps that use these tools in any way shard by workspace or user, so OT can scale very well. If you don't scale CRDT that way, by the way, you'd be relying too much on "eventual consistency" instead of "consistency as quickly as possible." > (For anyone who's read the papers, I'm conflating OT == the old Jupiter-based OT algorithm that's popular in Google Docs and others.) Similar to what I said before. I think limiting OT to an implementation that’s over three decades old doesn’t do OT justice.
- josephg 7mo ago> the fact that it stayed open for 6 YEARS before they started to try to fix it... This is all opensource software, provided for free by volunteers. If you want better bindings, go write them. Or pay someone else to do so.
- antics 7mo agoJust to be extremely clear: we pay for a lot of OSS software. We pay one individual project more than $10,000 a year. We would have paid for Yjs too, if we thought it was a good use of resources!