11 ms·
Zero-latency SQLite storage in every Durable Object
- kikimora 2y agoThis design does not handle hot partitions well and they are ubiquitous to so many domains.
- simonw 2y agoYour partition would have to be VERY hot for SQLite not to be able to handle it - anything up to several thousand writes per second would likely work fine. Since this is all running on Cloudflare you could scale reads with a 1 second cache TTL somewhere, which would drop your incoming read queries to around one per second no matter how much read traffic you had.
- kondro 2y agoDoes this mean SQLite for DO can lose up to 10 seconds of data in the event of a failing DO?
- stavros 2y ago> To ensure durability beyond that ten second window, writes are also forwarded to five replicas in separate nearby data centers as soon as they commit, and the write is only acknowledged once three of them have confirmed it. I think Simon meant "within", rather than "beyond", here.
- simonw 2y agoThanks, I've updated that word.
- stavros 2y agoThis is a really interesting design, but these kinds of smart systems always inhabit an uncanny valley for me. You need them in exactly two cases: 1. You have a really high-load system that you need to figure out some clever ways to scale. 2. You're working on a toy project for fun. If #2, fine, use whatever you want, it's great. If this is production, or for Work(TM), you need something proven. If you don't know you need this, you don't need it, go with a boring Postgres database and a VM or something. If you do know you need this, then you're kind of in a bind: It's not really very mature yet, as it's pretty new, and you're probably going to hit a bunch of weird edge cases, which you probably don't really want to have to debug or live with. So, who are these systems for, in the end? They're so niche that they can't easily mature and be used by lots of serious players, and they're too complex with too many tradeoffs to be used by 99.9% of companies. The only people I know for sure are the target market for this sort of thing is the developers who see something shiny, build a company (or, worse, build someone else's company) on it, and then regret it pretty soon and move to something else (hopefully much more boring). Does anyone have more insight on this? I'd love to know.
- jmtulloss 2y agoIf you're in #1, you talk to CloudFlare. They need some great customer stories and they have some great engineers that are most likely willing to work with you on how this will work/help you with bugs in exchange for some success stories. If it gets proven out this turns into a service relationship, but early on it's a partnership.
- gregwebs 2y agoThere are a lot of cases of low traffic applications that aren’t toys but instead are internal tools- this could be a great option for those. For higher traffic they are asking you to figure out how to shard your data and it’s compute. That’s really hard to do without hitting edge cases.
- stavros 2y agoWhy would you use this for an internal, low-traffic tool over Postgres?
- fracus 2y agoCould this be used to get a time edge in trading? I'm not an expert, just thinking out loud. I remember hearing about firms laying wire in a certain way because getting a microsecond jump on changing rates could be everything for them.
- crabmusket 2y agoI'm also no expert, but from reading around the subject a little (Flash Boys by Michael Lewis was pretty cool, also Jane Street's podcast has some fantastic information)... no. I doubt you'd be on a public cloud if low-latency trading is what you're doing.
- aldonius 2y agoAren't the HFT boxes usually stock exchange colocations? Each trader gets a rack (or multiple racks depending on size) in the exchange's datacenter, every rack has the same cable length to the switch, etc.
- simonw 2y agoOne thing I don't understand about Durable Objects yet is where they are physically located. Are they located in the region that hosted the API call that caused them to be created in the first place? If so, is there a mechanism by which a DO can be automatically migrated to another location if it turns out that e.g. they were created in North America but actually all of the subsequent read/write traffic to them comes from Australia?
- ko_pivot 2y agoDurable Objects have long term storage. They get hydrated from that storage, so in that sense, they can move to any Cloudflare DS. However, there is no API call to move a Durable Object. It has to have no connections and then gets recreated in the DS nearest to the next/first connection. Memory gets dropped when that happens, storage survives. (This is slightly out of date as they have some nuanced hibernation stuff that is recent).
- masterj 2y ago> Durable Objects do not currently change locations after they are created > Dynamic relocation of existing Durable Objects is planned for the future. https://developers.cloudflare.com/durable-objects/reference/data-location/#:~:text=Durable%20Objects%20do%20not%20currently,will%20be%20in%20close%20proximity https://developers.cloudflare.com/durable-objects/reference/.... IIRC Orleans (https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/Orleans-MSR-TR-2014-41.pdf https://www.microsoft.com/en-us/research/wp-content/uploads/...) allows actors to be moved between machines, which should map well to DOs being moved between locations.
- pests 2y agoAs actors in Orleans are virtual and persistent it can also be the case it is running nowhere. If it's stateless it could be running in multiple locations. I worry "Dynamic relocation of DOs" might be going a bit too granular, this should be something the runtime takes care of.
- mhart 2y agoBy default in the region you created them in, but you can alternatively specify a locationHint. Use "oc" for Australia. https://developers.cloudflare.com/durable-objects/reference/data-location/#supported-locations-1 https://developers.cloudflare.com/durable-objects/reference/... Note the "Dynamic relocation of existing Durable Objects is planned for the future"
- emadda 2y agoSome other interesting points: - The write api is sync, but it has a hidden async await: when you do your next output with a response, if the write fails the runtime will replace the response with a http failure. This allows the runtime to auto-batch writes and optimistically assume they will succeed, without the user explicitly handling the errors or awaits. - There are no read transactions, which would be useful to get a pointer to a snapshot at a point in time. - Each runtime instance is limited to 128mb RAM. - Websockets can hibernate and you do not have to pay for the time they are sleeping. This allows your clients to remain connected even when the DO is sleeping. - They have a kind of auto RPC ability where you can talk to other DOs or workers as if they are normal JS calls, but they can actually be calling another data center. The runtime handles the serialisation and parsing.
- crabmusket 2y agoThe RPC stuff is pretty interesting. More here: https://blog.cloudflare.com/javascript-native-rpc/ https://blog.cloudflare.com/javascript-native-rpc/
- skybrian 2y agoWithout a schema, I’m wondering about validation. I guess your server should use Zod or an equivalent library?
- deleted 2y ago[deleted]
- matharmin 2y agoJust wondering, do you have a specific use case for read transactions implemented on the database level here? In SQLite in general read transactions are useful since you can access the same database from multiple processes at a time. Here, only a single process can access the database. So you can get the same effect as read transactions either by doing all reads in one synchronous function, or implement your own process-level locking.
- 2y ago
- braden-lk 2y agoDurable objects seem so cool but the pricing always scares me. (Specifically, having to worry about getting hibernation right.) They’d be a great fit for our yjs document based strategy, but while everything in prod still works on plain ol redis and Postgres, it’s hard to justify an exploration.
- ignoramous 2y ago> Specifically, having to worry about getting hibernation right. As long as the client doesn't exchange websocket messages with DO, it'll hibernate. From what I can tell, ping/pong frames don't count towards uptime, if you're worried about that.
- attilakun 2y agoDoes CloudFlare have proper spending caps? If they have, I'd be open to try DOs but if they don't, it's a non-starter for an indie dev as I can't risk bankruptcy due to a bad for loop.
- viraptor 2y agoIt's not just the listed prices either. There was a story here not long ago where they essentially requested someone to migrate to an enterprise plan or get out. With AWS it's pretty common to get a refund for accidental abuse. From my contact so far and from stories here, I wouldn't expect anything close to that treatment from CF.
- anentropic 2y agoWhat scares me is it is super specific to Cloudflare What is your option if you want to eject to another cloud?
- paulgb 2y agoFor the specific example of a Yjs backend, I happen to be working on one that can be hosted either on Cloudflare or as a native process. We’ve had people running in production migrate from cloudflare to native just by swapping out the URL they connect to in their application config. https://github.com/jamsocket/y-sweet https://github.com/jamsocket/y-sweet
- pajeets 2y agowonder how this works with Pocketbase
- deleted 2y ago[deleted]
- tmikaeld 2y ago> ..each DO constantly streams a sequence of WAL entries to object storage - batched every 16MB or every ten seconds. Which also means it may take 10 seconds before you can (reliably) read the write globally. I keep failing to see how this can replace regionally placed database clusters which can serve a continent in milliseconds. Edit: I know it uses streams, but those are only to 5 followers and CF have hundreds of datacenters. There is no physical way to guarantee reads in seconds unless all instances of the SQLite are always connected and even then, packet latency will cause issues.
- kentonv 2y agoAs others have noted, you misunderstand how Durable Objects work. All traffic addressed to the same object is routed to a single machine where that object lives. That machine always has a consistent view of its SQLite database. You can have billions of objects, but each has its own separate database. There's no way to read from a database directly from a different machine than the one the DO is running on.
- neamar 2y agoThe writes are streamed in near real time to five followers, acknowledging it near instantly. The cloudflare blog article mention this more in depth. So writes remain fast, while still having durability.
- memothon 2y agoThose WAL entries streamed to object storage I think are just for backups. Each DO is globally unique (there's one DO with a given id running anywhere) and runs sqlite on its own local storage in that datacenter.
- firtoz 2y agoAFAIK the writes and reads are done only from the same process, so the long term storage will apply only if the current process is hibernated. When you write something and then read it, it's immediate, because the writes and reads are also updating the current process's state in memory. For another process (e.g. another DO or another worker) to access the data, they need to go through the DO which "contains" the data, so they'd be making a RPC or a HTTP request to the DO, and they'd get the latest information. + the hibernation happens after x seconds of inactivity, so it feels like the only time a data write to be unavailable as expected would be when the DO or worker crashes right after a write.
- segalord 2y agoNoticing CF pushing for devs to use DO for eveything over workers these days. Even websocket connections on workers get timed out after ~30s and the recommended way is to use DO for them
- rozenmd 2y agoDurable Objects have always been the recommended way to do websocket connections on Cloudflare Workers? (as far as I remember, anyway) The original chat demo dates back to 2020, using DOs + websockets: https://github.com/cloudflare/workers-chat-demo https://github.com/cloudflare/workers-chat-demo
- _1tem 2y agoWhat I don’t understand is why, in the example of flight seat mapping provided, you create a DO per flight. So does a DO correspond to a “model” in MVC architecture? What if I used DOs in a per-tenant way, so one DO per user. And then how do I query or “join” across all DOs to find all full flights? I guess you would have to design your DOs such that joins are not required?
- ec109685 2y agoThey support “function” calling between DOs, so you are able to compose a response from more than one DO.
- skrebbel 2y agoI really love the Durable Object design, particularly because it's easy to understand how it works on the inside. Unlike lots of other solutions designed for realtime data stuff, Durable Objects have a simplicity to them, much like Redis and Italian food. You can see all the ingredients. Given enough time and resources (and datacenters :) ), a competent programmer could read the DO docs and reimplement something similar. This makes it easy to judge the tradeoffs involved. I do worry that DOs are great for building fast, low-overhead, realtime experiences (eg five people editing a document in realtime), but make it very hard to make analyses and overviews (which groups of people have been which editing documents the last week?). Putting the data inside SQLite might make that even harder - you'd have to somehow query lots and lots of little SQLite instances and then merge the results together. I wonder if there's anything for this with DOs, because this is what keeps bringing me back to Postgres time and time again: it works for core app features and for overviews, BI, etc.
- parsadotsh 2y agoI think for those cases you're expected to use something like this: https://developers.cloudflare.com/analytics/analytics-engine/ https://developers.cloudflare.com/analytics/analytics-engine...
- jwblackwell 2y agoDoes anyone else struggle to wrap their head around a lot of this new cloud stuff? I have 15+ years experience of building for the web, using Laravel / Postgres / Redis stack and I read posts like this and just think, "not for me".
- djtango 2y agoFrom the article: > For useful background on the first version of Durable Objects take a look at Cloudflare's durable multiplayer moat by Paul Butler, who digs into its popularity for building WebSocket-based realtime collaborative applications. First apps that come to mind that have RT collaboration: - Google Docs/Sheets etc - Notion - Miro - Figma These are all global scale collaborative apps, I'm not sure a Laravel stack will support those use cases... Google had to in house everything and probably spearheaded the usage of CRDTs ( this is a guess!) but as the patterns emerge and the building blocks get SAASified, mass-RT collaboration no longer becomes a giant engineering problem and more and more interesting products get unlocked
- fastball 2y agoGoogle actually uses OT for their collab.
- jlokier 2y ago> Google had to in house everything and probably spearheaded the usage of CRDTs ( this is a guess!) Fwiw, Google Docs/Sheets etc don't use CRDTs, they use the more server-oriented Operational Transforms (OT). CRDTs were spearheaded by others.
- rcarmo 2y agoThe first thing I wondered was how this plays with data residency and privacy/regulatory requirements.
- blixt 2y agoI'm constantly impressed by the design of DOs. I think it's easy to have a knee-jerk reaction that something is wrong with doing it this way, but in reality I think this is exactly how a lot of real products are implicitly structured: a lot of complex work done at very low scale per atomic thing (by which I mean, anything that needs to be transactionally consistent). In retrospect what we ended up building at Framer for projects with multiplayer support where edits are replicated at 60 FPS while being correctly ordered for all clients is a more applied version of what DOs are doing now. We also ended up with something like a WAL of JSON object edits so in case a project instance crashed its backup could pick up as if nothing had happened, even if committing the JSON patches into the (huge) project data object didn't have time to occur (on an every-N-updates/M-seconds basis just like described here).
- 9dev 2y agoI would love to work with Durable Objects and all the other cool stuff from Cloudflare, but I’m really hesitant to make a single cloud providers technology the backbone of my application. If CF decides to pull the plug, or charge a lot more, the only way to migrate elsewhere would be rebuilding the entire app. As long as there aren’t any comparable technologies, or abstraction layers on top of DOs, I’m not going to make the leap of faith.
- deleted 2y ago[deleted]
- vlaaad 2y agoRe https://where.durableobjects.live/ https://where.durableobjects.live/ — why the hell are they still operating in Russia?
- esnard 2y agoFrom https://blog.cloudflare.com/steps-taken-around-cloudflares-services-in-ukraine-belarus-and-russia/ https://blog.cloudflare.com/steps-taken-around-cloudflares-s... : > Since the invasion, providing any services in Russia is understandably fraught. Governments have been united in imposing a stream of new sanctions and there have even been some calls to disconnect Russia from the global Internet. As discussed by ICANN, the Internet Society, the Electronic Frontier Foundation, and Techdirt, among others, the consequences of such a shutdown would be profound. > [...] > Beyond this, we have received several calls to terminate all of Cloudflare's services inside Russia. We have carefully considered these requests and discussed them with government and civil society experts. Our conclusion, in consultation with those experts, is that Russia needs more Internet access, not less.
- CyberDildonics 2y agoWhat is the difference between a "durable object" and a file?
- simonw 2y agoYou can read the full article for an answer to that: https://blog.cloudflare.com/sqlite-in-durable-objects/ https://blog.cloudflare.com/sqlite-in-durable-objects/ Short version: it's replicated to five data centers on every transaction, and backed up as a stream to object storage as well.
- CyberDildonics 2y agoSo it's a file that gets backed up like dropbox
- simonw 2y agoAt a very high level, yes. But the details matter here - you commit a transaction to the SQLite database and know that the commit has been pushed out to 3/5 replicas by the time the write API request returns - and that it will be logged to object storage (supporting 30 days of rollback) within ten seconds. AND it will live in a data center physically close to the user who caused it to be created.
- avinassh 2y agoI'd love to know how they have hooked VFS with WAL to monitor changes. The SQLite's WAL layer deals with page numbers where as VFS deals with file and byte offsets. I am curious to understand how they mapped it, how they get new writes to the WAL and read from the WAL.
- bluehatbrit 2y agoThis is probably a really stupid question, but how would one handle schema migrations with this kind of setup? My understanding is it's aimed at having a database per-tenant (or even more broken down than that). Is there a sane way of handling schema migrations, or is the expectation that these databases are more short-lived and so you support multiple versions of the db (DO) until it's deleted? In my head, this would be a fun way to build a bookmark service with a DO per user. But as soon as you want to add a new field to an existing table, you meet a pretty tricky problem of getting that change to each individual DO. Perhaps that example is too long lived though, and this is designed for more ephemeral usage. If anyone has any experience with this, I'd be really interested to know what you're doing.
- simonw 2y agoYou'd need to roll your own migrations. I have a version of that for SQLite written in Python, but I'm not sure if you could run that in Durable Objects - maybe via WASM and PyOdide? Otherwise you'd have to port it to JavaScript. https://github.com/simonw/sqlite-migrate https://github.com/simonw/sqlite-migrate
- bluehatbrit 2y agoAppreciate the response (and the blog post itself)! I probably worded my question poorly, but I'm more wondering about executing schema migrations against a large number of DO's as part of a deployment (such as 1 per customer). I suppose the answer is "it's easier to have 1 central database/DO", but it feels like this approach to data storage really shines when you can have a DO per tenant.
- simonw 2y agoA pattern where you check for and then execute any necessary migrations on initialization of a Durable Object would actually work pretty well I think - presumably you can update the code for these things without erasing the existing database?
- 2y ago