3 ms·
How does this work when there are many processes running livebook, behind a web server reverse proxy like nginx, there is no sticky sessions in place (because h
by ransackdev 3y ago
How does this work when there are many processes running livebook, behind a web server reverse proxy like nginx, there is no sticky sessions in place (because hot spots on certain upstreams kinda defeats the purpose of a load balancer), and the user that kicked off that async pdf job on PID 1, inside container 1, refreshes their browser and is now connecting to container 2, which is running a different Linux process than where the pdf task is running in a thread? Now add in the fact that container might be on a different node/vm in the cluster, different AZ, or different region/data center.
Background worker systems like Oban or sidekiq are where you want to run your async work imo. You can still do a pub/sub type notifier model to bubble up the job’s completion, or poll. Polling is perfectly acceptable for most apps and traffic levels and is simpler to mentally grok compared to websockets and persistent connections. Websockets require a central broker to pass messages through when you scale past one process and with elixir the options are the redis pubsub adapter, or for container based deployments or those wanting to use the built in distributed erlang functions, libcluster. More complexity. More points of failure. More potential centralized bottlenecks. More things to mentally model when debugging. More dev and prod drift. Just poll, until you can’t scale your web servers any more. It’s just an http request. It’s already implemented if you’re serving web traffic.
It’s not clear based on what you wrote but are the pdf’s contents being passed in the complete message? That sounds like an OOM termination that’s hard asf to track down just waiting to happen. PDF should go into blob storage somewhere and the task should bubble up a complete message, which then is used to signal the retrieval and proxying of the file to the user. If you have to background the pdf generation, it’s slow. If it’s slow to generate a document, it’s large. You don’t want that to be pushed through memory because your system stability is then coupled to not only the contents of a document that could exceed the memory of the system, but also is at the mercy of how many users/requests are all generating pdfs at the same time, and of course you’ll have that one guy spamming the “generate” button and queuing up 300 jobs because we didn’t need indicators on the UI and who adds uniqueness constraints to internal tools ;) Disks and file systems are for passing very large messages on a system
- iudqnolq 3y ago> How does this work when ... there is no sticky sessions in place (because hot spots on certain upstreams kinda defeats the purpose of a load balancer), and the user that kicked off that async pdf job on PID 1, inside container 1, refreshes their browser and is now connecting to container 2, which is running a different Linux process than where the pdf task is running in a thread? The user destroys their first liveview instance and gets a new liveview instance. You may not have properly set up bidirectional links so the destruction of this first liveview kills the first background job, or it may run to completion and be discarded. This might be exactly the semantics the user wants. If the user expects state to stick around that would need to be implemented separately. > If you have to background the pdf generation, it’s slow. If it’s slow to generate a document, it’s large. I'm not sure it's quite that slow. It's too slow to freeze the entire UI while it generates. But it's fast enough blocking a http response until it finishes often works. A simple system for transient background processing like Task.async might be appropriate. > Disks and file systems are for passing very large messages on a system We don't know what the type of the response message is. It could well be a filename in a temporary dirextory. That would be the simplest way to implement a background task that shells out.
- wmanley 3y ago> Websockets require a central broker to pass messages through when you scale past one process I believe that this is one of the big advantages of SSE vs. WebSockets. Unlike websockets, an SSE connection is just an HTTP request, just one that has a particular format and is conventionally long-lived. You don't need a broker and you can probably implement it yourself on top of whatever HTTP server library that you're already using. And it already contains functionality for dealing with network unreliability [1]. I've often thought that it would be really good to have better support for SSE in HTTP caching HTTP proxies. Then you could have nginx for example sitting infront of your SSE endpoints serving multiple clients with a single SSE stream to your backend. Just based on the caching headers and a little knowledge of `Last-Event-ID`. [1]: You can encode any state you require to resume the connection into Last-Event-ID.