4 ms·
Caveats: 1) this may not be what grandparent was trying to do and 2) I have limited experience debugging "regular" Docker, let alone firecracker or whatever Fly
by mostly_lurks 4y ago
Caveats: 1) this may not be what grandparent was trying to do and 2) I have limited experience debugging "regular" Docker, let alone firecracker or whatever Fly.io does.
With that said, I succesfully deployed a simple Rust/Actix/Sqlite app with a Dockerfile to Fly.io. I then thought I'd try out litestream. For reasons I'm still not sure about, having `ENTRYPOINT ["litestream replicate … -exec myapp"]` resulted in immediate kernel panics.[1]
As I was debugging, my instinct was to want to `fly ssh console` into a running container to see if running the same command from the shell produced any clues for further debugging. Though I understand this is the wrong mental model, the thought was something like "well, I know everything else works, I just need Fly to ignore this failing command so I can have a minute to poke around." To do this, I ended up just removing litestream from the ENTRYPOINT so the deploy would succeed, then I could SSH in and play around at the shell to see what was going on.
Again, I have no idea whether this is the same sort of problem the other person was having, but for my case, what would be helpful is probably not changing how `fly ssh console` behaves, but perhaps some documentation of suggested debugging techniques in case your app is failing to start.
[1] Putting the exact same command into a "start.sh" and making that the ENTRYPOINT worked fine so that's what I ended up doing.
- tptacek 4y agoOne of my biggest UX nits with Fly (I have no excuses, I have all the access I need to go fix this myself) is that we "kernel panic" when your entrypoint command fails. Of course, our kernel is not really panicking --- we've just run out of things for our `init` to manage, so it exits, and when pid 1 exits, so does the kernel. But you get the terrifying stack dump. We can clean this up, so that you get a clearer, simpler error ("your entrypoint exited, here's the exit code, there's nothing else for this VM to do so it's exiting, have a nice day"), and it's been on the docket for months. We'll get it done! We could conceivably add a flag for our `init` to hang around waiting for you to SSH in after your entrypoint exits. But that's clunky and complicated. Usually, you want your kernel to exit when your entrypoint fails, so that your service restarts! What you should do instead is push a container that has enough process supervision to hang around itself. Here's a doc: https://fly.io/docs/app-guides/multiple-processes/ https://fly.io/docs/app-guides/multiple-processes/
- mostly_lurks 4y agoThat all makes sense. I think context matters. Yes, when I have a service that's been up and running for some time, if that service fails, I want my service to restart. But when I'm trying to get a service running for the first time and I'm not sure that I have the right command in the entrypoint, the right arguments to that command, or the right supporting files in place, or the right libraries installed, or the right file permissions, …, well, then I don't want things to just blindly restart, I want a handle and some information so I can figure out why it isn't working. ETA: I recognize that your link to docs about running a supervisor addresses this problem. For me this raises some interesting questions. Like, I understand why Ben would implement `litestream exec` but maybe it would be better to steer users to a proper supervisor? Separately, what if it's the supervisor that's failing? Now I'm back to seeing kernel panics and not having error messages or a shell.
- tptacek 4y agoUsually, when I'm debugging a container, I start with a `tail -f /dev/null` entrypoint, or something like that, and then just shell into it to run the real entrypoint to see if it's working.
- CoolCold 4y agoThat's probably closer to sysadmins magic, not average developer way of thinking when debugging.
- rtpg 4y ago> We could conceivably add a flag for our `init` to hang around waiting for you to SSH in after your entrypoint exits. But that's clunky and complicated This is a common thing in CI platforms, and the way they usually expose this is "run tests with SSH enabled", and they keep it open for 30 minutes/2 hours/whatever until a session closes. So if I have some app failing, being able to run `fly restart --ssh-debug`, having that first just sit around waiting for the app to boot, and then dropping into ssh would be a very helpful piece of UX. The main thing is cleanup, but y'all charge for compute! You can be pretty loosy-goosy on that one honestly.