4 ms·
I've worked on a lot of sophisticated data pipelines, in robotics, built customized wysiwyg editors, distributed systems, append-only logs with replicated data
by jackcviers3 4y ago
I've worked on a lot of sophisticated data pipelines, in robotics, built customized wysiwyg editors, distributed systems, append-only logs with replicated data types, in around a dozen languages. While I have found occasionally remote profiling useful for optimizing hot code paths, I have always found that trace-level logs and structured logging with aggregation and indexing is superior to debugging in most applications.
Essentially - println debugging is superior to interactive debugging, especially for intermittent bugs, data races, deadlocks, resource leakage, modeling and serialization bugs, binary or source code incompatibility bugs, and bugs that are catastrophic and rare or pathological in nature.
The problem with interactive debugging is that you typically only have a view into not only a single process of a system, but also only a single concurrent execution thread of the single process.
When you have structured logging and aggregation and indexed log search, you have a much broader view of the systems under investigation and can see the conditions causing the bug with a much higher frequency than with interactive debugging.
For any non-trivial bug in any sufficiently large input or program you have to have insight into the root cause via intuition to cause it to occur quickly enough to inspect the program via the debugger. This obviously doesn't scale well. It's much better to have a detailed, structured log mode built into your systems that is capable of being enabled on demand - via environment variable, remote procedure call, request header, data packet, or command flag during deployment. This allows you to observe the system behavior and incur the cost of the detailed logging on demand, much like you would when connecting via remote debugging, but on a system of systems scale, increasing the probability of capturing the bug's root cause over stepping forward and backward through interactive execution.
Additionally, care should be taken to program in a value-centric way when you can get away with it performance or space-wise. Essentially - use FP principles of immutability and programs as values and composition to build your programs where you can. Values have a unique property in that they are serialization state that can be inspected or substituted in place of references to the value without changing the behavior of a program at runtime. This often simplifies debugging even in complex side-effecting concurrent algorithms.
When you can't, comprehensive argument and shared state change log events that can be shared amongst many members of the development, stakeholder, and operations teams asynchronously helps to narrow the surface area scan of the system you are debugging.
Formal methods during modeling, design, and development can help to increase constraints and enable more comprehensive test suites, but property-based testing with randomized generated valid inputs also serve as a sort of automated preemptive debugger. Things like chance in the js world, quickcheck/hedgehog in haskell and their ports or siblings in other languages, like hypothesis in Python, fall into this category of testing tools.
Interactive debugging is the tool of last resort, and still serves a useful purpose when all other mitigation and system state capture methods have failed, but I personally find myself believing that if a bug's root cause hasn't been discovered by this stage I have made some set of mistakes in the design of the system, either in the application of its logical rules or in the basic assumptions fom interpretation of business requirements to the choices of abstractions/tradeoffs made in development. It's usually a bad day when I have to resort to interactive debugging.
- mjw1007 4y agoI think a "Omnicient"/"Time-travel" debugger with strong tracing support would be interesting. I'd like to be able to take a recorded run and add tracepoints after the fact, and be able to toggle them on and off, or add or modify conditions on the output, watching the trace change as I do it. When I looked at the Pernosco demo it looked like it had the capabilities to do this sort of thing, but only a rather primitive user interface.
- acemarke 4y agoMentioned it up thread, but this is _exactly_ what our Replay time-traveling debugger for JS lets you do! We specifically let you add "print statements" to any line of code from the recorded app, let you toggle them or change conditions, and recalculate the logged output every time you make a change: https://docs.replay.io/reference-guide/print-statements https://docs.replay.io/reference-guide/print-statements Longer-term, we have plans to implement functionality for "persistent object IDs" and tracing individual values over time. No ETA on that, but from what our runtime team has said it's feasible. (We actually have a partial hardcoded implementation of that that we're using to track anything that looks like a React internal "Fiber" data structure, and I used that to build some of our support for the React DevTools: https://blog.replay.io/how-we-rebuilt-react-devtools-with-replay-routines https://blog.replay.io/how-we-rebuilt-react-devtools-with-re... )
- jackcviers3 4y agoReply to Grandparent and this: Very cool and definitely useful to record and replay the session and modify breakpoint conditions for replay. I wonder if session recording like this could be safely brought to backend systems. You'd almost have to stub every remote call to allow for replayability in isolation of external services, and the potential for data leakage from a downloaded session is also a concern. I'm glad someone else is working on problems and tools like this. I should note I'm not disparaging the use of debuggers in my comments, but lamenting the nature of the process of debugging the types of systems we deploy today, and the largely non-interactive nature tradeoffs we've made for scalability, reliability, availability, and operability. Sometimes I long for the simplicity of a fully self-contained local system, but I know we just don't live in that world anymore, and haven't for a long time now.