9 ms·
I'm adamantly pro dry-run and like OP I've found that you really need to design for it. From my experience, OP's two main ideas are spot on: librarification and
by amirkdv 5y ago
I'm adamantly pro dry-run and like OP I've found that you really need to design for it. From my experience, OP's two main ideas are spot on: librarification and pushing dry-run logic deeper into library code.
Here are two other things I've found:
1. Regardless of where you push your run/dry-run dispatch to, the underlying "do it" logic really needs to be factored properly for individual side effects (think functional). Otherwise, you inevitably end up with loads of "if dry-run / else" soup or worse, bugs _caused_ by your dry-run support.
2. You still need safety nets for when someone is doing dangerous operations without dry-run. Pointing and calling [0] is a great trick for this. For example, say your tool is about to delete X, Y, and Z. Instead of a simple "Yes" confirmation which the user would quickly end up doing on auto-pilot, you could have a more involved confirmation like "Enter the number of resources listed above to proceed" and then only proceed to delete if the user enters "3".
Very curious to hear about design idioms folks have come up with!
[0] https://en.wikipedia.org/wiki/Pointing_and_calling https://en.wikipedia.org/wiki/Pointing_and_calling
- kbenson 5y agoUsually I just use dry run directives to make sure stuff is working right when I first develop it, so all that is worked out from the very beginning. I also like to develop with a lot of debug output gated behind debug as well, and just leave that there for the inevitable but at some point that would make me add it if not already present. What I don't like is passing those as params as args. Currying this stuff around is very error prone and cumbersome. I prefer to set environment variables for DRY_RUN and DEBUG (with debug accepting higher integer levels for more debugging). This works wonderfully for me since I never have to worry I didn't pass something along correctly, I'm always asking the global set at runtime.
- Buttons840 5y agoSo, for a top level bit of code to pass arguments to lower level bits of code you set environment variables? Seems a bit odd to me Using environment variables for the user to configure their environment or pass information into a program seems normal, but code passing arguments to other bits of code in the same language with environment variables seems odd.
- raziel2p 5y agoIn many programming languages, environment variables are a mutatable dictionary/map, which makes it the ultimate place to store global variables. It's definitely my guilty go-to-hack when I'm not up for refactoring everything to be more functional and/or take a dry_run function parameter everywhere.
- Buttons840 5y agoIt's a global mutable map that multiple uncoordinated processes can change. If you want a global map, just make a global map. Or just make global variables, since the variable namespace itself is a map.
- kbenson 5y agoAll that assumes you have some stuff to set up that global map, and that means that setup code is a requirement. I end up writing lots of library code. Sometimes that library code is called from within a command line utility I created, sometimes it's called from a web service, sometimes it's just a small driver script because I'm not developing or testing the utility, but the library code itself. I can make that include to set up the shared global and try to make sure it's included in all instances and all ways I want to use the code, or make all the call sites resilient to it not existing, or I can just use the included OS mechanism for doing this and since that's always available, I get it for free. Also, dry run mode isn't necessarily something you want set in a config. It's generally something you run once or twice prior to running for real (the normal case) or while in development/debugging. It's not something you would want to set in code and accidentally forget and push live, and generally a good dry run mode will look like it succeeded without actually succeeding, mocking responses that would fail along the way, because you aren't testing one small thing you're testing a workflow of some sort generally which has a few steps. That said, I fully admit the trade-off might go a different way for different languages. Using a compiled strongly typed language may mean there's enough bits to check that you need to write a debug/dry run helper function to make it convenient, so there's not a lot lost by requiring setup in that as well. But for something like Perl (and I assume Python and Ruby and JS, to almost the same degree) where I can do: warn "Calling out to foo() with args: " . Dumper($args) if $ENV{DEBUG} and $ENV{DEBUG} >= 2; foo($args) if not $ENV{DRY_RUN}; and it will be completely valid, obvious and idiomatic with zero additional work, there's a real draw to using environment vars for these two specific cases (even if not for all config).
- d4mi3n 5y agoThere are other approaches to this as well: 1. Passing around a configuration or context (similar to your gripe around params) 2. Referencing some kind of global configuration (a la env vars) 3. Referencing a local or scoped configuration (for example, a method of an object can check instance variables that dictate behavior) I personally prefer either a passed context or a local configuration; I find both easier to test in isolation. Global contexts have their uses, but tend to become problematic when they clash with other libraries or tools that may also be present in the execution environment.
- kbenson 5y agoThose are good approaches, and I wouldn't try to use manually set env vars for most things. I do tend to think they work very well for debug an dry run options though, because those are generally ephemeral, and things you might want to set ad-hoc in different environments easily without changing the config for everything in that environment, or passing around a special config which may not be updated when the real one is. That said, I'm not married to it, if I saw something that seemed obviously better, I would switch. I also suspect that different languages may make one approach easier/better than others based on their capabilities, idioms, etc. In many scripting languages, accessing an environment variable is extremely easy. In some compiled or more strictly typed languages, the access and conversion to the expected type might be cumbersome enough to do on site that it's worse, and if you are standardizing in come parsing routine, that might tip the benefits in favor of some global context that is used instead.
- hinkley 5y agoI have some code that lacks permissions to run to completion except on Bamboo or other servers, so I practically have to have a dry-run mode anyway.
- IgorPartola 5y agoWhat I like is a system where you first create a plan for what you will do, then there is an executor that will execute the plan. So --dry-run just doesn't run the executor. Of course, depending on what you are working with that might not be possible, but if you can design it like so, do it. It also makes everything nicely decoupled.
- cortesoft 5y agoYeah, not sure how this would work with any workflow that caLLs out to another service and uses the result of that to determine the next step, if that intermediate call changes state somewhere.
- dundarious 5y agoThen you can only really dry run that first phase anyway.
- minitoar 5y agoIndeed, the other service then needs to support dry runs.
- ItsMonkk 5y agoI think a large part of the next 10 years is us figuring out that frameworks are an anti-pattern and that everything is going to have to move to libraries that abide by CQRS.
- rocqua 5y agoDid not know what CQRS is, did some light googling[1]: CQRS (Command and Query Responsibility Segregation) is an alternative to a simple CRUD-based database interface. CQRS separates reads and writes into different models, using commands to update data, and queries to read data. Commands should be task-based, rather than data centric. ("Book hotel room", not "set ReservationStatus to Reserved"). Commands may be placed on a queue for asynchronous processing, rather than being processed synchronously. Queries never modify the database. A query returns a DTO that does not encapsulate any domain knowledge. The models can then be isolated, as shown in the following diagram, although that's not an absolute requirement. [1]
- whateveracct 5y ago> Regardless of where you push your run/dry-run dispatch to, the underlying "do it" logic really needs to be factored properly for individual side effects (think functional). Otherwise, you inevitably end up with loads of "if dry-run / else" soup or worse, bugs _caused_ by your dry-run support. --dry-run is one of the best motivations for using granular, extensible effects when you constrain literally every bit of IO your program is allowed to do, you can now 100% know you've stubbed them all out when doing a dry run
- lytefm 5y ago> 2. You still need safety nets for when someone is doing dangerous operations without dry-run. Another important property is idempotency. Especially if the script involves network requests or moving files around, you'll want to reach the goal by re-running in case something breaks half-way.
- catlifeonmars 5y agoIs idempotence the right word for this? Honest question. This is more like eventual consistency after a small number of operations, where each individual operations is idempotent.
- GuB-42 5y agoDry run is among the things you really need to design for, others are: - cancel - undo/redo - progress bars Progress bars with accurate timing are notoriously difficult, or even impossible to get right. But even regardless of timing, having a progress bars that really shows progress and doesn't freeze is hard. Every slow operation has to have some sort of callback mechanism to update progress, and you have to know the number of steps in advance. Undo requires, for each operation, to know how to roll back. You also need to have every operation go through the undo stack. Another option is to have an efficient snapshot system, which may be just as hard or even harder. Cancel is actually the hardest to do right because it combines the difficulties of both the progress bar and undo. You have to have a way to interrupt a long operation at any time and get back to before it started. And because "cancel" is most often used when things go wrong (ex: disk full, bad connection, ...) you have to be very careful with error handling.
- monsieurbanana 5y agoThe thing I've gather from this whole thread is that I made the right decision by choosing a functional, immutable language.
- hinkley 5y ago> I've found that you really need to design for it It's an architectural step a lot of people skip, and by doing so you often end up in a situation where the only way to make things faster is via caching, and chasing caching bugs for the rest of your tenure. One of the less appreciated aspects of model-view-controller is that usually it demands that you plan out your action before you do it - especially when that action is read-only. Like a cooking recipe, you gather up all of the ingredients at the beginning and then use them afterward. By fetching all of the data early, you reduce the length of the read transactions to the database, allowing MVCC to work more efficiently. You also paint a picture of the dependency tree in your application, up high where people can spot performance regressions as they show up in the architecture - where there's a better opportunity to intercede, and to create teachable moments. To make dry-run work you are best served by book-ending your reads and your writes, because if the writes are spread across your codebase how do you disable all of the writes? And if any reads are after writes, how do you simulate that in your dry-run code? It won't be easy, that's for sure. The problem is that we tend to write code stream-of-consciousness, grabbing things just as we need them instead of planning things out. This results in a web of data dependencies that is scattered throughout the code and difficult or impossible to reason about. This is where your caching insanity really kicks into high gear. To my eye, Dependency Injection was sort of a compromise. You still get a good deal of the ability to statically analyze the code but you can work more stream-of-consciousness on a new feature. But it does rob you of some more powerful caching mechanisms (memoization, hand tuning of concurrency control, etc)
- Smaug123 5y agoThis is the first comment I've used HN's "favorite" feature on. How wonderfully insightful.
- suls 5y agoI think this is a case where the Free Monad could help a lot. I've only ever built toy-examples but I really liked the way it allows you to structure the program's plan separately from interpreting it. A dry-run mode could be implemented as a side-effect free interpreter .. without ever having to touch the plan or the program's grammar.
- 5y ago
- ryanbrunner 5y agoOne thing that I've used in the past for safety that has been extremely helpful is to always log results to a detail level that allows reversal of whatever you did. For example, for a big DB update, create a CSV (or just write to STDOUT) all affected records along with their old data. If you get into this pattern enough, and your operations are similar enough, you can even standardize on the format so you can have a common tool that can undo an operation for you given the log. Certainly it's not something that can be applied for every use case, but when it can, it has so many benefits - an audit trail of what actually happened, ability to undo, and a common structure and approach for what dry runs should do (just output the log)
- throwaway_2047 5y agoI have yet to see anywhere implemented like this, but IMO mutation/deletion should be on dry run mode by default.