10 ms·
Best Practices for Working with Configuration in Python Applications
- alanfranz 6y agoWhile most points are valid, I feel some pieces are missing in this article. What should I finally do? How to put everything together without creating an hard-to-maintain mess of casting/parsing/configuration? Should I manually cast strings to integers (or other types) for all and each value I parse? Where do I keep default values? It's cumbersome to have them embedded in code for get(). I usually want a) a default configuration kept in a file b) a way to override that config with other files, but only for certain parts (I don't want to rewrite the configuration every time), c) a way to override the config at launch time (e.g. from cmdline) The fact that Python is dynamically typed/type hinted only makes it harder than statically typed languages at configuration, where most configuration libraries instantiate an adequate type conversion function to put a string somewhere. I ultimately found that the latest solution (parse from json) is good enough for most use cases for points a) and b); since json is typed, a decent conversion can happen for 95% of the use cases (for the others, just use a string and manually parse). A sidenote to the author if he/she's reading: datetime.date objects, just as python's own naive datetime objects, are dangerous objects that can lead to unpredictable results when used with actual time-handling code. I wouldn't use them anywhere in my Python code.
- ggolu2 6y agoWhat’s naive and dangerous about Python’s Datetime objects?
- alanfranz 6y agoYour question implies that you don't know about the nuisances of the datetime library :-) (see https://docs.python.org/3/library/datetime.html https://docs.python.org/3/library/datetime.html it's the first paragraph!) Python datetime objects, by design, can be naive or timezone-aware. Timezone-aware datetime objects are OK; they identify a certain instant in time. Naive datetime objects are Python-only abstractions (AFAIK) that don't identify anything in the real world; they're highly error prone, because there's no "right" way to use them. They sort of work properly only if used in a very limited scope - e.g. your own code only, for small sections - but they're risky because they're not a different type (with respect to tz-aware), and it's hard to tell what any code accepting a datetime does if passed a naive object. Some libraries like java.time DO have a similar concept (e.g. LocalTime, LocalDate) but they keep it well separate from the "real" concept (e.g. Instant or Date in Java) so you can't use them accidentally. Example: you pass a naive datetime object to any library which must translate it to an instant, like an ISO string with a well defined timezone. What does the library do? Throw an exception? Associate an arbitrary timezone (e.g. UTC)? Associate the local, current timezone? There's no "correct" behaviour.
- aucontraire 6y agoI agree that python datetime objects are problematic, but for the opposite reason. It is tzinfo that is the sneaky disaster, the plain datetimes are fine. Transparent timezone awareness always fail, unless you are 100% certain that a tzaware datetime object will remain uncoverted from the very top to the very bottom of the stack and all the way up again no matter who is reading and what they are doing. For longterm minimization of pain, bugs and effort, you convert datetimes to UTC as early as possible and take them back to some localized version as late as possible (in the frontend, for a webapp, so that the backend never needs to know there is such a thing as timezones (except for separate validation and correction routines, since timezone definitions always end up being incorrect to some degree when you use them at scale)). If the localization of the datetime is an essential aspect (such as the departure time of a ship leaving port), you store a UTC value together with a record of the location. Only at the latest possible moment of processing, should you do a lookup on the location data to make a local time. Obviously, there will be exceptions to this rule. If you batch process billions of timestamps under a tight deadline and must do calculations in local time, it might make sense to have the values persisted localized.
- alanfranz 6y ago> I agree that python datetime objects are problematic, but for the opposite reason. It is tzinfo that is the sneaky disaster, the plain datetimes are fine. Why naive datetimes should be fine? How are they fine? What do they represent? > For longterm minimization of pain, bugs and effort, you convert datetimes to UTC This COULD work if naive objects had an IMPLIED UTC in their contract - e.g. naive objects are declared as ALWAYS UTC. Your argument fails as soon as you pass a naive datetime object somewhere in a library/framework and it gets accepted, and/or you try serializing it without augmenting with a TZ. As I said, naive datetime only works if you control 100% of their usage. No libraries, no external points of contact. And, the reason because tz-aware objects sometimes fail for the opposite reason (e.g. libraries assume naive objects) is a fault of the API design (they're not distinct types), but the problem lies in the existence of the naive version, not vice-versa. For the records: in the backend I always use tz-aware datetime objects with a fixed UTC timezone. That's the best way IMHO not to get crazy with time problems in Python. So, your points about datetime handling are all valid and correct (timezone is mostly a "UI problem" and should not leak into the backend) but they don't prove your "naive works better" argument.
- kissgyorgy 6y agoBasically everything.
- thomk 6y agoWho is the targeted editor of a config file? If it's another programmer, just use config.py and make life easy on yourself.
- alanfranz 6y agoNo. Using a programming language for configuration is terrible. If code can be used, it will be eventually used and it will be impossible to understand what's happening until runtime. Configuration should be clear and not subject to modifications to the software. If somebody wants to use an external preprocessing tool, if using a standard format (json) he/she can.
- thomk 6y agoConfiguration, by definition, modifies the software. How much it modifies the software is determined by the software author. If there aren't limits imposed in your example (json) you're still going to get unexpected results. You'll have to limit what's accepted. Also, what if I allow a config.py file then translate that to .json then back to python? That is no better than just allowing the consumption of a python source file and simply ignoring any disallowed code (again, limiting what is accepted). I know it's vogue to say configuration as code is bad; that simply has not been my experience.
- alanfranz 6y agoI don't know whether it's vogue or not, I think it's just a bad idea, because you can't restrict HOW MUCH is is modifying the software, and then people needs to know the nuisances of your programming language. What if I need some configuration to be user (not admin) configurable? What if I don't want to learn programming language X just to use a software? > If there aren't limits imposed in your example (json) you're still going to get unexpected results. It's very unlikely for a JSON string to be parsed as a class, or to execute any code, unless you do eval() or do unsafe deserialization. A Python config could do anything. > You'll have to limit what's accepted. This is much, MUCH easier said than done. How would you do that in Python? Google App Engine had a sort of "restricted python" idea, and it was hard to implement AFAIK. Same thing for Zope/Plone (there was some templating with restrictions, don't remember the precise system). Then, you'd need do document such "restricted python". > Also, what if I allow a config.py file then translate that to .json then back to python? I suppose the first config.py is under the control of the user, and the second is under control of a software author. This is ok, because the "second python version" can perform validation and object construction. Python configs can be OK (but can get messy) if the software you're building is mostly internal, and few people modify it and they all know what they're doing. As soon as you've got a large enough team and or userbase, IMHO executable configurations are painful.
- zelphirkalt 6y agoI think argparse covers most of the points mentioned as desirable in this article. * validate at start (using the type keyword argument for add_argument) * access by name as identifier, not string And what is more is, that default values are stored with the configuration, plus you add a help text for telling everyone what the argument is for. One downside is that you command line call gets longer.
- xiwenc 6y agoPretty good guidelines. I like the idea of having a configuration close to the class (perhaps module-level depending on the project size) that uses it. With dataclass the class definition is fairly clean. In addition to that, I'd consider using https://docs.python.org/3/library/dataclasses.html#post-init-processing https://docs.python.org/3/library/dataclasses.html#post-init... for business-specific validations.
- hathym 6y agoit's worth looking at python's own configparser[1] before rolling your own. [1] https://docs.python.org/3/library/configparser.html https://docs.python.org/3/library/configparser.html
- kingosticks 6y agoIsn't that basically the same end result as using json.loads except a different format (that has no actual spec).
- frumiousirc 6y agoJSON does not support comments nor string interpolation. Python ConfigParser language does.
- kingosticks 6y agoTrue, you do get string interpolation but the comment support in ConfigParser isn't very good. Although actually they may have fixed some of that in Python 3 but I'm still using workarounds. To be clear, I am not suggesting using JSON for config, I think that would be my last choice. My point is that ConfigParser isn't really an alternative to rolling your own if you want decent validation etc (those spec files are horrible to use). You very quickly need to start extending ConfigParser to the point where you've started rolling your own. And at that point you'd be better off with one of the other (tested) solutions already suggested.
- nomel 6y agoWhat's wrong with the comment support? You can't have comments at the end of a line, but that's sort of the nature of supporting arbitrary strings as values. I don't want my users to have to quote or escape special characters if they happen to want to use them. They're not programmers. # The note to display note = Our #1 customer! rather than name = Our \#1 customer # The note to display. or name = "Our #1 customer" # The note display
- fermigier 6y agoQuick list of Python libraries that help with application configuration: - Python application configuration -> https://github.com/edaniszewski/bison https://github.com/edaniszewski/bison - Configuration with env variables for Python -> https://github.com/hynek/environ_config https://github.com/hynek/environ_config - Configuration library for python projects -> https://github.com/willkg/everett https://github.com/willkg/everett - Strict separation of config from code -> https://github.com/henriquebastos/python-decouple https://github.com/henriquebastos/python-decouple This is from my personal notes. See also: https://github.com/vinta/awesome-python#configuration https://github.com/vinta/awesome-python#configuration Anything else that we've missed ?
- andyljones 6y agoThe two big ones in machine learning at least are Google's Gin: https://github.com/google/gin-config https://github.com/google/gin-config and Facebook's Hydra: https://hydra.cc/ https://hydra.cc/
- kyawzazaw 6y agothe name usage of "gin" is quite common, it seems: https://github.com/gin-gonic/gin https://github.com/gin-gonic/gin
- ungawatkt 6y agoMarshmallow. You can use its schema validation for any dict/json, which makes it a nice fit for validating json config files (which mitigates some of the json concerns from the article). Just immediately move the json.reads through a schema validate, build some classes around it for different config files. marshmallow.readthedocs.io/en/
- mianos 6y agoI use a similar library called Schema at https://pypi.org/project/schema/ https://pypi.org/project/schema/ I love the expressive nature of it. Just this week I was validating both yaml and json configuration data with it. One thing missing from this article is to always use a proper external key value store for everything but the stage (dev test prod etc) and the connections to the kv store. Config files on disk and even in environment variables suck for anything but the most trivial platforms.
- jry1234 6y agoKensho has probably the best solution to this problem I've seen so far: https://github.com/kensho-technologies/grift https://github.com/kensho-technologies/grift Handles typing really well, as well as config defaults and fallbacks, giving you the ability to configure your app a few ways, and fall back on other configs if something isn't specified.
- anentropic 6y agoShout out here for pydantic BaseSettings https://pydantic-docs.helpmanual.io/usage/settings/ https://pydantic-docs.helpmanual.io/usage/settings/ That provides typed and validated auto-loading from env vars. I have been quite happy with that in conjunction with an optional .toml file, to do flexible config cleanly and simply like: import toml from myproj.conf.types import Settings # a pydantic BaseSettings model try: _config = toml.load('myproj.toml') except FileNotFoundError: _config = {} settings = Settings( **{key.upper(): val for key, val in _config.items()} )
- thomk 6y agoUnless the end user is not technical, use a .py file and force them to subclass your Configuration class which has an __init_subclass__ method so you can enforce rules. When you are ready to move to a more generic solution, your .config or .yml file can generate these. The advantage here is both flexibility (it's Python) and control (allow/disallow whatever you want). If you need nested items, use nested classes.
- ralston3 6y agoThis. Until you the app reaches the level of advanced yaml config files for cloud deployments, it’s really hard to beat a “config.py” that does a single read of all your ENV_VARS at startup
- markus23 6y agoAs already written by others, the article does not go very deep and is missing many essentials. What I was mostly missing is more about keeping configuration parameters as simple as possible. A much more detailed best practices can be found here: https://www.libelektra.org/ftp/elektra/slides/cm/ https://www.libelektra.org/ftp/elektra/slides/cm/
- mixmastamyk 6y agoAlong these lines and unsatisfied with current solutions I started this project, "Turtle Config." It is format-agnostic and supports type checking as well: https://github.com/mixmastamyk/tconf https://github.com/mixmastamyk/tconf I'll see if I can add any advice this article gives; feedback would be helpful.
- okomestudio 6y agoI have been working on a minimalistic application config library for Python, aiming to consolidate config loading from files, environment variable, and command-line argument parsing. It's an alpha, so please feel free to provide feedback if app configuration has been your pain point. https://github.com/okomestudio/resconfig https://github.com/okomestudio/resconfig
- jayjader 6y agoFYI that sounds like exactly the same feature set of 'python-decouple'.
- okomestudio 6y agoIndeed "python-decouple" looks like it serves similar niche. (I didn't know of the package. Thanks for letting me know.) I think I'd like to target a smaller niche though, someone writing a small applications, with a little more flexibility in things like YAML support and dynamic loading. Unless "decouple" eventually supports similar features, I want to keep experimenting.
- _pastel 6y agoAs an ML engineer working in Python, I keep running into a problem: If parameters are defined close to usage, and strongly typed, then it's hard to cleanly search for good configurations of the parameters. Especially for fancier search strategies, you want all parameter lookups to go through a single file. On the other hand, there's a lot of code churn until an ML pipeline is finished. And errors from typos and type violations will often only show up after hours of training. So it's also painful to try to keep a separate, loosely-typed parameter file in sync. So far, my compromise is to: (1) on a first pass, define all parameters as global variables at the top of the files they are used in (2) once mostly code-complete, pull them into a separate file that tracks initial values, current values, and search ranges. Make all usages go through a lookup where the key is an enum, but the value is untyped: def param(name: ParamName) -> Any: return params[name].current_value Which is not ideal. Does anyone else keep running into this problem and have a better solution?
- silviogutierrez 6y agoAlas, MyPy doesn't have the concept of "keyof" and mapped types like TypeScript. So in place of that I would: 1. Define a variable called Params = Any; 2. Liberally use Param["foo"] and Param["bar"] anywhere. 3. Once you stabilize, reimplement Params as a TypedDict. You'll get failures if accessing any invalid key. You can also use a NamedTuple if you prefer. If you insist passing a param name, then you'll have to create a big list of Literals for param with every key in your dict. So Param = Literal["foo", "bar"] etc.
- _pastel 6y agoThank you! It's surprising that mypy can typecheck based on string keys like that. Cool!