4 ms·
I agree with avoiding *kwargs, but I wonder what we should do rather than to pass around dataframes if we have a program that works with dataframes? I get that
by bjornasm 3y ago
I agree with avoiding *kwargs, but I wonder what we should do rather than to pass around dataframes if we have a program that works with dataframes? I get that there is a downside as the actual types are hidden inside the dataframes, but I am unsure if the tradeoff is worth it by unpacking the dataframe into arrays or objects for every function call and return, if the functions require them to be dataframes?
- sgarland 3y agoYou can define nested types for args and return values: def foo(bar: dict[int, dict[int, str]]) -> list[tuple[str]] In older (< 3.10 ?) versions you have to import the classes from the typing module, but other than that it’s the same. You can also define types with the Generic class if you’d rather. With live checks while you write, it’s pretty easy to do this as you go.
- mjr00 3y agoSadly I haven't found any good ways to strongly type dataframes, so my solution is mainly cultural: just make people aware that they shouldn't use dataframes unless it makes sense, as opposed to Python codebases where dataframes are the default data type that gets passed around. I've tried to limit dataframe usage only to things like reading data from CSVs, then converting them to typed dataclasses before further processing. It also depends on your use case. pandas is obviously very flexible and easy to use, so it might be worth the tradeoff to keep using it if your main use case is interactive Jupyter notebooks run by data scientists, or dealing with highly variable input data. But most of my Python is backend data pipelines that work on predictable input data, where reliability is important, so I default to dataclasses, only using pandas in rare situations with a lot of validation code surrounding it.