19 ms·
Python at Scale: Strict Modules
- avip 7y agoIt's very important to think about objects lifecycle management. It's also important to... use pytest fixtures instead of arbitrarily patching around in tests.
- brenden2 7y agoIt still blows my mind that people don't use strongly typed languages in the first place and spare themselves from all this future pain. My guess (based on my experiences) is that companies wind up in this position from having inexperienced people building early versions of products instead of hiring experienced engineers (who are usually more expensive).
- b3orn 7y agoPython is strongly typed, just not static.
- brenden2 7y agoPython uses duck typing: https://en.wikipedia.org/wiki/Duck_typing https://en.wikipedia.org/wiki/Duck_typing I would categorize it as a subset of dynamic typing, and that's what Wikipedia says too.
- rectangletangle 7y agoDynamic typing vs. static typing is on a different axis than strong vs. weak typing. Python is a strong dynamically typed language, with some "static lite" features introduced in Python 3. Dynamic typing means that types can be changed arbitrarily at runtime, compared to statically typed languages which define all types at compile time. Strong/weak means that type coercions rarely/never happen automatically. For instance JS has some interesting behavior enabled by weak typing `[] + [] -> ""`. Whereas Python rarely coerces things for you. The division operator in Python 2 was strongly typed, while they changed it to weak typing in Python 3 (inline with the practicality vs. purity convention).
- tom_mellior 7y ago"strong typing": everything has a type and cannot be accessed at some other type; "static typing": everything's type can be determined statically (according to one definition) In Python everything has a type, and you can't use a float as a list, for instance. It's correct to call it both strongly typed and dynamic, those are not antonyms.
- ainar-g 7y agoIt's a constant struggle against the current. Dynamically-typed languages are often “good enough for the time being”. I have the same issue explaining to our C/C++/Obj-C team why they should use static (Clang-Tidy, Infer, PVS-Studo) and dynamic (ASan, MSan, UBSan) analysis tools. They just keep giving me basically the same response of “I am a good programmer, and my code is good, and shame on you for even daring to think that a mere machine could find bugs in my code!”. I don't know what kind of status anxiety causes it. It also makes me think about what kind of other I am missing because of the was I keep thinking that I do that thing well-enough myself.
- ken 7y agoI'm confused. It should be easy to demonstrate the benefit, if there is one. Just show them the bugs! For me, it's not "status anxiety". It's simply not worth the effort. The last couple static analysis tools I ran on my programs, I spent a while getting the tool to not-crash (because even though the authors obviously had a static analysis tool themselves, they either didn't bother to run it on their own code, or it wasn't good enough to find actual issues). These tools flagged only a couple issues, and almost all of them were places where it couldn't really cause any problems, but the type system was not strong enough for me to prove why it couldn't go bad. So I spent a while sorting through false-positives. I'm not going to spend hours with a tool to find only a couple (real) bugs, which no user has ever reported seeing, and which I've gotten no automated crash reports about. I have much better uses for my time.
- ainar-g 7y agoSee, that's another thing that a lot of people don't understand about static analysis. It's not just there to find bugs in existing code, it's there to find bugs as you write or edit the code! Of course it won't find a lot in a tested code base. It's tested after all. But it immensely shortens debug time as you develop, and thus reduces testing time as well.
- zallarak 7y agoThis article is among the best argument for using a typed language I’ve yet seen.
- kbd 7y agoThis has nothing to do with types. It's more about static guarantees the language gives about module import behavior.
- nothrabannosir 7y agoIn OP's defence: > So that's a third pain point for us. Mutable global state is not merely available in Python, it's underfoot everywhere you look: every module, every class, every list or dictionary or set attached to a module or class, every singleton object created at module level. It requires discipline and some Python expertise to avoid accidentally polluting global state at runtime of your program. > One reasonable take might be that we’re stretching Python beyond what it was intended for. It works great for smaller teams on smaller codebases that can maintain good discipline around how to use it, and we should switch to a less dynamic language. > But we’re past the point of codebase size where a rewrite is even feasible. And more importantly, despite these pain points, there’s a lot more that we like about Python, and overall our developers enjoy working in Python. So it’s up to us to figure out how we can make Python work at this scale, and continue to work as we grow. Those are literal quotes from the article. That is quite damning. How did they get to this point? By starting when Python was appropriate, and taking it day by day.
- kbd 7y agoHow is "our developers really like Python even on a million-line codebase, despite its global mutable state requiring discipline" "quite damning"?
- iso-8859-1 7y agoDepends what you mean by "type". A type in e.g. Haskell specifies whether there are side effects.
- 7y ago
- allan_s 7y ago> This means that just by importing this module, we're mutating global state somewhere else. Yes, this ! That's why I hate Django and some flask app the most for, the fact that by importing a module, you're implicitly creating a database connection, and a lot of other magic stuff, which mean that now I can't import a constant defined in said module outside of `python manage.py` Also as said below in the article, suddenly it's much harder to handle smoothly the "the database is momentary unavailable" (because someone has put the line starting the database connection in the global space of a module somewhere) I much prefer frameworks/modules for which code is executed only once you invoke their "setup" function
- tummybug 7y agoI'm not sure about django but flasks Application object has a before_first_request method which takes a function designed to do this type of initialization operations.
- rectangletangle 7y agoI'm a huge fan of Django, but I always felt that this was true. I wish there was more of a push to decouple parts of the framework. Keep the magic, but allow usage without it.
- nerdponx 7y agoI much prefer frameworks/modules for which code is executed only once you invoke their "setup" function Django _does_ have a "setup" function. You can't import and use Django database connections outside of a running application without it. Flask also has a "run" method and does no i/o without it.
- heavenlyblue 7y agoEvery time I hear a comment alike parent's, it makes me think how many times a day I actually read a comment in the same fashion, but about something I actually know nothing about.
- 7y ago
- ben509 7y ago> How do we know that the log_to_network or route functions are not safe to call at module level? We assume that anything imported from a non-strict module is unsafe, except for certain standard library functions that are known safe. It's hard to know anything about the stdlib as it can be monkey patched, e.g. [1] That said, you could solve this with diagnostics; calculate signatures of stdlib functions and classes to find any known safe ones that were patched. Run that check in your test suite to find problematic imports. > If the utils module is strict, then we’d rely on the analysis of that module to tell us in turn whether log_to_network is safe. I like this. It seems far more usable than proposals like adding const decorators.[2] [1]: https://github.com/gevent/gevent/blob/master/src/gevent/monkey.py https://github.com/gevent/gevent/blob/master/src/gevent/monk... [2]: https://github.com/python/typing/issues/242 https://github.com/python/typing/issues/242
- ledauphin 7y agoI love the idea, but it feels like just an idea at this point. I'd rather read about them releasing their 'compile-time' analyzer and revealing their measurements for how much startup time it saves. In our codebase, we have pretty strict developer-enforced rules about not doing I/O at the module level, usually through the use of simple "Lazy" wrappers for module-level objects. I'd be curious to know what other approaches people have taken with Python here.
- rectangletangle 7y agoIt is an interesting approach, though I feel like this could introduce some nasty unintended consequences given how dynamic and introspective Python can be (admittedly I haven't studied this particular implementation). I always treated this a bit like single underscore private functions/methods, i.e., follow a convention that produces code that's easy to reason about, even if it's not strictly enforced by the language/compiler. So in practice this equates to separating out modules that mutate global state, and placing the majority of logic in "strict" modules that only declare a bunch of "pure" classes/routines. So the "non strict" code is really just a thin layer of wiring gluing everything together. For instance my Celery task files tend to be very thin.
- ledauphin 7y agowell, we also heavily use static typing, so you end up with something like my_db_conn: Lazy[DbConn] = Lazy(lambda: make_db_conn(...)) and MyPy will tell you if you're doing something silly when you try to use it. EDIT: After typing up this response and submitting I realize you were talking about their strict approach rather than ours. whoops :)
- timothycrosley 7y agoMore and more I want someone to create a new language that amounts to a strict subset of Python, with mypy built-in, and is compilable into machine code. Python has by far my favorite syntax, community, and in my experience leads to the greatest productivity. There just happens to be a lot of overly dynamic features, that aren't even used by most, but used just enough to hold back optimization and structural improvement.
- TylerE 7y agonim
- timothycrosley 7y agoI've been tracking nim, and would agree it's the most promising so far! I feel though that it's trying to be too flexible in many ways. Examples of this include allowing multiple different garbage collectors and encouraging heavy ast manipulation. I'm also afraid it is different enough to keep it from attracting a significant amount of developers from the Python community. Nonetheless, it's something I plan on using and contributing to, since it's the best option so far.
- nimmer 7y ago> allowing multiple different garbage collectors How's that a problem?
- carapace 7y agoCython? Nuitka?
- timothycrosley 7y agoI use Cython a lot! But mostly to speed up existing Python code, and build C-extensions faster. I don't see it as a strict subset of Python or a new language to build a community around. Nuitka I just started experimenting with to build standalone Python executable, and I really like the direction and roadmap they are following. In the end though both of these technologies seem like ways to somewhat speedup existing Python code and not attempts to introduce a strict language subset that would allow the greatest amount of optimization, and finally fix long running issues, like the inability to have multiple versions of a package installed.
- miki123211 7y agoThis is yet another example of the divide between wizarding and engineering[1]. When you're a small startup, what matters is the expressiveness of your language, and the ability do do a lot of things very very quickly. Type safety, performance, readability, those things don't matter. You're just a bunch of engineers who know the whole codebase inside out, you're pretty certain of what you're doing. In short, you're wizarding. If you grow big enough, this approach slows you down greatly, and you need to switch to engineering. You sacrifice some speed for making the codebase more understandable to a larger group of people, you can no longer assume everyone knows all the code, you write unit tests, need types and dislike metaprogramming because of the confusion it creates. This is why languages like Python, Ruby, Lisp or Smalltalk are amazing for small startups, but Java is what enterprises use. They're different ends of the wizarding/engineering spectrum. I wish there was a language that let you move gradually from one end to the other, exactly when you need to. [1] https://www.tedinski.com/2018/03/20/wizarding-vs-engineering.html https://www.tedinski.com/2018/03/20/wizarding-vs-engineering...
- ben_jones 7y agoWhat you call Wizarding I call "ordinary Software Development". A software developer spends ~70% of their time writing features and the rest mixed between organization/planning/roadmapping etc. A software engineer spends ~30% of their time writing features and the rest of it managing technical debt and making long-term investments towards better features and processes. Too many companies need devs but have engineers, or they need engineers but only have devs :/
- heartbreak 7y agoI’ve never known anyone to distinguish between the two roles, and they seem to be used interchangeably in the industry. (Except those people who claim software engineers aren’t real engineers)
- TylerE 7y agoWho is starting large-scale new projects in Java in 2019?
- carapace 7y ago> Instagram Server is a several-million-line Python monolith That's bananas. Nothing Instagram does requires that much code. Also, that much Python code means you're doing it wrong.
- carapace 7y agoNo, I'm seriously you guys. Python is too expressive to require mega-LoC for that site. You could implement an OS, relational DB, spreadsheet, and optimizing compiler all in less than that.
- orf 7y agoYou have no idea about their codebase, the implementation details of their features nor how they counted the lines (comments included?). So stating that it’s dumb is beyond ridiculous. You are right in that it’s certainly a high LoC count for Python, but still...
- carapace 7y agoI didn't say "dumb" I said "bananas". And yes, knowing nothing else about their code base than A) It's in Python, and B) it's several million lines of code, I feel very confident that there is at least an order of magnitude too much of it. Instagram is just not doing anything that complicated. (I should mention I specialize in maintaining and refactoring legacy Python code. I know what I'm talking about here.)
- ClippyIO 7y agoFeatures that are "not complicated" can actually very easily be "very complicated" at scale. Which Instagram does have. 500 million users, every single day.
- depressedpanda 7y agoFewer LOC would actually benefit them at scale. If you need several millions of lines of Python to do what Instagram server does, the code is bloated. My bet is that they let too many Java devs loose on the code base, without experienced Python devs reviewing the commits and managing the deluge of unnecessary classes. I've seen it happen before.
- tahdig 7y ago> ... many of whom are new to Python. well, if you ask me to write language X, I would definitely make mistakes for the first couple of weeks/months/years, that is why you need code review, mentoring and education plans for your hires. > Here’s another thing we often find developers doing at import time: fetching configuration from a network configuration source. MY_CONFIG = get_config_from_network_service() I am pretty sure this an anti-pattern, if this code passed the code review, you should make your review process more strict. def myview(request): SomeClass.id = request.GET.get("id") > Likely you’ve already spotted the problem Well, yes, why would you do this? why would this pass code review? why do we we have linters and other checks for dynamic languages > It works great for smaller teams on smaller codebases that can maintain good discipline around how to use it, and we should switch to a less dynamic language. It seems we are here blaming python for shortcomings of a monolith also, instead of chunking out specific businesses modules to separate services/micro-services. TO be honest the strict mode seems interesting, but I believe the problems they seem to be facing can be solved by a couple of changes to their pocess and code: - everyone gets a mentor if they are not experienced in python or django - code review atleast by two experienced python developers(does not count if you have coded for Java for 20 years) - teams should try to move their logic outside the monolith(it sounds like they have a monolith) - write CI tests to measure how much time it takes to import a file, if it takes more than T(line count * LINE_PROCESSING_THRESHOLD) you have to fix your code. - prepare config and load it before running the actual server, no network call for getting config All in all, python is suitable for big companies also, the thing is if don't care about the best practices, you would also have problems when you are a small startup, but in a big co it would make it impossible to move forward, trick is to independent of the company size follow best practices and have code review.
- scrollaway 7y agoThat's a long post to say "do more code review instead of investing into technical solutions to technical problems". Clearly, Instagram's solution saves them time. That means faster code reviews which incidentally makes them more accurate. Your post doesn't really make sense.
- zestyping 7y agoI like this a lot.
- jedberg 7y agoIt's interesting to me that they are going down this path instead of the microservices path. This seems like something ripe for slowly breaking down into microservices. Someone made a change that took down production because of non-deterministic outcomes? How about break out whatever they were changing into it's own service? With proper fallbacks, breaking that part shouldn't take down all of production again. To be clear, I'm not saying microservices will solve all their problems or be less work. I'm just saying that with an equal level of effort, they would probably get more overall reliability by having multiple services, they'd be able to use multiple languages, whatever is suited to the task at hand, be able to deploy even more often with less risk, and be able to isolate these types of "change on import" behavior to a much smaller surface on any given deployment.
- coldtea 7y ago>Someone made a change that took down production because of non-deterministic outcomes? How about break out whatever they were changing into it's own service? With proper fallbacks, breaking that part shouldn't take down all of production again. Yeah, now you'll have 10 interconnected services, 10x the complexity, and everything will have the ability to take down all of large parts of production, plus all the extra pain points of a distributed system...
- jedberg 7y agoYou won't have 10 times the complexity if you are taking a monolith and making each section services. You'll have to same dependency graph, it will just use the network to make calls between them instead of being local. You'll have added complexity with the network calls, which is why I said it wouldn't be any less work, just different work.
- coldtea 7y ago>You won't have 10 times the complexity if you are taking a monolith and making each section services. You'll have to same dependency graph, it will just use the network to make calls between them instead of being local. Merely "use the network to make calls between them instead of being local" will add 10 times the complexity -- you suddenly have a distributed system, latency, delays, parts that can be on or off, de-centralized configuration (which can also get out of sync), and so on.
- konschubert 7y agoI have a question about a detail in the article: > But if we moved the log_to_network call out into the outer log_calls function, [...] this would no longer compile as a strict module. My current understanding is that the log_calls method would NOT get executed during module load time!?! Why would having a side effect in this function violate the intention of __strict__ ?
- scrollaway 7y ago> My current understanding is that the log_calls method would NOT get executed during module load time!?! That's incorrect. log_calls gets executed on import because it's a decorator, so equivalent to `hello_world = log_calls(hello_world)` at the top-level (which does also get executed). log_to_network in the _wrapped() definition doesn't get executed until hello_world gets called; but outside of the definition of _wrapped does get executed.
- konschubert 7y agoRight! I missed the fact that log_calls is used as a decorator further down.
- marcoseliziario 7y agohttps://docs.python.org/3/library/importlib.html#importlib.util.LazyLoader https://docs.python.org/3/library/importlib.html#importlib.u...
- tln 7y agoAvoiding module side effects and making classes and modules immutable seem like two separate concerns
- bjoli 7y agoNot really. Mutation in general, and in modules in particular, inhibit a lot of reasoning about the code, and thus stops a whole lot of optimizations from being possible. Guile (a scheme dialect) recently got declarative modules for that reason, where a top level binding cannot change (i.e. you cannot set! a binding, but you can wrap in it a mutable container and change the contents of that container). This makes procedure calls and variable lookup a lot faster. Andy Wingo wrote about it here: https://wingolog.org/archives/2019/06/26/fibs-lies-and-benchmarks https://wingolog.org/archives/2019/06/26/fibs-lies-and-bench... . Those optimizations won't mean much for cpython, since Cpython doesn't try to run things fast, but for something like pypy this could be a big deal.
- bjoli 7y agoTo quote the article (from.memory): "adding static modules is probably the single most important optimization guile can do in the near future". The quote is probably wrong, but it is right in spirit.
- jbmsf 7y agoI like the idea, but it feels a bit heavy handed outside of a very large team. I think the first step here is to get away from the assumption that importing a module will have "interesting" side effects. This is not only a problem with Python... I tend to create mini "dependency injection" frameworks that create a pattern for loading module code at some point well after import. This patterns tends to reduce to wrapping whatever code you have in the module in a function/closure instead of just running whenever. Again, I like the idea of enforcing constraints with code, but I don't think it's a substitute for educating developers to avoid certain patterns and giving them infrastructure that makes the alternative easy.
- accidentaldev 7y agowho would have thought Instagram is a python monolith. ?
- ianamartin 7y agoTry Zope.Interface and Pyramid for a framework. You'll be really happy.
- alexchamberlain 7y agoI think this is an interesting idea, which appears to embed a stricter subset of Python within Python itself. Have the Instagram engineers tried floating this with the wider community via established channels like Python-Ideas or discuss.python.org?
- gtirloni 7y ago> That means 20-60 seconds between a developer making a change and being able to see the results of that change in their browser, or even in a unit test. This, unfortunately, is the perfect amount of time to get distracted by something shiny and forget what you were doing /me laughs in Ansible/terraform
- time4tea 7y agoWow. Talk about solving the wrong problem! Millions of lines of code in a monolith. 20s start up time. Meta monkey patching. One unit test per process... Yikes! Software architecture, anyone? Maybe Instagram should get a copy of Michael Feathers' book...
- k_sze 7y agoAnother thing that I would like to see in some kind of strict mode is the ability to mark explicit exports like in JavaScript modules. I often want to import multiple things globally at the top of a module because they are shared by multiple class or function definitions that I am writing. However, such imports end up being exposed to and usable by the consumers of my module, even though the consumers should really have imported those things at their source instead of via my module. There are currently maybe two ways to tackle this “problem”, without a strict mode: 1. Don’t import at the global module scope; but that’s a bit tedious. 2. Import with rename, like `import os as _os`, and then leave it to the principle of “we’re all consenting adults”. I.e. if anybody imports and used things that start with an underscore, it’s clearly their fault, not mine.
- andreareina 7y ago3. Import as normal, and leave it to the principle of "we're all consenting adults"; unless something is explicitly called out as being part of the public API I consider Law of Demeter[1] "violation" the same as accessing _var. [1] https://en.wikipedia.org/wiki/Law_of_Demeter https://en.wikipedia.org/wiki/Law_of_Demeter
- rurban 7y agoI like that idea, it's just not that easy. How to do define module versions and inheritance, when you are not allowed to do global assignments in the module. declarations only, and no IO or global side effect is fine, but declaring versions and inheritance need to be allowed in global scope. I added these ideas here: https://github.com/perl11/cperl/issues/406 https://github.com/perl11/cperl/issues/406