4 ms·
I think we learned a long time ago that though codegen is one path to speedy code, it is not the only path. The thriftpy module dynamically generates a Python
by pixelmonkey 11y ago
I think we learned a long time ago that though codegen is one path to speedy code, it is not the only path.
The thriftpy module dynamically generates a Python module that is based on parsing the thrift schema. That happens once: at module import time. After that, Python has cached the module in sys.modules and it doesn't need to be evaluated again during the program's runtime.
The module contains efficient Python classes that are Python bytecode just the same way "compiled" classes would be. But, thriftpy has Cython implementation of Thrift protocols and transports that compile down to C, not Python bytecode. Thus, the thriftpy impl's are not only more dynamic and less cumbersome than the codegen alternative, but they are also faster.
In the Python community, we prefer to take other approaches to achieve execution speed.
Dynamically generating a Python module is not "metaprogramming". It is not black magic. You can generate a module yourself simply by instantiating a Python module type. Hooking the import statement is an officially supported part of the language, and done by many libraries.
This may all seem very worrying to a C++ or Java programmer, but in the Python community we have been doing dynamism with execution speed for about 10 years now, and we haven't regretted any of the results!
- haberman 11y agoFor what it's worth, I prefer the approach you describe. I have advocated for it for a long time. I've been writing C extensions to accelerate parsing in dynamic languages for 10 years, and I'm very familiar with this sort of dynamism. It's interesting though that you describe this approach as more Pythonic, because one of the most outspoken critics of this approach is a hard-core Python guy that I work with who has lots of experience with the Python ecosystem. He feels very strongly that it is more Pythonic to have very flat/concrete generated code that is transparent to the reader, and really does not like the idea of hooking import and generating everything at runtime. This is exactly what I am describing about how it is hard to please everybody. Different people appear to have fiercely different opinions about what is idiomatic. When we wrote the Ruby protobuf implementation this year, we took an approach much more along the lines of what I personally prefer. The extension is mostly implemented in C, and it's very easy to build types at runtime. It doesn't directly import .proto files (which I would have preferred) because there was still some desire from others to have some kind of code generation. So the approach we took was to use a Ruby-like DSL for describing the protobuf schema. But really this is nothing but a translation of the .proto file into the Ruby DSL. ie. take the .proto file: syntax = "proto3"; message Test { int32 foo = 1; double bar = 2; Test test = 3; } The "generated code" for this is simply: require 'google/protobuf' Google::Protobuf::DescriptorPool.generated_pool.build do add_message "Test" do optional :foo, :int32, 1 optional :bar, :double, 2 optional :test, :message, 3, "Test" end end Test = Google::Protobuf::DescriptorPool.generated_pool.lookup("Test").msgclass