5 ms·
The dictionary implementation changes are the biggest speed up I can think of. They touch nearly everything in the language. One of the interesting side-effect
by rectangletangle 8y ago
The dictionary implementation changes are the biggest speed up I can think of. They touch nearly everything in the language.
One of the interesting side-effects of this, is in 3.7+ dictionaries are now ordered by insertion, which I'm sure will help fix tons and tons of subtle non-deterministic bugs (and probably expose a few too).
Although I'm not 100% on the performance characteristics of the new unicode implementation. The unicode first string handling is a massive improvement when it comes to non-English languages. So IMO it's worth a hefty performance hit, based solely on it's own merits. In practice, it hasn't seemed to make a meaningful difference to me. So I'd wager it's good enough, unless you're hitting some sort of use case specific bind. In which case, it's probably time to leverage the C bindings anyway.
- danso 8y agoIs there a good, in-depth technical write up of the changes made to dict implementation?
- dralley 8y agoLong video, but it's Raymond Hettinger, so it's worth it. https://www.youtube.com/watch?v=p33CVV29OG8 https://www.youtube.com/watch?v=p33CVV29OG8
- philipov 8y agoNot a write-up, but how about a pycon talk from one of the senior core devs? https://www.youtube.com/watch?v=npw4s1QTmPg https://www.youtube.com/watch?v=npw4s1QTmPg
- danso 8y agoThanks, I remember Hettinger's video making the rounds last year but never got around to watching it. Guess it's about time to make time for it. edit: Ah, I remember. His talk uses slides that were created in sphinx/readthedocs, but IIRC, the link to those docs never worked for me: https://twitter.com/raymondh/status/867021035719110656 https://twitter.com/raymondh/status/867021035719110656
- rectangletangle 8y agohttps://mail.python.org/pipermail/python-dev/2012-December/123028.html https://mail.python.org/pipermail/python-dev/2012-December/1... and https://www.python.org/dev/peps/pep-0468/ https://www.python.org/dev/peps/pep-0468/ The implementation is based off of PyPy's "compact" dictionary implementation.
- mark-r 8y agoJust to be clear, I wasn't talking about the change making Unicode strings universal in 3.0. I was talking about the change to the internal string representation in 3.3, described by PEP 393. The change to dictionaries seems a lot more likely to make a difference, although it was introduced much later.
- blattimwind 8y ago> is in 3.7+ dictionaries are now ordered by insertion, 3.6 > The unicode first string handling is a massive improvement when it comes to non-English languages. So IMO it's worth a hefty performance hit, based solely on it's own merits. In practice, it hasn't seemed to make a meaningful difference to me. So I'd wager it's good enough, unless you're hitting some sort of use case specific bind. In which case, it's probably time to leverage the C bindings anyway. That's because everyone had to use Unicode strings on Python 2 already, so the en/decoding and Unicode handling overhead was already there. Applications not using unicode strings were mostly just buggy or didn't work outside ASCII+. The major difference is that Python 3 string code is a lot less fragile than Python 2 code, because it either works and does so for non-English text, too, or it doesn't. Meanwhile Python 2 code would appear to work fine until you started to drop those sweet umlauts and got exceptions all over the place.
- hultner 8y agoWell it’s like that by accident in 3.6 because cpython36 does it this way and no other implementation exists. In 3.7 it’s officially a part of the spec thus required by all Python 3.7 compatible implementations and not just cpython.
- deleted 8y ago[deleted]
- UncleEntity 8y ago>> is in 3.7+ dictionaries are now ordered by insertion, >3.6 It was an implementation detail in 3.6, 3.7 made it part of the language spec. Which is good, can finally get rid of the metaclass I hacked together for my spark (earley parser that used to be used in the python build process to parse ASDL) shenanigans.
- albertzeyer 8y agoDo you have a link for the rationale behind this? So it is not a hashmap anymore? Or in addition to the hashmap, it stores the insertion index? I wonder if it is a good idea to have this part of the language spec, as they might want to use a faster (non-deterministic) implementation later at some point, which would then break again lots of code.