7 ms·
PyPy for low-latency systems
- raymondh 8y agoThis is a nice bit a progress and addresses a major concern about using PyPy in real-time systems.
- alex-wallish 8y agoThis is awesome!
- dajonker 8y agoSeems quite useful, allowing the developer to guide the garbage collector into the right direction by carefully placing statements to tell the garbage collector when it is (not) ok to run. But make sure to add some comments describing the purpose of those statements, and how to profile your code to check whether it's still working correctly. Don't want to accidentally stop garbage collection altogether, either.
- zokier 8y agoThis could possibly be combined with multiprocessing for great effect? I'm imaging something like having a pool of workers executing tasks (/reacting to events/serving requests/etc), and only running gc after the task has been done, but before indicating readyness to the supervising process (/loadbalancer/etc).
- aidenn0 8y agoEven a very simple mode with two processes that never have GC enabled at the same time would greatly improve things.
- htfy96 8y agoNice examples and graphs. What really confuses me is the definition of "low-latency" nowadays. The meaning suffers a slippery slope in recent years. It used to refer to a scale of microseconds at HFT shops, then it came to web request latency at a scale of milliseconds. Now every GC-based language claims they are "low-latency" because 90%/95%/99% GC stop is within 16.67ms/50ms/whatever. Therefore today some HFT developers invented words like "ultra-low latency"[0] to name their work. [0]: https://en.wikipedia.org/wiki/Ultra-low_latency_direct_market_access https://en.wikipedia.org/wiki/Ultra-low_latency_direct_marke...
- aidenn0 8y agoPrior to HFT shops, embedded use cases might consider 100s of us to be the threshold for low-latency (40 instructions on a typical 8-bit micro running in single-digit MHz is under 100us, with no cache-misses to muck up timings).
- jacobolus 8y ago> What really confuses me is the definition of "low-latency" nowadays There has never been a strict definition of “low latency”. It is heavily context-dependent. Something which is not “low latency” enough for live audio or a car’s steering controller might be plenty fast for an internet text chat service.
- gpderetta 8y agoA better term is soft or hard real-time.
- baragiola 8y agoThat's been the case for operating systems and it's quite effective
- zokier 8y agoWhile not completely orthogonal, I feel like real-time (soft or hard) is distinct from low latency; real-time can be still relatively high-latency as long as the latency is bounded/controlled/predictable. Of course generally in practice real-time systems strive to have low latency too.
- jacobush 8y agoAbsolutely. Think about turning a super tanker or something other huge. It must happen on time. It doesn't matter if it takes one or two seconds to initiate.
- true_religion 8y agoI thought the difference between soft vs hard realtime wasn't one of magnitude, but context. Soft: After X time, usability of the data degrades with age. Hard: After X time, the data is useless (example: realtime car control systems; if you take too long the car has already crashed)
- pstrateman 8y agoBut a new event could be submitted at anytime. This is certainly an improvement, but not a complete solution.
- viraptor 8y agoIdeally you have enough copies of the server process to handle the events that come when another process is running GC. You already have to have enough of them to handle events while other workers are busy.
- cakoose 8y agoRelevant: "Blade: A Data Center Garbage Collector" (2015, https://arxiv.org/pdf/1504.02578.pdf https://arxiv.org/pdf/1504.02578.pdf) Terrible title, but basically the same idea.
- kahseng 8y agoReminds me of a time at Quora in 2011 where we saw Python GC impact 99th percentile server-side site speed. So drawing from HFT inspiration where some companies would disable JVM GC during trading hours and perform them offline, I thought about how to take some backends periodically offline in order to have GC not happen on user requests. A simpler operational solution emerged though where I just had to disable GC on user requests and make it happen only on a special "/_gc" endpoint. I then dual purposed the frequent nginx/haproxy backend health-check functionality to use that endpoint, thereby ensuring all backends had frequent GC and the time spent there only impacting the health check requests, and not that of end users. edit: added more details I remembered later
- JanisL 8y agoThis is a very interesting approach, what happened with memory footprint when you did this?
- kahseng 8y agoThanks, don't think I saw much impact at all in aggregate - our memory consumption on these web servers were dominated by objects we intentionally stored per request or globally, and not temporary/unreferenced python objects.
- JanisL 8y agoNice! I'll have to give this a try at some point if I run into GC related latency issues and see if it works on my systems.
- cma 8y agoEven while GC is delayed Python (CPython at least) will free some stuff through reference counting. Only circularly referenced stuff should stick around until the next GC run. So that can avoid lots of stack temporaries and stuff.
- Doxin 8y ago
- genjipress 8y agoI'd be curious to see if any of the work done here could be applied back to the main CPython project. I doubt it could happen immediately -- at least, not with the way GC is currently implemented -- but PyPy has been a source of innovation for CPython in the past (see: new dict implementation).
- throwaway12iii 8y agoCPython already has a lower latency GC than PyPy, gc.disable() already works, and allowed manual memory management when needed. Reference counting allows you (if needed) to keep references to memory in your python code, and free them in the right spots. This is PyPy becoming useful for a lot more production use cases. From web APIs that have a latency SLA, to audio, games. In many cases peak performance is not important, it's the minimum performance.
- mattip 8y agoRefcounting comes with its own in-thread gc pauses whenever you exit a block or context and the local variables are collected.
- throwaway12iii 8y agoYeah. However you have the option to not pause if it is important. You can control where the memory management happens. You can either keep references to the memory, and call gc.disable(). When you are ready you can let go the references and enable the gc. PyPy now lets you control where memory management happens. Making it possible to control worst case performance. For many production apps this is a big deal.
- mattip 8y agoYou can never prevent the GC cycle in CPython at the end of a block (context). You can only prevent the GC that tries to break reference cycles. If your class does crazy things at destruction, like "time.sleep(10)", and you create an instance of the class inside a function, when that function returns you will pause CPython even if you call gc.disable() You also cannot disable the minor collections in PyPy, only the major collections, but once the JIT kicks in PyPy can prevent some of the object churn by optimizing instances away.
- a_imho 8y agoCombining these two functions, it is possible to take control of the GC to make sure it runs only when it is acceptable to do so. I'm conflicted on this. My gut tells me if I'm going to manually take control over the garbage collector I should reconsider my design decisions. Disclaimer, I don't know the first thing about PyPy or Gambit Research, so probably this is the right approach for them?
- Fragoel2 8y agoIt's explained in the article that this is to solve one specific issue that Gambit Research had: in some parts of their code they need to take action with very low latencies (<10ms) and hence they can't wait for the GC. This way they manually execute the GC in other sections of the code were timing requirements are relaxed.
- Doxin 8y agoIt's a fairly common thing to do for e.g. games as well. Disable the GC during all your code, and then run it manually at the end of your frame. All you're doing is moving the gc runs to predictable points. You could even skip a collection if you've got a slow to render frame or two, evening out the spikes in framerate.
- firethief 8y agoThe advantages of a typical GC are that it avoids the development costs of manual memory management, and it allows high throughput. The main disadvantage is usually latency spikes. Using this feature decreases maximum latency spikes by orders of magnitude, with only a small cost in cognitive burden and throughput. If without this feature GC might have been a good tradeoff if not for latency, there is only a sliver of design space in which using this feature would push the throughput out of acceptable bounds. (And cognitive burden is still way lower than any other form of memory management.)