7 ms·
STW has to die. I love that the goal is 10ms but in some environments the world has passed you by in 10ms and with STW you've timed out every single connection
by kator 12y ago
STW has to die. I love that the goal is 10ms but in some environments the world has passed you by in 10ms and with STW you've timed out every single connection.
I want badly to fall in love with Go, I've enjoyed using it for some of my projects but I've seen serious challenges with more then four cores at very low latency (talking 9ms network RTT) and high QPS.. I cheated in one application and pinned the go process to each cpu and lied to Go and told it there was only one CPU but then you loose all the cool features of channels etc.. That helped some but a Nginx/LuaJIT implementation of the same solution still crushed it on the same box, identical workload.
It would be nice if we have to have STW to have it configurable, in some environments swapping memory for latency is fine and should be configurable.
The way Zing handles GC for Java acceleration is quite brilliant, not sure how much of that is open technology, but it would be cool to see the Go team reviewing what was learned in the process of maturing Java for low-latency high qps systems.
- f2f 12y ago> lied to Go and told it there was only one CPU but then you loose all the cool features of channels etc.. channels work just fine in a single-cpu context. concurrency is not parallelism.
- dinkumthinkum 12y agoThat struck me as odd. I questioned the entire poster's results based on that statement.
- Jare 12y agoI took it to mean he launched several copies of the process, each copy pinned to one (and only one) core.
- kator 12y agoCorrect it was the only way to get Go to scale to a reasonable response rate for this application and it thus meant my CPU loading was not even since it was based on the randomness of socket connections that persisted on one cpu or another.
- coldtea 12y agoHe's not talking about merely being able to use channels in one core setup. This misunderstanding made me question your whole comment.
- scott_s 12y agoIt sounds like he wanted parallelism for performance.
- georgemcbay 12y agochannels work fine in a single-cpu context, but in the setup OP described you won't be able to use them to have the various processes talk to each other. At least not without using one of the various netchan-alikes or rolling your own channel<->ipc solution, neither of which is likely to result in an elegant solution (note how the original netchan was abandoned by the Go team).
- kyrra 12y agoI believe what your parent was looking for is being able to pin a specific goroutine to a single CPU. If you are doing network traffic analysis, if a goroutine will jump between CPUs, you get the cache-miss overhead. This is the reason snort[0] (all in C) doesn't even use threads and sticks to a single process. To be able to use multiple threads when network processing you need hardware support (or something like PF_RING[1]) to distribute your load between cores. You want to be able t keep a data stream (TCP stream) pinned to a given CPU core, and doing that in Go is near impossible with how go-routines are scheduled. [0] https://www.snort.org/ https://www.snort.org/ [1] http://www.ntop.org/products/pf_ring/ http://www.ntop.org/products/pf_ring/
- shurcooL 12y ago> You want to be able t keep a data stream (TCP stream) pinned to a given CPU core, and doing that in Go is near impossible with how go-routines are scheduled. Doesn't http://godoc.org/runtime#LockOSThread http://godoc.org/runtime#LockOSThread let you do that?
- rayiner 12y ago> STW has to die. Depending on your usage profile, there is very little more efficient from a throughput standpoint than STW collection, except maybe manual memory management with clever use of pool allocation. Even malloc()/free() will be slower in situations where you're creating a lot of short-lived garbage.
- Locke1689 12y agoConcurrent Gen1 seems fine. You still have to worry about STW Gen2, but it's much better. FWIW, no one serious about allocations in native code uses naive malloc()/free(). My favorite trick is in game programming where you have a current frame pool that you just reset every frame.
- jjoonathan 12y agoThe thing I like about pools is that they seem to actually reduce the cognitive load required to consistently get memory management "right" in a large application. With manual memory management you have to worry about the ownership convention of each chunk of code, with GC you have to worry about architecting strong/weak references so as not to inadvertently retain everything (not to mention latency issues), and with reference counting you have to worry about ownership cycles. In practice I haven't come up with a better strategy than enforcing some sort of top-down hierarchy which effectively smashes most of the differences in cognitive load between manual/refcounting/GC. GC has slightly less upfront busywork but in practice the tooling tends to be poor so it's a wash (if that). In GC and refcounting it's easy for one inexperienced/tired/sloppy individual to create a massive leak completely out of proportion to the footprint of their immediate code. Pools, in contrast, allow the same top-down approach with the very significant benefits that I don't have to think about the hierarchy at a finer level than the pool itself (which I have to do for the 3 other approaches) and that memory management mistakes don't typically lead to the globally-connected-component of the dependency graph sticking around indefinitely.
- pcwalton 12y agoOne big problem with pools is that you have to deallocate the pool sometime (or else your program's working set will continually grow over time), and when you do, all the pointers into that pool become dangling. Another problem with pools is that you can't deallocate individual objects inside a pool. This is bad for long-lived, highly mutable data stores (think constantly mutating DOMs, in which objects appear and disappear all the time, or what have you).
- nulltype 12y agoHere's the paper on the Zing JVM's C4 GC: http://www.azulsystems.com/sites/default/files/images/c4_paper_acm.pdf http://www.azulsystems.com/sites/default/files/images/c4_pap... I'm not sure how much of it applies to Go, but I'm guessing it's quite a bit more complicated than the current collector. It also may only work on 64 bit X86 hardware (which is fine by me, but looks like a subset of what go currently supports). There's a simpler overview in http://www.azulsystems.com/sites/default/files//images/wp_pgc_zing_v5.pdf http://www.azulsystems.com/sites/default/files//images/wp_pg...
- hedgehog 12y agoOne of the C4 authors posted some advice on golang-dev a while back: https://groups.google.com/d/msg/golang-dev/GvA0DaCI2BU/1EpYa8HbxdIJ https://groups.google.com/d/msg/golang-dev/GvA0DaCI2BU/1EpYa... The good news is Go has been chipping away at that stuff.
- kator 12y agoThat was a great thread thanks for linking it! If Go wants to be a serious solution in many applications pauseless GC needs to happen and until it does it will be just like all the other solutions that can't play at scale and low latency. If it stays in that zone it'll have to compete with a very large set of perfectly fine solutions with massive libraries and lots of people who know how to code in those languages.
- tomp 12y agoThe problem with C4 is that it requires a kernel module, so it's obviously not suitable as the primary GC.
- fmstephe 12y agoThat's a good point and often passed over when talking about C4. It's a real shame, because that is a fairly large sticking point. There was an attempt to get the kernel changes needed by C4 pushed into the standard kernel, but I remember it got a lot of push-back by the kernel devs.
- FreezerburnV 12y agoNot sure if this makes it any better, but a bit later on they talk more about how the STW scheme will work. They call each thread which is not the concurrent GC a mutator, and part of the GC scheme will be ensuring that a mutator thread be stopped for less than 1ms, and that the 1ms pause should not stop any other mutator threads. If they can manage it, this scheme sounds promising for most soft real time programs.
- 6cxs2hd6 12y agoThat's not how I understood it. It says they're doing a hybrid STW and concurrent GC. (a) A mutator may be stopped for up to 10 ms of every 50 ms by the STW collector. (b) If more collection is needed, it will transition to a concurrent collector for the remaining 40 ms. Concurrent may pause stop a mutator for up to an additional 1 ms. Therefore the worst case is 10 + 1 = 11 ms, not 1 ms.
- bradleyjg 12y agoI read them to mean that the one GC thread would stop the world for no more than 10 out every 50 ms, and in addition the concurrent garbage collector could ask any mutator thread to pause and do no more than 1 ms worth of GC work during any 40 ms non-STW period. So in terms of real time, your deadlines could be no less than 11 ms.
- vbit 12y ago> That helped some but a Nginx/LuaJIT implementation of the same solution still crushed it on the same box, identical workload. That's very interesting because LuaJIT has GC as well. Could you please reveal a bit more about the type of application, number of lines, and if the code is public?
- yoklov 12y agoNot sure if it's related, but LuaJIT can remove allocations all-together in a lot of cases. This[0] talks about how it works some. [0]: http://wiki.luajit.org/Allocation-Sinking-Optimization http://wiki.luajit.org/Allocation-Sinking-Optimization
- lucian1900 12y agoMost JITs do that. Static compilation can do it to some degree as well, but in most mutation-happy languages (Go included) it cannot prove the allocation does not escape frequently enough to matter.
- yoklov 12y agoNot really. Most do escape analysis/SROA. Allocation sinking is much more general. It works even if the allocation escapes (the allocation is sunk to the point where it escapes). Additionally, all allocations can be sunk, not just stuff like the point class used in the example on that page. Growing dynamic buffers, string concatenation, etc, are all sinkable. My understanding is that it's fairly tricky to implement (requires special consideration in runtime design) and so is not very widely used. Last I checked, the Dart compiler is the only other place I've heard of it being used. FWIW its not really a thing Go would need anyway, given that you actually have control over your allocations.
- kator 12y agoSadly it's not public but yes I was very happy to see how well LuaJIT with Nginx peformed for this problem set. I developed several solutions in C, C++, Java, Java+Zing, LuaJIT and Go. The Nginx/LuaJIT (openresty to be exact) solution had a great ratio of "approachable code" to "performance" factor. I was able to get it built and transfer its use to several other teams without them having to know C and still be able to make changes effectively.
- p0nce 12y agoCouldn't you avoid doing allocations in Go?
- scythe 12y ago>I cheated in one application and pinned the go process to each cpu and lied to Go and told it there was only one CPU but then you loose all the cool features of channels etc.. That helped some but a Nginx/LuaJIT implementation of the same solution still crushed it on the same box, identical workload. If your primary motivation for using Go is lightweight concurrency, it has been developed for Lua, though in fairness not as extensively as goroutines: http://github.com/askyrme/luaproc http://github.com/askyrme/luaproc http://www.jucs.org/jucs_14_21/exploring_lua_for_concurrent/jucs_14_21_3556_3572_skyrme.pdf http://www.jucs.org/jucs_14_21/exploring_lua_for_concurrent/...
- fleitz 12y agoYeah if you want to hit 60 fps, 2/3rds of your processing time just evaporated.
- waps 12y ago60fps won't matter one bit if it's 120fps for one second followed by 0fps for another second. That's why people are asking for pauseless.
- fleitz 12y agoFully agree. I thought the GC only paused for 10 ms, which would be bad enough, if it's a full second... that's insane.
- jblow 12y agoAnd it's worse, because... "only 60fps?" 60fps was fine pre-VR, but now you want to do 90fps times 2 eyes. That's a lot of rendering.
- reality_czech 12y agoIf only someone would invent a co-processor that we could offload our rendering to. A "graphics processing unit," so to speak. Oh well. One day! If you want to update some physics model at 60 FPS, you will be updating every 16 ms. An occassional 10 ms pause is not going to prevent that. Anyway, any reasonable game design needs to deal with pauses that are caused by network congestion intelligently, which can often be longer than 16 ms anyway.
- cma 12y agoA pause due to network congestion isn't nearly on the same level as a pause in updating your head orientation in VR.
- jblow 12y agoDude you have NO idea what you are talking about. An occasional 10ms pause absolutely will "prevent that". Those of us who work on software that does a lot of rendering using these fabled GPUs you mention know that feeding the GPU properly is a big problem and requires a lot of code to be running on the, err, CPU. I don't even know what you are talking about wrt network congestion. What are you talking about??
- dschiptsov 12y agoIt seems that concurrent (non-STW) GC for non-functional languages (randomly accessed mutable data) is a myth. Yes, Java claims to have an efficient GC (on paper). In reality, everyone know how it works (it bloats, and GC pauses are increasing dramatically). On the contrary, Haskell has less sophisticated GC but much smaller and more predictable GC timings. btw, some really smart people are suggesting that it is the code organization, not GC algorithm is a crucial factor. http://home.pipeline.com/~hbaker1/LinearLisp.html http://home.pipeline.com/~hbaker1/LinearLisp.html