8 ms·
Why Discord is switching from Go to Rust (2020)
- midrus 4y agoOr they just got bored and wanted to try some shinier toy. I've seen this happen dozens of time, all the bullshit for justifying it is just that, bullshit. Not saying this is the case here but highly likely.
- phendrenad2 4y ago(2020) (Anyone know if they're still using Rust?)
- eatonphil 4y agoTheir blog doesn't list all articles on a single page (so you could ctrl-f) and doesn't have its own search and doesn't have it's own domain (so googleing `site:discord.com rust` returns a mix of Discord communities and blog posts). Makes it pretty hard to find stuff!
- ajot 4y agoYou can search "site:discord.com/blog/ rust", it appears to work for me on DuckDuckGo or Google. It seems TFA is the latest article mentioning Rust.
- eatonphil 4y agoWoah I didn't realize you could filter on paths. I thought `site:` was domain only.
- tedunangst 4y agoHN site listing is also useless because they host the blog on the same domain as everything else. https://news.ycombinator.com/from?site=discord.com https://news.ycombinator.com/from?site=discord.com
- faitswulff 4y agoLooks like they're still hiring for it: https://www.google.com/search?q=rust+site%3Adiscord.com%2Fjobs https://www.google.com/search?q=rust+site%3Adiscord.com%2Fjo...
- jhgg 4y agoYes we are using rust in a big way. We have multiple teams now full time working on Rust. It is being used on both the client and server, as native modules, web assembly, and also native rust services and NIFs that embed themselves in our elixir services. It has been an incredible success. I plan to blog more about it in the coming months. Our usage of Rust is continuing to grow, and if you check out our jobs page, you might notice all backend / infra jobs list Rust in them now :) I think probably 40% of requests are handled directly by rust services now, with the rest involving one or more rust service called from our Python API layer.
- SirGiggles 4y agoBit of a late reply, but how does Elixir fit into the overall strategy? Is it still like how it is described in previous engineering blogs where it acts as a kind of orchestrator for guilds?
- abeltensor 4y agoI love Rust elixir Nifs. Gives you the best of both worlds to be honest. Highly fault tolerant code with fast computation. Only downside is that it can't really handle extreme crashes like a native process can.
- wut42 4y agoErlang/Elixir + Rust is an awesome couple. For the downside you mentioned, depending on the use case, it could be interesting to use Rust as a node: https://github.com/sile/erl_dist https://github.com/sile/erl_dist
- abeltensor 4y agoVery nice. Nodes are great if you want a long running system along side your elixir app.
- alberth 4y agoIsn’t this a function of them being such heavy Erlang users and are writing NIFs (via Rustler) in Rust.
- technological 4y agoPrevious Discussion https://news.ycombinator.com/item?id=22238335 https://news.ycombinator.com/item?id=22238335
- pfraze 4y agoI’m told the Go GC has gotten better in recent years. Has anybody run a similar program in Go lately that can confirm that?
- erikbye 4y agoNot discrediting Rust, but I've noticed you rarely hear "we improved performance" by rewriting our implementation using the same language... Although this, too, can yield similar performance improvements.
- ceeplusplus 4y agoIt does look like with some GC tuning (e.g. manually triggering GC's at a smaller interval than the Go automatic GC threshold) they might've mitigated the spikes, although I don't think they would have gotten the level of perf improvement they did. Golang assembly code IME is not very optimized compared to Rust/C++. edit: reading comprehension skills are lacking, please see comment below for why I'm wrong
- jholman 4y agoI don't understand.... isn't this idea (triggering GC more often) explicitly discussed in the article?
- ceeplusplus 4y agoMy understanding is they tried to tune the GC percent to make the automatic heuristic do GC sooner, but they didn't allocate enough to have that make a difference. However, Go has a way to manually trigger a GC which they could've set on a timer on a goroutine. If they weren't actually generating that much garbage then pause times should theoretically be pretty short if you're doing a GC every 5 seconds or something like that. That being said it's not something that 100% is guaranteed to fix the issue so maybe they did test this and just didn't mention it in the blog.
- jholman 4y agoOkay, I see what you're saying about the timer. That isn't in TFA, agreed. But I still don't understand, because.... (NB: I'm not a GC expert, just a curious amateur, so my apologies if there are errors in the following, and the opportunity to be corrected in these errors is part of why I'm posting this.) Regarding the "not much garbage => theoretically times would be shorter", my understanding is that this is actually not how GC works. The GC time is a function of the size of the GC pool, because GC works by walking ("tracing") the tree of live references. So the only way to make GC faster is to have not less garbage, but less stuff allocated at all. Multi-generational GC works by dividing the whole pool into smaller pools, so that most GC passes only visit the high-churn nursery, but even then some GC passes need to read the TFA mentions this, where they say "the spikes were huge not because of a massive amount of ready-to-free memory, but because the garbage collector needed to scan the entire [thing we were keeping track of]". That is, they had virtually no garbage to collect, and that wasn't speeding up the GC. Which is consistent with how all tracing GC works, as far as I know. Comments/corrections/clarifications are requested!!
- mc4ndr3 4y agoHow illuminating. From CloudFlare posts, I had been under the impression that Go's gc was incredibly unintrusive, near-real time performance for applications operating in increments of a few hundred milliseconds. For example, CloudFlare uses Go to analyze network traffic. Yes, Rust provides a more predictable, faster memory management model than Go. At the expense of unpredictable, expensive memory leaks triggering application termination. Curious how much time and effort was dedicated to improving gc, which is a useful endeavor in its own right.
- whoisburbansky 4y agoWhere do you get that sense that Rust results in memory leaks? Is that just an assumption you’re making about languages without garbage collection, or are there examples you’re aware of Rust applications having to deal with runaway memory consumption?
- Xylakant 4y agoRust as a language does not protect against memory leaks - std::mem::forget even explicitly does so. Generally garbage collected languages do, so trading go for rust increases the risk of having memory leaks.
- bsder 4y agoYou are invoking std::mem::forget which is explicitly for circumventing destructor execution and then complaining about leaks. Okay. The documentation page for std::mem::forget goes through all the alternatives you should try before resorting to std::mem::forget. Now, perhaps std::mem::forget should be marked unsafe. However, you don't just "accidentally" run std::mem::forget. BTW, one of the problems with GC languages is the fact that you never know when your destructor might get run (ie. object gets reclaimed) so your GC thinks life is just fine but ... oops ... you just ran out of file descriptors because they are all waiting to be reclaimed.
- lillecarl 4y ago
- boxingrock 4y agoisn't the tldr on this that Go let them scale up for years before it became the bottleneck? a natural progression for any successful project...
- verdagon 4y agoI suspect that GC'd languages could mitigate this problem by introducing regions; separate areas of memory that cannot point at each other. Pony actors [0] have them, and Cone [1] and Vale [2] are trying new things with them. If golang had this, then it might not ever need to run its GC because it could just fire up a new region for every request. The request will likely end and blast away its memory before it needs to collect, or it could choose to collect only when that particular goroutine/region is blocked. Extra benefit: if there's an error in one region, we can blast it away and the rest of the program continues! [0] https://tutorial.ponylang.io/types/actors.html#concurrent https://tutorial.ponylang.io/types/actors.html#concurrent [1] https://cone.jondgoodwin.com/fast.html https://cone.jondgoodwin.com/fast.html [2] https://verdagon.dev/blog/seamless-fearless-structured-concurrency https://verdagon.dev/blog/seamless-fearless-structured-concu...
- jatone 4y agoor you know; just pace the GC mark and sweep algorithm. which is what go is doing now.
- verdagon 4y agoCorrect me if I'm wrong, but IIRC pacing would still cause a latency spike, it would just be a more strategically-timed latency spike.
- jatone 4y agosure - https://github.com/golang/go/issues/44167 https://github.com/golang/go/issues/44167. you'll see the new design the CPU util only increases in GC CPU utilizations when you're actually allocating heavily. which makes sense you're doing more work. this should completely resolve the problem discord had; since their system was in a steady state.
- abeltensor 4y agoYou are always going to have some kind of latency spike with a sweeping GC; even if that spike is tiny.
- butterisgood 4y agoWell... is this still true? Go's had a lot of perf improvements in the last two years.
- mohanmcgeek 4y agoIt wasn't true even when they wrote the blog. Realistically this should read 2018 because apparently they waited two years before writing this blogpost.
- Andys 4y agoThere was a major performance boost in Go GC just after this happened
- loudtieblahblah 4y agomeh. I'm switching from Discord to Guilded.
- sys_64738 4y agoC with coroutines?