5 ms·
Sorry, I must be missing something in this blog post because the requirements here sound incredibly minimal. You just needed an HTTP service (sitting behind an
by meritt 6y ago
Sorry, I must be missing something in this blog post because the requirements here sound incredibly minimal. You just needed an HTTP service (sitting behind an Envoy proxy) to process a mere 500 requests/second (up to 1MB payload) and pipe them to Kinesis? How much data preparation is happening in Rust? It sounds like all the permission/rate-limiting/etc happens between Envoy/Redis before it ever reaches Rust?
I know this comes across as snarky but it really worries me that contemporary engineers think this is a feat worthy of a blog post. For example, take this book from 2003 [1] talking about Apache + mod_perl. Page 325 [2] shows a benchmark: "As you can see, the server was able to respond on average to 856 requests per second... and 10 milliseconds to process each request".
And just to show this isn't a NodeJS vs Rust thing, check out these webframework benchmarks using various JS frameworks [3]. The worst performer on there still does >500 rps while the best does 500,000.
It's 2020, the bar needs to be much higher.
[1] https://www.amazon.com/Practical-mod_perl-Stas-Bekman/dp/0596002270 https://www.amazon.com/Practical-mod_perl-Stas-Bekman/dp/059...
[2] https://books.google.com/books?id=i3Ww_7a2Ff4C&pg=PT356&lpg=PT356 https://books.google.com/books?id=i3Ww_7a2Ff4C&pg=PT356&lpg=...
[3] https://www.techempower.com/benchmarks/#section=data-r19&hw=ph&test=db&l=zik0sf-1r https://www.techempower.com/benchmarks/#section=data-r19&hw=...
- BubRoss 6y agoI hope some day doing something trivial using rust will no longer warrant a hacker news post to promote a startup.
- lostcolony 6y agoThey list out what is being done by the service - "It would receive the logs, communicate with an elixir service to check customer access rights, check rate limits using Redis, and then send the log to CloudWatch. There, it would trigger an event to tell our processing worker to take over." That sounds like a decent amount of work for a service, and without more detail it's very hard to say whether or not a given level is efficient or inefficient (we don't know exactly what was being done; we can assume that they're using pretty small Fargate instances though since the Node one came in at 1.5G). They also give some number; 4k RPM was their scaleout point for Node (that's not necessarily the maximum, but the point they felt load was sufficiently high to warrant a scaleout; certainly, their graph shows an average latency > 1 second). Rewriting in Rust, that number was raised to 30k RPM; 100 mb of memory, < 40ms average latency (and way better max), and 2.5% of CPU. Given all that, it sounds like, yes, GC was the issue (both high memory and CPU pressure), and with the Rust implementation (no GC) they're nowhere near any CPU or memory limit, and so the 30k is likely a network bottleneck. That said, while I agree that sounds like a terrible metric on the face of it, with what data they've provided (and without anything else), it also sounds like it may be due to they're just operationally dealing with very large amounts of traffic. They may want to consider optimizing the network pipe; not familiar enough with Fargate, but if it's like EC2, there may be a sizing of cpu/memory that also gives you a better network connection (EC2 goes from 1 GBPS to a 10 GBPS network card at one instance type)
- theikkila 6y agowasn't that 4k requests per minute?
- wahern 6y ago> That sounds like a decent amount of work for a service 5+ years ago I wrote a real-time transcoding, muxing streaming radio service that did 5000 simultaneous connections with inline, per-client ad spot injection (every 30 seconds in my benchmark). Using C and Lua. On 2 Xeon E3 cores--1 core for all the stream transcoding, muxing, and HTTP/RTSP setup, 1 core for the Lua controller (which was mostly idle). The ceiling was handling all the NIC IRQs. While I think what I did was cool, I know people can eke much more performance out of their hardware than I can. And I wasn't even trying too hard--my emphasis is always on writing clear code and simple abstractions (though that often translates into cache-friendly code). At my day job, in the past two months I've seen two services in a "scalable" k8s clusters fall over because the daemons were running with file descriptor ulimits of 1024. "Highly concurrent" Go-based daemons. For all the emphasis on scale, apparently none of the engineers had yet hit the teeny, tiny 1024 descriptor limit. We really do need to raise our expectations a little. I haven't written any Rust but I have recently helped someone writing a concurrent Rust-based reverse proxy service debug their Rust code and from my vantage point I have some serious criticisms of Tokio. Some of the decisions are clearly premature optimization chosen by people who probably haven't actually developed and pushed into production a process that handles 10s of thousands of concurrent connections, single-threaded or multi-threaded. At least not without a team of people debugging things and pushing it along. For example, their choice of defaulting to edge-triggered instead of level-triggered notification shows a failure to appreciate the difficulties of managing backpressure, or debugging lost edge-triggered readiness state. These are hard lessons to learn, but people don't often learn them because in practice it's cheaper and easier to scale up with EC2 than it is to actually write a solid piece of software.
- lostcolony 6y agoAll I'm saying is that without some example of the payloads they're managing, and the logic they're performing, it's hard to say "this is inefficient". And, as I mentioned, if their CPU and memory are both very low, it's likely they're hitting a network (or, yes, OS) limit. I've seen places hit ulimit limits...I've also seen places hit port assignment issues, where they're calling out to a downstream that can handle thousands of requests with a single instance, so there are two, and there aren't enough port identifiers to support that (and the engineers are relying on code that isn't reusing connections properly). Those are all things worth learning to do right, agreed, and generally doing right. I'm just reluctant to call out someone for doing something wrong unless I -know- they're doing something wrong. The numbers don't tell the whole story.
- jeffbee 6y agoSeriously, 500 qps was something we used to do in interpreted languages on the Pentium Pro. But this kind of blog post is a whole genre: How [ridiculous startup name] serves [trivial traffic] using only [obscenely wasteful infrastructure] in [trendy runtime framework that's a tiny niche or totally unknown in real industry].
- jen20 6y ago> totally unknown in real industry Microsoft, Apple, Amazon, Oxide, Mozilla, Dropbox and CloudFlare would like a word...
- jeffbee 6y agoEngineering by press release? A fun fact for you: Rust is forbidden at Dropbox for new development.
- blub 6y agoForbidden for new development? Ouch. That sounds pretty serious, do you have some more info?
- staticassertion 6y agoI worked there, it's totally false.
- jeffbee 6y ago_When_ did you work there?
- staticassertion 6y agoMy last day was sometime in February, if I recall correctly. I was fulltime until September, contracting part time after. I'm aware of at least two ongoing projects they're writing in Rust, which I obviously won't comment on in detail publicly.
- hu3 6y agoIt took your comment to make me notice it wasn't 30k requests/second but minute instead. 500 requests per second is what I would expect of a default PHP + Apache installation on a small Ubuntu server. I too have a hard time grasping whats special here. For example I saw cached Wordpress setups handle 400 to 500 requests per second. And Wordpress isn't known for performance even with caching plugins.
- deleted 6y ago[deleted]
- varjag 6y agoTo put it further into perspective, C10K challenge was circa 1999: http://www.kegel.com/c10k.html http://www.kegel.com/c10k.html
- deleted 6y ago[deleted]
- deleted 6y ago[deleted]