10 ms·
Serving a half billion requests per day with Rust and CGI
- deleted 1y ago[deleted]
- andrewstuart 1y agoHow meaningful is “per day” as a performance metric?
- diath 1y agoNot at all, it may be a useful marketing metric, but not a performance one. The average load does not matter when your backend can't handle the peaks.
- xnx 1y agoTrue, though a lot higher spec'ed systems couldn't handle the minimum 5000 requests/second this implies.
- dspillett 1y agoAs a comparison between implementations it can be useful. It is more than a big enough number that, if the test was actually done over a day, temporary oddities are dwarfed. If the test was done over an hour and multiplied then it is meaningless: just quote the per hour figure. Same, but more so, if the tests were much shorter than an hour.
- hu3 1y agoI work on a system for a client that averages 50 requests per second but handles 6k req/s during peaks and we have SLA of P99% <= 50ms. So I'd say per day is not very meaningful.
- kragen 1y agoIt was traditional 30 years ago to describe web site traffic levels in terms of hits per day, perhaps because "two hundred thousand hits per day" sounds more impressive than "2.3 hits per second". Consequently a lot of us have some kind of intuition for what kind of service might need to handle a thousand hits per day, a million hits per day, or a billion hits per day. As other commenters have pointed out, peak traffic is actually more important.
- shrubble 1y agoIn a corporate environment, for internal use, I often see egregiously specced VMs or machines for sites that have very low requests per second. There's a commercial monitoring app that runs on K8s, 3 VMs of 128GB RAM each, to monitor 600 systems; using 500MB per system, basically, just to poll it each 5 minutes, do some pretty graphs, etc. Of course it has a complex app server integrated into the web server and so forth.
- RedShift1 1y agoYep. ERP vendors are the worst offenders. Last deployment for 40-ish users "needed" an 22 CPU cores and 44 GB of RAM. After long back and forths I negotiated down to 8 CPU cores and 32 GB. Looking at the usage statistics, it's 10% MAX... And it's cloud infra so paying a lot for RAM and CPU sitting unused.
- ted537 1y agoHaha yes -- like what do you mean this CRUD app needs 20 GB of RAM and half an hour to startup?
- simonw 1y agoI really like the code that accompanies this as an example of how to build the same SQLite powered guestbook across Bash, Python, Perl, Rust, Go, JavaScript and C: https://github.com/Jacob2161/cgi-bin https://github.com/Jacob2161/cgi-bin
- Bluestein 1y agoThis is a veritable Rosetta stone of a repo. Wow.-
- masklinn 1y agoChecked the rust version, it has a toctou error right at the start, which likely would not happen in a non-cgi system because you’d do your db setup on load and only then would accept requests. I assume the others are similar. This neatly demonstrates one of the issues with CGI: they add synchronisation issues while removing synchronisation tooling.
- simonw 1y agoHad to look that up: Time-Of-Check-to-Time-Of-Use Here's that code: let new = !Path::new(DB_PATH).exists(); let conn = Connection::open(DB_PATH).expect("open db"); // ... if new { conn.execute_batch( r#" CREATE TABLE guestbook( So the bug here would occur only the very first time the script is executed, IF two processes run it at the same time such that one of them creates the file while the other one assumes the file did not exist yet and then tries to create the tables. That's pretty unlikely. In this case the losing script would return a 500 error to that single user when the CREATE TABLE fails. Honestly if this was my code I wouldn't even bother fixing that. (If I did fix it I'd switch to "CREATE TABLE IF NOT EXISTS...") ... but yeah, it's a good illustration of the point you're making about CGI introducing synchronization errors that wouldn't exist in app servers.
- kragen 1y agoThat sounds correct to me, but I think I would apply your suggested fix.
- jchw 1y agoHonestly, I'm just trying to understand why people want to return to CGI. It's cool that you can fork+exec 5000 times per second, but if you don't have to, isn't that significantly better? Plus, with FastCGI, it's trivial to have separate privileges for the application server and the webserver. The CGI model may still work fine, but it is an outdated execution model that we left behind for more than one reason, not just security or performance. I can absolutely see the appeal in a world where a lot of people are using cPanel shared hosting and stuff like that, but in the modern era when many are using unmanaged Linux VPSes you may as well just set up another service for your application server. Plus, honestly, even if you are relatively careful and configure everything perfectly correct, having the web server execute stuff in a specific folder inside the document root just seems like a recipe for problems.
- p2detar 1y agoFor smaller things, and I mean single-script stuff, I pretty much always use php-fpm. It’s fast, it scales, it’s low effort to run on a VPS. Shipped a side-project with a couple of PHP scripts a couple of years ago. It works to this day.
- jchw 1y agophp-fpm does work surprisingly well. Though, on the other hand, traditional PHP using php-fpm kinda does follow the CGI model of executing stuff in the document root.
- g-mork 1y agoprocessless is the new serverless, it lets you fit infinite jobs in RAM thus enabling impressive economies of scale. only dinosaurs run their own processes
- taeric 1y agoI thought the general view was that leaving the CGI model was not necessarily better for most people? In particular, I know I was at a bigger company that tried and failed many times to replace essentially a CGI model with a JVM based solution. Most of the benefits that they were supposed to see from not having the outdated execution model, as you call it, typically turned into liabilities and actually kept them from hitting the performance they claimed they would get to. And, sadly, there is no getting around the "configure everything perfectly" problem. :(
- deleted 1y ago[deleted]
- rokob 1y agoI’m interested why Rust and C have similarly bad tail latencies but Go doesn’t.
- twh270 1y agoOP posited SQLite database contention. I don't know enough about this space to agree or disagree. It would be interesting, and perhaps illuminating, to perform a similar experiment with Postgres.
- bracketfocus 1y agoThe author guessed it was a result of database contention. I’d also be interested in getting a concrete reason though.
- scraptor 1y agosqlite resolves lock contention between processes with exponential backoff. When the WAL reaches 4MB it stops all writes while it gets compacted into the database. Once the compaction is over all the waiting processes probably have retry intervals in the hundred millisecond range, and as they exit they are immediately replaced with new processes with shorter initial retry intervals. I don't know enough queuing theory to state this nicely or prove it, but I imagine the tail latency for the existing processes goes up quickly as the throughput of new processes approaches the limit of the database.
- rokob 1y agoThat is interesting, I’ll have to look into that further. I would expect Go to have similar issues because the RPS isn’t that much less. But maybe there is some knife edge here.
- neilv 1y agoOne reason to use CGI is legacy systems. A large, complex, and important system that I inherited was still using CGI (and it worked, because a rare "10x genuinely more productive" developer built it). Many years later, to reduce peak resource usage, and speed up a few things, I made an almost drop-in replacement library, to permit it to also run with SCGI (and back out easily to CGI if there was a problem in production). https://docs.racket-lang.org/scgi/ https://docs.racket-lang.org/scgi/ Another reason to use CGI is if you have a very small and simple system. Say, a Web UI on a small home router or appliance. You're not going to want the 200 NPM packages, transpilers and build tools and 'environment' managers, Linux containers, Kubernetes, and 4 different observability platforms. (Resume-driven-development aside.) A disheartening thing about most my recent Web full-stack project was that I'd put a lot of work into wrangling it the way Svelte and SvelteKit wanted, but upon finishing, wasn't happy with the complicated and surprisingly inefficient runtime execution. I realized that I could've done it in a fraction of the time and complexity -- in any language with convenient HTML generation, a SQL DB library, and an HTTP/CGI/SCGI-ish library, plus a little client-side JS).
- ptsneves 1y agoI found that ChatGPT revived vanilla javascript and jquery for me. Most of the chore part is done by chatgpt and the mental model of understanding what it wrote is very light and often single file. It is also easily embedded in static file generators. On the contrary Vue/React have a lot of context required to understand and mentally parse. On react the useCallback/useEffect/useMemo make me need to manually manage dependencies. This really reminds me of manual memory management in C, with perhaps even more pitfalls. On vue the difference between computed, props and vanilla variables. I am amazed that the supposed more approachable part of tech is actually more complex than regular library/script programming.
- maxwell 1y agoI've had a similar experience. Generating Vue/React scaffolding is nice, but yeah debugging and refactoring require the additional context you described. I've been using web components lately on personal projects, nice to jump into comprehensible vanilla JS/HTML/CSS when needed.
- oxcabe 1y agoIt'd be interesting to compare the performance of the author's approach to an analogous design that changes CGI for WASI, and scripts/binaries to Wasm.
- IshKebab 1y agoWould it? It would be exactly the same but a bit slower because of the WASM overhead.
- kragen 1y agoNo, Linux typically takes about 1ms to fork/exit/wait and another fraction of a millisecond to exec, and was only getting about 140 requests per second per core in this configuration, while creating a new WASM context is closer to 0.1ms. I suspect the bottleneck is either the web server or the database, not the CGI processes.
- pkal 1y agoI have recently been writing CGI scripts for the web server of our universities computer lab in Go, and it has been a nice experience. In my case, the Guestbook doesn't use SQLite but I just encode the list of entries using Go's native https://pkg.go.dev/encoding/gob https://pkg.go.dev/encoding/gob format, and it worked out well -- and critically frees me from using CGO to use SQLite! But in the end efficiency isn't my concern, as I have almost not visitors, what turns out to be more important is that Go has a lot of useful stuff in the standard library, especially the HTML templates, that allow me to write safe code easily. To test the statement, I'll even provide the link and invite anyone to try and break it: https://wwwcip.cs.fau.de/~oj14ozun/guestbook.cgi https://wwwcip.cs.fau.de/~oj14ozun/guestbook.cgi (the worst I anticipate happening is that someone could use up my storage quota, but even that should take a while).
- kragen 1y agoHow do you protect against concurrency bugs when two visitors make guestbook entries at the same time? With a lockfile? Are you sure you won't write an empty guestbook if the machine gets unexpectedly powered down during a write? To me, that's one of the biggest benefits of using something like SQLite.
- pkal 1y agoThat is exactly what I do, and it works well enough because if the power-loss were to happen, I wouldn't have lost anything of crucial value. But that is admittedly a very instance-specific advantage I have.
- kragen 1y agoThere's a fsync/close/rename dance that ext4fs recognizes as a safe, durable atomic file replacement, which is often sufficient for preventing data loss in cases like this.
- masklinn 1y agoFWIW POSIX requires that rename(2) be atomic, it's not just ext4, any POSIX FS should work that way. However this still requires a lockfile because while rename(2) is an atomic store it's not a full CAS, so you can have two processes reading the file concurrently, doing their internal update, writing to a temp file, then rename-ing to the target. There will be no torn version of the reference file, but the process finishing last will cancel out the changes of the other one. The lockfile can be the "scratch" file as open(O_CREAT | O_EXCL) is also guaranteed to be atomic, however now you need a way to wait for that path to disappear before retrying.
- 0xbadcafebee 1y ago> No one should ever run a Bash script under CGI. It’s almost impossible to do so securely, and performance is terrible. Actually shell scripting is the perfect language for CGI on embedded devices. Bash is ~500k and other shells are 10x smaller. It can output headers and html just fine, you can call other programs to do complex stuff. Obviously the source compresses down to a tiny size too, and since it's a script you can edit it or upload new versions on the fly. Performance is good enough for basic work. Just don't let the internet or unauthenticated requests at it (use an embedded web server with basic http auth).
- kragen 1y agoEasy uploading of new versions is a good point, and I agree that the likely security holes in the bash script are less of a concern if only trusted users have access to it. However, about 99% of embedded devices lack an MMU, much less 50K of storage, which makes it hard to run Unix shells on them.
- 0xbadcafebee 1y agoBusybox runs MMU-less and has ash built in. It also has a web server! It can be a little chonky but you can remove unneeded components. Things like wireless routers and other devices that have a decent amount of storage are a good platform for it
- kragen 1y agoYeah, a lot of wireless routers would have no trouble. A lot of them do in fact have MMUs. I wonder if you could get Busybox running on an ESP32? Probably not most 8051s, though, or AVR8s.
- 0xbadcafebee 1y agoLooks like the ESP32-S3 model works with modern Linux (it's so bloated compared to the old 2.0/2.2/2.4 branches...) The other option seems to be Apache NuttX as an RTOS (runs on all ESP32), and then Busybox w/hush or Toybox w/toysh. The more shell features you need, the more space it's gonna take, but technically 64 kB flash is possible.
- kragen 1y agoThis is a followup to Gold's previous post that served 200 million requests per day with CGI, which Simon Willison wrote a post about, which we had a thread about three days ago at https://news.ycombinator.com/item?id=44476716 https://news.ycombinator.com/item?id=44476716. It addresses some of the misconceptions that were common in that thread. Summary: - 60 virtual AMD Genoa CPUs with 240 GB (!!!) of RAM - bash guestbook CGI: 40 requests per second (and a warning not to do such a thing) - Perl guestbook CGI: 500 requests per second - JS (Node) guestbook CGI: 600 requests per second - Python guestbook CGI: 700 requests per second - Golang guestbook CGI: 3400 requests per second - Rust guestbook CGI: 5700 requests per second - C guestbook CGI: 5800 requests per second https://github.com/Jacob2161/cgi-bin https://github.com/Jacob2161/cgi-bin I wonder if the gohttpd web server he was using was actually the bottleneck for the Rust and C versions?
- dengolius 1y agoWhat is the reason to choose gohttpd? I mean there are a lot of non standard libraries for go that are pretty fast or faster then gohttpd - https://github.com/valyala/fasthttp/ https://github.com/valyala/fasthttp/ as example
- exabrial 1y agoCurrently in Europe. Earlier, was trying to use the onboard wifi on a train, which has frequent latency spikes as you can imagine. It never quite drops out, but latency does vary between 50ms-5000ms on most things. I struggled for _15 mins_ on yet another f#@%ng-Javascript-based-ui-that-does-not-need-to-be-f#@%ng-Javascript, simply trying to reset my password for Venmo. Why... oh why... do we have to have 9.1megabytes of f#@*%ng scripts just to reset a single damn password? This could be literally 1kb of HTML5 and maybe 100kb of CSS? Anyway, this was a long way of saying I welcome FastCGI and server side rendering. Js need to be put back into the toys bin... er trash bin, where it belongs.
- carodgers 1y agoLooks like CGI was recently removed from python 3. https://docs.python.org/3/library/cgi.html https://docs.python.org/3/library/cgi.html What is a modern python-friendly alternative?
- kragen 1y agoPython has a policy against maintaining compatibility with boring technology. We discussed this at some length in this thread the other day at https://news.ycombinator.com/item?id=44477966 https://news.ycombinator.com/item?id=44477966; many people voiced their opposition to the policy. The alternatives suggested for the specific case of the cgi module were: - wsgiref.handlers.CGIHandler, which is not deprecated yet. gvalkov provided example code for Flask at https://news.ycombinator.com/item?id=44479388 https://news.ycombinator.com/item?id=44479388 - use a language that isn't Python so you don't have to debug your code every year to make it work again when the language maintainers intentionally break it - install the old cgi module for new Python from https://github.com/jackrosenthal/legacy-cgi https://github.com/jackrosenthal/legacy-cgi - continue using Python 3.12, where the module is still in the standard library, until mid-02028
- antoineleclair 1y agoWe used CGI to add support for extensions in Disco (https://disco.cloud/ https://disco.cloud/). It's so simple and it can run anything, and it was also relatively easy to have the CGI script run inside a Docker container provided by the extension. In other words, it's so flexible that it means the extension developers would be able to use any language they want and wouldn't have to learn much about Disco. I would probably not push to use it to serve big production sites, but I definitely think there's still a place for CGI. In case anyone is curious, it's happening mostly here: https://github.com/letsdiscodev/disco-daemon/blob/main/disco/endpoints/cgi.py https://github.com/letsdiscodev/disco-daemon/blob/main/disco...
- kragen 1y agoThis is an interesting idea!
- hedgehog 1y agoCGI still makes a lot of sense when there are many applications that each only get requests at a low rate. Pack them onto servers, no RAM requirement unless actively serving a request. If the most of the requests can be served straight from static files by the web server then it's really only the write rate that matters, so even a high traffic sites could be a good match. With sendfile and kTLS the static content doesn't even need to touch user space.
- arzookanak 1y ago[dead]