7 ms·
Serving 200M requests per day with a CGI-bin
- Traubenfuchs 1y agoTry an apache tomcat 11 next. You can just dump .jsp files or whole java servlet applications as .war file via ssh and it will just work! One shared JVM for maximum performance! It can also share db connection pools, caches, etc. among those applications! Wow!
- atemerev 1y agoI miss this so much. Deployment should be just copying the file (over ssh or whatever). Why people overcomplicated it so much?
- rexreed 1y agoPHP can work the same way. Push / FTP / SFTP PHP file to directory, deployed.
- Twirrim 1y agoWe used to use symlinks to enable atomic operations, too. e.g. under /var/www/ we'd have /var/www/webapp_1.0, and have a symlink /var/www/webapp pointing to it. When there was a new version, upload it to /var/www/webapp_1.1, and then to bring it live, just update the symlink. Need to roll back? Switch the symlink back.
- trinix912 1y agoWouldn't that cause problems when someone would find the old version and corrupt the data with it? Or would only the current version be accessible from the outside?
- indigodaddy 1y agoHow would an external user find the old version?
- Twirrim 1y agoYour apache/whatever config would be pointed to the symlink location. No one would be able to get at the old versions of the site. We'd use this approach not just for webapps, but versions of applications we'd build in house, bundles of scripts, whatever.
- mrweasel 1y ago> Why people overcomplicated it so much? Because a lot of production software is half-baked. If you have to hand over an application to an operations team you need documentation, instrumentation, useful logging, error handling and a ton of other things. Instead software is now stuffed into containers that never receive security updates, because containers make things secure apparently. Then the developers can just dump whatever works into a container and hide the details. To be fair most of that software is also way more complex today. There are a ton of dependencies and integrations and keeping track of them is a lot of work. I did work with an old school C programmer that complained that a system we deployed was a ~2GB war file, running on Tomcat and requiring at least 8GB of memory and still crashed constantly. He had on multiple occasions offered to rewrite the how thing in C, which he figured would be <1MB and requiring at most 50MB of RAM to run. Sadly the customer never agreed, I would have loved to see if it had worked out as he predicted.
- nickjj 1y agoDocker helps with this nowadays. Of course you need to understand setting things up the first time you do it but once you know, it can apply to any tech stack. I develop and deploy Flask + Rails + Django apps regularly and the deploy process is the same few Docker Compose commands. All of the images are stored the same with only tiny differences in the Dockerfile itself. It has been a tried and proven model for ~10 years. The core fundamentals have held up, there's new features but when I look at Dockerfiles I've written in 2015 vs today you can still see a lot of common ideas.
- atemerev 1y agoDocker makes things opaque. You deploy black boxes and have no idea how the components there operate. Which is fine for devops, but as a software engineer, I prefer to work without Docker (and having to use Docker to install something on a local machine is an abomination, of course).
- stackskipton 1y agoOps here, I mean you still can if you use something like Golang or Java/.Net self-contained. However, the days of "Just transfer over PHP files" ignore the massive setup that Ops had to do to get web server into state where those files could just be transferred over and care/feeding required to keep the web server in that state. Not to mention endless frustration any upgrades would cause since we had to get all teams onboard with "Hey, we are upgrading PHP 5, you ready?" and there was always that abandoned app that couldn't be shut down because $BusinessReasons. Containers have greatly helped with those frustration points and languages self-hosting HTTP have really made stuff vastly better for us Ops folks.
- mrkeen 1y agoPerhaps. Over SSH? With a password or with a key? Do all employees share the same private key or do keys need to get added and removed when employees come and go. Is there one server or three (Are all deployment instructions done manually in triplicate?). When tomcat itself is upgraded, do you just eat the downtime? What about the system package upgrades or the OS? Which file should be copied over - whatever a particular Dev feels is the latest?
- miroljub 1y agoIt depends on the application usage pattern. For heavily used applications, sure, it's an excellent choice. But imagine having to host 50 small applications each serving a couple of hundreds requests per day. In that case, the memory overhead of Tomcat with 50 war files is much bigger than a simple Apache/Nginx server with a CGI script.
- whartung 1y agoThe other issue with Tomcat is that a single bad actor can more easily compromise the server. Not saying that can't happen with CGI, but since Tomcat is a shared environment, it's much more susceptible to it. This is why shared, public Tomcat hosting never became popular compared to shared CGI hosting. A rogue CGI program can be managed by the host accounting subsystem (say, it runs too long, takes up too much memory, etc.), plus all of the other guards that can be put on processes. The efficiency of CGI, specifically for compiled executables, is that the code segments are shared in virtual memory, so forking a new one can be quite cheap. While forking a new Perl or PHP process shares that, they still need to repeatedly go through the parsing phase. The middle ground of "p-code" can work well, as those files are also shared in the buffer cache. The underlying runtime can map the p-code files into the process, and those are shared across instances also. So, the fork startup time, while certainly not zero, can be quite efficient.
- grandiego 1y agoI believe even today there's no way to control/isolate memory leaks on a per-war basis.
- lenkite 1y agoWell, James Gosling was working on the Java Isolates spec, but then Sun experienced financial difficulties and most of the future-thinking JSR (java specification request) work got frozen. Oracle had different priorities after acquisition - moving away from big, fat, enterprise app servers was a big no-no.
- immibis 1y agoTomcat/Jakarta EE/JSP is a surprisingly solid stack. I only tried it once. Everything mostly just worked, and worked pretty well. You get to write pages PHP-style (interspersed HTML and code) but with the full power of Java instead of a hack language like PHP. Of course that paradigm may not suit everyone, but you also don't have to handle requests that way as you can also install pure Java routes. It supports websockets. You can share data between requests since it's a single-process multi-threaded model, so you can write something with real-time communication. You can also not do that; JSP code (and of course local variables) is scoped to a request. Deployment is very easy: drop the new webapp (a single file) in the webapps directory, by any method you like e.g. scp, and when Tomcat notices the new file is there, it transparently loads the new app and unloads the old one. You do have to watch out for classloader leaks that would prevent the old app being garbage-collected, though - downside of a single-process model.
- palmfacehn 1y agoI'm still happily using Jetty for webapp backends.
- sugarpimpdorsey 1y agoSurprised with the choice of Apache. There are better choices for serving CGI nowadays. The only reason for still running Apache is you have legacy cruft that requires Apache (like .htaccess).
- mrweasel 1y agoApache is still a solid option. It does everything, works with everything and is easy to configure. Performance is perfectly fine for ~99% of everything.
- chgs 1y agoI host 1200 vhosts off apache as an authenticating proxy, and run Al sorts of random scripts. This is all internal use though, I don’t need to scale to hundreds of concurrent users let alone thousands. Apache and cgi bin is fine.
- immibis 1y agoI think the point of the experiment was to see how fast the old-school tech stack would go on modern hardware.
- Twirrim 1y agoApache's httpd is great, reliable, fast, and feature-full, and you don't have to deal with Nginx's ongoing conflict over the the open source vs commercial offerings. That conflict has caused needless pain, like e.g. that quirk around dns resolution where if you put the hostname under proxy_pass it only used to resolve it on start-up and ignored TTL (not sure if it's still doing that). There were work-arounds on the open source version, like using a variable instead, but that wasn't necessary in the commercial offering.
- jacob2161 1y agoIn the post I also showed results for a little `gohttpd` program running the CGI program: https://github.com/Jacob2161/cgi-bin/blob/main/gohttpd/main.go https://github.com/Jacob2161/cgi-bin/blob/main/gohttpd/main.... See as "Benchmarking (writes|reads) using Go net/http" It was faster but not by very much. Running CGI programs is just forking processes, so Apache's forking model works just about as well as anything else.
- simonw 1y agoI got my start in the CGI era, and it baked into me an extremely strong bias against running short-lived subprocesses for things. We invented PHP and FastCGI mainly to get away from the performance hit of starting a new process just to handle a web request! It was only a few years ago that I realized that modern hardware means that it really isn't prohibitively expensive to do that any more - this benchmark gets to 2,000/requests a second, and if you can even get to a few hundred requests a second it's easy enough to scale across multiple instances these days. I have seen AWS Lambda described as the CGI model reborn and that's a pretty fair analogy.
- pjc50 1y ago> We invented PHP and FastCGI mainly to get away from the performance hit of starting a new process just to handle a web request! Yes! Note that the author is using a technology that wasn't available when I too was writing cgi_bin programs in the 00's: Go. It produces AOT compiled executables but is also significantly easier to develop in and safer than trying to do the same with C/C++ in the 00's. Back then we tended to use Perl (now basically dead). Perl and Python would incur significant interpreter startup and compilation costs. Java was often worse in practice. > I have seen AWS Lambda described as the CGI model reborn and that's a pretty fair analogy. Yes, it's almost exactly identical to managed FastCGI. We're back to the challenges of deployment: can't we just upload and run an executable? But of course so many technologies make things much, much more complicated than that.
- shrubble 1y agoI know of two large telecoms that internally develop with Perl, and a telecom product sold by Oracle that heavily relies on Perl. For text munging etc. it is still used, though I grant that other languages like Python are more popular.
- geocar 1y agoI think you might have found that CGI scripts deployed as statically-linked C binaries, with some attention given to size, you might've not been so disappointed. The "performance hit of starting a new process" is bigger if the process is a dynamically-linked php interpreter with gobs of shared libraries to load, and some source file, reading parsing compiling whatever, and not just by a little bit, always has been, so what the author is doing using go, I think, would still have been competitive 25 years ago if go had been around 25 years ago. Opening an SQLite database is probably (surprisingly?) competitive to passing a few sockets through a context switch, across all server(ish) CPUS of this era and that, but both are much faster than opening a socket and authenticating to a remote mysql process, and programs that are not guestbook.cgi often have many more resource acquisitions which is why I think FastCGI is still pretty good for new applications today.
- cb321 1y agoEven back in the 1990s, CGI programs written in C were lightning fast. It just was (is) an error prone environment. Any safer modern alternative like the article's Go program or Nim or whatever not making database connections will be very fast & low latency to localhost - really similar to a CLI utility where you fork & exec. It's not free, but it's not that expensive compared to network latencies then or now. People/orgs do tend to get kind of addicted to certain technologies that can interact poorly with the one-shot model, though. E.g., high start up cost Python interpreters with a lot of imports are still pretty slow, and people get addicted to that ecosystem and so need multi-shot/persistent alternatives. The one-shot model in early HTTP was itself a pendulum swing from other concerns, e.g. ftp servers not having enough RAM for 100s of long-lived, often mostly idle logins.
- foobiekr 1y agoYou know, CGI with pre-forking (for latency hiding) and a safer language (like Rust) would be a great system to work on. Put the TLS termination in a nice multi-threaded web server (or in a layer like CloudFront). No lingering state, very easy to dump a core and debug, nice mostly-linear request model (no callback chains, etc.) and trivially easy to scale. You're just reading from stdin and writing to stdout. Glorious. Websockets adds a bit of complexity but almost none. The big change in how we build things was the rise of java. Java was too big, too bloated, too slow, etc. so people rapidly moved into multi-threaded application servers, all to avoid the cost of fork() and the dangers of C. We can Marie Kondo this shit and get back to things that are simple if we want to. I don't even like Rust and this sounds like heaven to me. Maybe someone will come up with a way to make writing the kind of web-tier backend code in Rust easy by hiding a lot of the tediousness and/or complexity in a way that makes this appealing to node/js, php and python programmers.
- cb321 1y agoThis is not to disagree, but to agree adding some detail... :-) Part of the Java rise was C/C++ being error prone and syntax similarity with such, but this was surely intermingled with a full scale marketing assault by Sun Microsystems who at the time had big multi-socket SMP servers they wanted to sell with Solaris/etc. and part of that was the Solaris/Java threading. Really for a decade or two prior to that the focus was on true MMU-based hardware-enforced isolation with OS kernel clean-up (more like CHERI these days) not the compiler-enforced stuff like Rust does. I think you could have something more ergonomic than Perl/Python ever was and as practically fast as C/Rust with Nim (https://nim-lang.org/ https://nim-lang.org/). E.g., I just copied that guy's benchmark with a Nim stdlib std/cgi and got over 275M CGI/day to localhost on a 2016 CPU doing only 2 requesters & 2 http server threads. With some nice DSL easily written if you don't like any current ones you could get the "coding overhead" down to a tiny footprint. In fairness I did zero SQLite whatever, but also he was using a computer over 4x bigger and probably a GHz faster with some IPC lift as well. So, IF you had the network bandwidth (hint - usually you don't!), you could probably support billions of hits/day off a single server. To head off some lazy complaints, GC is just not an issue with a single threaded Nim program whose lifetime is hoped/expected to be short anyway. In many cases (just as with CLI utilities!) you could probably just let the OS reap memory, but, of course, it always "all depends" on a lot of context. Nim does reference counting anyway whereas most "fighting the GC" is actually fighting a "separate GC thread" (Java again, Go, D, etc.) trashing CPU caches or consuming DIMM bandwidth and so on. For this use, you probably would care more about a statically linked binary so you don't pay ld.so shared library set up overhead on every `exec`.
- petee 1y agoA fastcgi comparison would be interesting
- 10000truths 1y agoThe nice thing about CGI is that you don't have to reinvent isolation primitives for multi-tenant use cases. A bug in one request doesn't corrupt another request, due to process isolation. An infinite loop in one request doesn't DoS other requests, due to preemptive scheduling. You can kill long-running requests with rlimit. You can use per-tenant cgroups to fairly allocate resources like memory, CPU and disk/network I/O. You can use namespaces/jails and privilege separation to restrict what a request has access to.
- bob1029 1y ago> These days, we have servers with 384 CPU threads. Even a small VM can have 16 CPUs. The CPUs and memory are much faster as well. With this hardware, if you reach for Kestrel you can easily do a few trillion requests per day. The development experience would be nearly identical - You can leverage the string interpolation operator for a PHP-like experience. LINQ and String.Join() open the door to some very terse HTML template syntax for tables and other nested elements. The hard part is knowing how to avoid certain landmines in the ecosystem (MVC/Blazor/EF/etc.). The whole thing can live in one top-level program file that is ran on the CLI, but you need to know the magic keywords - "Minimal APIs" - or you will find yourself in the middle of the wrong documentation.
- anoojb 1y agoThe amount of Director/VP level promos that have come from creating abstractions on top of core technology no one gets rewarded for is amazing.
- lukeasrodgers 1y agoDoes anyone know if the benchmarking tool the author uses, plow, avoids coordinated omission (https://www.scylladb.com/2021/04/22/on-coordinated-omission/ https://www.scylladb.com/2021/04/22/on-coordinated-omission/)? I didn’t see any mention in the docs, and haven’t been able to peruse the source code yet.
- p0w3n3d 1y agoThis was something that I've been suspecting for some time. We're moving towards complicated architecture while having possibility to use good ol' tech with newest CPUs. I've been asked about architecture of a stocks ticker that would serve millions of clients to show them on their phone the current stock price. First thought was streams, Kafka, pubsub etc but then I came up with static files on a server. I wonder how much would it cost though
- nine_k 1y agoAFAICT the latency of any non-trivial web API is determined by the latency of DB queries, ML model queries, and suchlike. The rest is trivial in comparison, even when using slow languages like Python. If all you need is to return rarely-changing data, especially without wasting time on authorization, you can easily approach the limits of your NIC.
- firefoxd 1y agoI've created a visualizer for apache requests with the workers, queues and whatnot [0]. You can load the demo to view real traffic comic from HN earlier this year. [0]: https://www.ibrahimdiallo.com/reqvis https://www.ibrahimdiallo.com/reqvis Note: works best on desktop browsers for now.
- znpy 1y agoOp is probably missing the point: 2400 requests/second is abysmally low on a modern 16 core cpu. At the very least go for FastCGI, for christ’s sake…
- EasyMark 1y agoI think he was just simply trying to revisit CGI and see how it compares from back in the day. Doing FastCGI would have been pointless if that was his primary goal.
- olcarl75 1y ago2400 requests is somewhat decent. Having seen companies like amazon, where their website backend can only handle a few hundred RPS per box, 2400 ain't that bad.
- apgwoz 1y agoYeah, instead of relooking at this, we went off and built a new paradigm “serverless functions.” Obviously, serverless functions, like those via Lambda, have some other safety mechanisms in place (eg micro vms), but you could probably get pretty far with CGI and adjusting capabilities and such, with far less complexity.
- kiitos 1y ago> I used plow to make concurrent HTTP requests and measure the results. If this refers to https://github.com/six-ddc/plow https://github.com/six-ddc/plow then -- oops! lots of issues in that repo, no tests, etc. etc. The results in the README are also pretty clearly unsound! In both scenarios, writes were faster than reads? _edit_: I guess because the writes all returned 3xx, oops again! Probably don't take this article's claims at face value...
- jacob2161 1y ago(I didn't downvote you) plow may not be the best tool that exists but it does make concurrent HTTP requests and generate metrics for them successfully. The writes returned 3xx because the handler returns a redirect, so this is expected.
- kiitos 1y ago> plow may not be the best tool that exists but it does make concurrent HTTP requests and generate metrics for them successfully. HTTP load testing is a problem area that is much more subtle than it seems. I've no doubt that plow does what you're saying here, but, without any tests whatsoever, I have serious doubts that it does so correctly, particularly if/when the load test starts bumping up against any of the numerous bottlenecks that can affect results and their measurements. > The writes returned 3xx because the handler returns a redirect, so this is expected. Yeah, but, unless `plow` actually follows that redirect, it's not really measuring the actual end-to-end latency, and further the guestbook.cgi returns 301 See Other for both valid requests (that performed a write) and invalid requests (that didn't).
- jacob2161 1y agoYou have a fair point about plow. I've used it enough to know that it basically works at spamming HTTP requests at a specified concurrency. When I'm doing something where I want better accuracy I've tended to use K6 and Vegeta. > Yeah, but, unless `plow` actually follows that redirect, it's not really measuring the actual end-to-end latency... Not in this case, since the purpose here was to measure the POST (write) not the subsequent GET (read) that a browser would do after the redirect.
- TZubiri 1y agoNonono, that can't be right, you need Kubernetes and Kafka and RabbitMQ and Graphana and a 300K$/day bill
- HocusLocus 1y agoand a DeLorean and a flux capacitor
- perlgeek 1y agoCGI scripts were one of the reason that perl was optimized for a quick startup time. I just did a `time perl -e ''` (starting perl, executing an empty program), it took 5ms. 33ms with python3, 77ms with ruby.
- cb321 1y agoWhile all you say is true, it bears note that it didn't need to be decisive. The current mob branch of tcc is such that a `#!/bin/tcc -run` "script" is about 1.3x faster than perl</dev/null on two CPUs I tried. Besides your two slower examples, Julia and Java VMs and else thread PHP also have really big start up times. As I said up top, people just get addicted to "big environments". Lisp culture would do that with images and this is part of where the "emacs is bloated" meme came from. Anyway, at the time getline wasn't even standardized (that was 2008 POSIX - and still not in Windows; facepalm emoji), but you could write a pretty slick little library for CGI in a few hundred..few thou' lines of C. Someone surely even did. But things go by "reputation" and people learn what their friends tell them to, by and large. So, CGI was absolutely the thing that made the mid to late 90s "Perl's moment".
- dolmen 1y agoCGI is not anymore a major use case for Perl. The modern Perl web apps are built on PSGI and production deployment is as a long lived Perl process. I wouldn't be surprised to learn that Perl startup time has drifted. Need benchmark.
- cb321 1y agoThe major modern use case I know of is command-line utilities which also benefit from low start-up. Of course, that doesn't mean there hasn't been "perf rot" over the decades as you say. Such rot should never surprise anyone. :-) Some perl5 lover should take the time to compile all those 5.6 to 5.42 versions on the same host OS/CPU and do a performance comparison and create a nice chart for the world to trap and maybe correct such performance regressions. I just tried getting 5.8.9 to compile on modern Linux with gcc-15, and it seemed like a real PITA. (Earlier didn't even ./Configure -des right.)
- riobard 1y agoCan we please stop using request/day as a perf metric?
- mediumsmart 1y agoThat is the future. Static sites no JavaScript, +SQLite +cgi-bin, NavMobile toggle is a page and done.
- ksec 1y ago> I ran these benchmarks on an older 16-thread AMD 3700X The 3700X is an 8 Core Zen 2 CPU. Or about 150 RPS per vCPU. By EOY we will have 256 Core Zen 6c EPYC, on a Dual Socket that is 512 Core or 1024 Thread. 153,600 RPS. And in the old days a RPS is a single pageview, which isn't the case in the modern world. CPU Core still have a healthy 5 - 10 years cost reduction roadmap. I wonder if CGI will make a comeback someday.
- josephernest 1y agoWhy would it be limited to ~ 100 connections on a 1-4 GB RAM server? Out of curiosity if we fork() httpd and exec() the cgi handler, it doesn't take the same RAM as the parent process and it could just take a few KB or MB, is that right? So I guess 1000+ concurrent connections even on a small server is possible.
- monkeyelite 1y agoYep. In the age of security and reliability, we should revisit process level encapsulation like this.
- tdiff 1y ago> It’s almost never going to be the best choice these days, but it’s definitely viable. Could someone please explain what the better options are?
- notcrazylol 1y agoCool. Kinda what serverless architecture does but way easier... and cheaper. Is anyone doing this for a real business usecase?
- sporkland 1y agoOver a decade ago I was starting a java process up with local MySQL and getting 45k rps read, request per thread with load wasn't hard to achieve. Not sure why this is an accomplishment.
- bhollis 1y agoSlow startup was definitely one reason to have long lived servers, but I’m surprised not to see the major other reasons: - Keep-alive/pooled connections to remote services can significantly reduce average latency for making those calls. - In-memory caches that allow amortizing repeated lookups across requests. Just those two alone mean that a serious high performance server probably couldn’t get away with just CGI even ignoring startup time.