3 ms·
Show HN: Python file streaming 237MB/s on $8/M droplet in 507 lines of stdlib
Quick Links:
- PyPI: https://pypi.org/project/axon-api/ https://pypi.org/project/axon-api/
- GitHub: https://github.com/b-is-for-build/axon-api https://github.com/b-is-for-build/axon-api
- Deployment Script: https://github.com/b-is-for-build/axon-api/blob/master/examples/deployment_scripts/deploy-axon.sh https://github.com/b-is-for-build/axon-api/blob/master/examp...
Axon is a 507-line, pure Python WSGI framework that achieves up to 237MB/s file streaming on $8/month hardware. The key feature is the dynamic bundling of multiple files into a single multipart stream while maintaining bounded memory (<225MB). The implementation saturates CPU before reaching I/O limits.
Technical highlights:
- Pure Python stdlib implementation (no external dependencies)
- HTTP range support for partial content delivery
- Generator-based streaming with constant memory usage
- Request batching via query parameters
- Match statement-based routing (eliminates traversal and probing)
- Built-in sanitization and structured logging
The benchmarking methodology uses fresh Digital Ocean droplets with reproducible wrk tests across different file sizes. All code and deployment scripts are included.
- SkiFire13 1y ago> The implementation saturates CPU before reaching I/O limits. Is this supposed to be a pro?
- b_llc 1y agoGood question! Yes, CPU saturation is the desired behavior here. The multipart streaming workload is inherently expensive. The cost of generating boundaries and constructing headers scales with request count and payload size. The architecture demonstrates efficient resource utilization: bounded memory usage (<225MB) while maximizing CPU throughput. CPU saturation with bounded memory means performance scales predictably with processing power. On multicore systems, you can leverage multiple processes to effectively utilize all cores. Alternatively, you can distribute the workload horizontally using droplets as cost-efficient instances.
- nomel 1y agoBy I/O limits, do you mean memory size limits? If so, this wording will lead to much confusion, since addressing limits (which is what a memory limit is) is a somewhat unusual use of "I/O limits" which, in a streaming context, most would perceive as a bandwidth limit (either memory or network).
- b_llc 1y agoBy I/O limits, I meant network bandwidth and disk throughput limits, not memory capacity. Thanks for pointing out the ambiguity.
- nomel 1y agoNow I'm more confused. An infinitely efficient system would saturate the network. An infinitely inefficient system would saturate the CPU. " The implementation saturates CPU before reaching I/O limits." is true infinitely inefficient system, but false for an infinitely efficient system. That means it's an undesirable. The metric that actually matters is efficiency of the task, given a hardware constraint. In this context, that's entirely network throughput (streaming ability/hardware, with hardware being constant, you can just compare streaming ability directly). For a litmus test of the concept, if you rewrote this in C or Rust, would the CPU bottleneck earlier or later? Would the network throughput be closer or further from its bottleneck?
- b_llc 1y agoYou're right - this represents computational duress, not optimal efficiency. The 1 CPU struggles to handle the 50 concurrent user scenario and was chosen to demonstrate worst-case behavior rather than peak performance. I intended to stress test the framework. I did not mean to indicate that CPU saturation is ideal but rather highlight that performance remained predictable even at the limits. Lower-level languages would certainly offer higher performance. I was hoping to showcase how Python can perform when architecture is restrained. The goal was to show that careful design choices (bounded memory, generator-based streaming) can maintain predictable behavior even when computational resources are exhausted.
- skyzouwdev 1y ago[dead]
- sc68cal 1y ago> The implementation saturates CPU before reaching I/O limits So, I did look over the code and the thing that I walked away asking was "isn't this sort of the reason why sendfile(2) was developed?"
- b_llc 1y agoAxon generates dynamic multipart responses with boundaries and headers to bundle files specified via query parameters. sendfile handles "serve this specific file" but does not handle "bundle these N files into a multipart response." For static file serving, sendfile would be the better choice.
- nomel 1y agoThis appears to be an AI codebase, with AI written replies in the comments here.
- b_llc 1y agoThe code is 100% mine. The architecture and code evolved over 5-years as Axon adapted for a number of projects. I started to revamp the project 2-3 weeks ago. Throughout that process, I used Claude to check for logical errors and simplification opportunities. Claude generated the readme and shell scripts in entirety and commented the codebase for clarity. Edit: The hello-world.html was also generated.
- b_llc 1y agoI'd like to thank everyone for the feedback. I'm going to reflect on the feedback, rewrite the article with better explanations of the problem domain, and its limitations. I've clearly failed to articulate the core capability: dynamic file bundling via query parameters into a unified multipart stream. I'll check back over the next few days but won't be monitoring constantly. If you want to discuss this project privately, I’m available by email.