18 ms·
I use zip bombs to protect my server
- deleted 1y ago[deleted]
- codingdave 1y agoMildly amusing, but it seems like this is thinking that two wrongs make a right, so let us serve malware instead of using a WAF or some other existing solution to the bot problem.
- cratermoon 1y agoSomething like https://xeiaso.net/notes/2025/anubis-works/ https://xeiaso.net/notes/2025/anubis-works/
- xena 1y agoI did actually try zip bombs at first. They didn't work due to the architecture of how Amazon's scraper works. It just made the requests get retried.
- cookiengineer 1y agoDid you also try Transfer-Encoding: chunked and things like HTTP smuggling to serve different content to web browser instances than to scrapers?
- wiredfool 1y agoAmazon's scraper has been sending multiple requests per second to my servers for 6+ weeks, and every request has been returned 429. Amazon's scraper doesn't back off. Meta, google, most of the others with identifiable user agents back off, Amazon doesn't.
- toast0 1y agoIf it's easy, sleep 30 before returning 429. Or tcpdrop the connections and don't even send a response or a tcp reset.
- cratermoon 1y agoThat's a good way to self-DOS
- toast0 1y agoThat's why I said, if it's easy. On some server stacks it's no big deal to have a connection open for an extra 30 seconds; others, you need to be done with requests asap, even abuse. tcpdrop shouldn't self DOS though, it's using less resources. Even if other end does a retry, it will do it after a timeout; in the meantime, the other end has a socket state and you don't, that's a win.
- deathanatos 1y agoSo first, let me prefix this by saying I generally don't accept cookies from websites I don't explicitly first allow, my reasoning being "why am I granting disk read/write access to [mostly] shady actors to allow them to track me?" (I don't think your blog qualifies as shady … but you're not in my allowlist, either.) So if I visit https://anubis.techaro.lol/ https://anubis.techaro.lol/ (from the "Anubis" link), I get an infinite anime cat girl refresh loop — which honestly isn't the worst thing ever? But if I go to https://xeiaso.net/blog/2025/anubis/ https://xeiaso.net/blog/2025/anubis/ and click "To test Anubis, click here." … that one loads just fine. Neither xeserv.us nor techaro.lol are in my allowlist. Curious that one seems to pass. IDK. The blog post does have that lovely graph … but I suspect I'll loop around the "no cookie" loop in it, so the infinite cat girls are somewhat expected. I was working on an extension that would store cookies very ephemerally for the more malicious instances of this, but I think its design would work here too. (In-RAM cookie jar, burns them after, say, 30s. Persisted long enough to load the page.)
- lcnPylGDnU4H9OF 1y ago> Neither xeserv.us nor techaro.lol are in my allowlist. Curious that one seems to pass. IDK. Is your browser passing a referrer?
- cycomanic 1y agoJust FYI temporary containers (Firefox extension) seem to be the solution you're looking for. It essentially generates a new container for every tab you open (subtabs can be either new containers or in the same container). Once the tab is closed it destroys the container and deletes all browsing data (including cookies). You can still whitelist some domains to specific persistent containers. I used cookie blockers for a long time, but always ended up having to whitelist some sites even though I didn't want their cookies because the site would misbehave without them. Now I just stopped worrying.
- xena 1y agoYou're seeing an experiment in progress. It seems to be working, but I have yet to get enough data to know if it's ultimately successful or not.
- theandrewbailey 1y agoWAF isn't the right choice for a lot of people: https://news.ycombinator.com/item?id=43793526 https://news.ycombinator.com/item?id=43793526
- codingdave 1y agoAt least, not with the default rules. I read that discussion a few days ago and was surprised how few callouts there were that a WAF is just a part of the infrastructure - it is the rules that people are actually complaining about. I think the problem is that so many apps run on AWS and their default WAF rules have some silly content filtering. And their "security baseline" says that you have to use a WAF and include their default rules, so security teams lock down on those rules without any real thought put into whether or not they make sense for any given scenario.
- chmod775 1y agoTruly one my favorite thought-terminating proverbs. "Hurting people is wrong, so you should not defend yourself when attacked." "Imprisoning people is wrong, so we should not imprison thieves." Also the modern telling of Robin Hood seems to be pretty generally celebrated. Two wrongs may not make a right, but often enough a smaller wrong is the best recourse we have to avert a greater wrong. The spirit of the proverb is referring to wrongs which are unrelated to one another, especially when using one to excuse another.
- deleted 1y ago[deleted]
- zdragnar 1y ago> a smaller wrong is the best recourse we have to avert a greater wrong The logic of terrorists and war criminals everywhere.
- BlackFingolfin 1y agoAnd sometimes one man's terrorist is another's freedom fighter.... (Not to defend terrorism, but it's just not that simple)
- _Algernon_ 1y agoAnd also how fuctioning governments work: https://en.m.wikipedia.org/wiki/Monopoly_on_violence https://en.m.wikipedia.org/wiki/Monopoly_on_violence Do you really want to live in a society were all use of punishment to discourage bad behaviour in others? That is a game theoretical disaster...
- toss1 1y agoDefense and Offense are not the same. Crime and Justice are not the same. If you cannot figure that out, you ARE a major part of the problem. Keep thinking until you figure it out for good.
- impulsivepuppet 1y agoI admire your deontological zealotry. That said, I think there is an implied virtuous aspect of "internet vigilantism" that feels ignored (i.e. disabling a malicious bot means it does not visit other sites) While I do not absolve anyone from taking full responsibility for their actions, I have a suspicion that terrorists do a bit more than just avert a greater wrong--otherwise, please sign me up!
- imiric 1y agoThe web is overrun by malicious actors without any sense of morality. Since playing by the rules is clearly not working, I'm in favor of doing anything in my power to waste their resources. I would go a step further and try to corrupt their devices so that they're unable to continue their abuse, but since that would require considerably more effort from my part, a zip bomb is a good low-effort solution.
- bsimpson 1y agoThere's no ethical ambiguity about serving garbage to malicious traffic. They made the request. Respond accordingly.
- joezydeco 1y agoThis is William Gibson's "black ICE" becoming real, and I love it. https://williamgibson.fandom.com/wiki/ICE https://williamgibson.fandom.com/wiki/ICE
- gherard5555 1y agoThis book was so far ahead of its time
- petercooper 1y agoBased on the example in the post, that thinking might need to be extended to "someone happening to be using a blocklisted IP." I don't serve up zip bombs, but I've blocklisted many abusive bots using VPN IPs over the years which have then impeded legitimate users of the same VPNs.
- java-man 1y agoI think it's a good idea, but it must be coupled with robots.txt.
- cratermoon 1y agoAI scraper bots don't respect robots.txt
- jsheard 1y agoI think that's the point, you'd use robots.txt to direct Googlebot/Bingbot/etc away from countermeasures that could potentially mess up your SEO. If other bots ignore the signpost clearly saying not to enter the tarpit, that's their own stupid fault.
- reverendsteveii 1y agoThe ones that survive do
- forinti 1y agoI was looking through my logs yesterday. Bad bots don't even read robots.txt.
- extraduder_ire 1y agoThe worst ones treat it as a target.
- zzo38computer 1y agoI also had the idea of zip bomb to confuse badly behaved scrapers (and I have mentioned it before to some other people, although I did not implemented it). However, maybe instead of 0x00, you might use a different byte value. I had other ideas too, but I don't know how well some of them will work (they might depend on what bots they are).
- ycombinatrix 1y agoThe different byte values likely won't compress as well as all 0s unless they are a repeating pattern of blocks. An alternative might be to use Brotli which has a static dictionary. Maybe that can be used to achieve a high compression ratio.
- zzo38computer 1y agoI meant that all of the byte values would be the same (so they would still be repeating), but a different value than zero. However, Brotli could be another idea if the client supports it.
- dspillett 1y agoCompressing a sequence of any single character should give almost identical results length-wise (perhaps not exactly identical, but the difference will be vanishingly small). For example, with gzip using default options: me@here:~$ pv /dev/zero -s 10M -S | gzip -c | wc -c 10.0MiB 0:00:00 [ 122MiB/s] [=============================>] 100% 10208 me@here:~$ pv /dev/zero -s 100M -S | gzip -c | wc -c 100MiB 0:00:00 [ 134MiB/s] [=============================>] 100% 101791 me@here:~$ pv /dev/zero -s 1G -S | gzip -c | wc -c 1.00GiB 0:00:07 [ 135MiB/s] [=============================>] 100% 1042069 me@here:~$ pv /dev/zero -s 10M -S | tr "\000" "\141" | gzip -c | wc -c 10.0MiB 0:00:00 [ 109MiB/s] [=============================>] 100% 10209 me@here:~$ pv /dev/zero -s 100M -S | tr "\000" "\141" | gzip -c | wc -c 100MiB 0:00:00 [ 118MiB/s] [=============================>] 100% 101792 me@here:~$ pv /dev/zero -s 1G -S | tr "\000" "\141" | gzip -c | wc -c 1.00GiB 0:00:07 [ 129MiB/s] [=============================>] 100% 1042071 Two bytes difference for a 1GiB sequence of “aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa…” (\141) compared to a sequence of \000.
- altairprime 1y agoSee also (2017) HN, https://news.ycombinator.com/item?id=14707674 https://news.ycombinator.com/item?id=14707674
- wewewedxfgdf 1y agoI protected uploads on one of my applications by creating fixed size temporary disk partitions of like 10MB each and unzipping to those contains the fallout if someone uploads something too big.
- sidewndr46 1y agoWhat? You partitioned a disk rather than just not decompressing some comically large file?
- gchamonlive 1y agohttps://github.com/uint128-t/ZIPBOMB https://github.com/uint128-t/ZIPBOMB 2048 yottabyte Zip Bomb This zip bomb uses overlapping files and recursion to achieve 7 layers with 256 files each, with the last being a 32GB file. It is only 266 KB on disk. When you realise it's a zip bomb it's already too late. Looking at the file size doesn't betray its contents. Maybe applying some heuristics with ClamAV? But even then it's not guaranteed. I think a small partition to isolate decompression is actually really smart. Wonder if we can achieve the same with overlays.
- sidewndr46 1y agoWhat are you talking about? You get a compressed file. You start decompressing it. When the amount of bytes you've written exceeds some threshold (say 5 megabytes) just stop decompressing, discard the output so far & delete the original file. That is it.
- gchamonlive 1y agoThose files are designed to exhaust the system resources before you can even do these kinds of checks. I'm not particularly familiar with the ins and outs of compression algorithms, but it's intuitively not strange for me to have a a zip that is carefully crafted so that memory and CPU goes out the window before any check can be done. Maybe someone with more experience can give mode details. I'm sure though that if it was as simples as that we wouldn't even have a name for it.
- ChuckMcM 1y agoI sort of did this with ssh where I figured out how to crash an ssh client that was trying to guess the root password. What I got for my trouble was a number of script kiddies ddosing my poor little server. I switched to just identifying 'bad actors' who are clearly trying to do bad things and just banning their IP with firewall rules. That's becoming more challenging with IPV6 though. Edit: And for folks who write their own web pages, you can always create zip bombs that are links on a web page that don't show up for humans (white text on white background with no highlight on hover/click anchors). Bots download those things to have a look (so do crawlers and AI scrapers)
- bjoli 1y agoI am just banning large swaths of IPs. Banning most of Asia and the middle east reduced the amount of bad traffic by something like 98%.
- johnisgood 1y agoSame, using ipsets, and a systemd {service,timer} for updating the lists.
- 1970-01-01 1y agoWhy is it harder to firewall them with IPv6? I seems this would be the easier of the two to firewall.
- echoangle 1y agoMaybe it’s easier to circumvent because getting a new IPv6 address is easier than with IPv4?
- firesteelrain 1y agoI think they are suggesting the range of IPs to block is too high?
- CBLT 1y ago
- sgc 1y agoI am ignorant as to how most bots work. Could you have a second line of defense for bots that avoid this bomb: Dynamically generate a file from /dev/random and trickle stream it to them, or would they just keep spawning parallel requests? They would never finish streaming it, and presumably give up at some point. The idea would be to make it more difficult for them to detect it was never going to be valid content.
- shishcat 1y agoThis will waste your bandwidth and resources too
- sgc 1y agoThe idea is to trickle it very slowly, like keeping a cat occupied with a ball of fluff in the corner.
- uniqueuid 1y agoCats also have timeouts set for balls of fluff. They usually get bored at some point and either go away or attack you :)
- CydeWeys 1y agoYeah but in the mean time it's tying up a connection on your webserver.
- jeroenhd 1y agoIf the bot is connecting over IPv4, you only have a couple thousand connections before your server starts needing to mess with shared sockets and other annoying connectivity tricks. I don't think it's a terrible problem to solve these days, especially if you use one of the tarpitting implementations that use nftables/iptables/eBPF, but if you have one of those annoying Chinese bot farms with thousands of IP addresses hitting your server in turn (Huawei likes to do this), you may need to think twice before deploying this solution.
- stavros 1y ago
- _QrE 1y agoThere's a lot of creative ideas out there for banning and/or harassing bots. There's tarpits, infinite labyrinths, proof of work || regular challenges, honeypots etc. Most of the bots I've come across are fairly dumb however, and those are pretty easy to detect & block. I usually use CrowdSec (https://www.crowdsec.net/ https://www.crowdsec.net/), and with it you also get to ban the IPs that misbehave on all the other servers that use it before they come to yours. I've also tried turnstile for web pages (https://www.cloudflare.com/application-services/products/turnstile/ https://www.cloudflare.com/application-services/products/tur...) and it seems to work, though I imagine most such products would, as again most bots tend to be fairly dumb. I'd personally hesitate to do something like serving a zip bomb since it would probably cost the bot farm(s) less than it would cost me, and just banning the IP I feel would serve me better than trying to play with it, especially if I know it's misbehaving. Edit: Of course, the author could state that the satisfaction of seeing an IP 'go quiet' for a bit is priceless - no arguing against that
- KTibow 1y agoIt's worth noting that this is a gzip bomb (acts just like a normal compressed webpage), not a classical zip file that uses nested zips to knock out antiviruses.
- deleted 1y ago[deleted]
- d--b 1y agoZip libraries aren’t bomb proof yet? Seems fairly easy to detect and ignore, no?
- harrison_clarke 1y agoit'd be cool to have a proof of work protocol baked into http. like, a header that browsers understood
- layer8 1y agoBack when I was a stupid kid, I once did ln -s /dev/zero index.html on my home page as a joke. Browsers at the time didn’t like that, they basically froze, sometimes taking the client system down with them. Later on, browsers started to check for actual content I think, and would abort such requests.
- artursapek 1y ago[flagged]
- koolba 1y agoI hope you weren’t paying for bandwidth by the KiB.
- santoshalper 1y agoNah, back then we paid for bandwidth by the kb.
- slicktux 1y agoThat’s even worse! :)
- sandworm101 1y agoDevide by zero happens to everyone eventually. https://medium.com/@bishr_tabbaa/when-smart-ships-divide-by-zer0-uss-yorktown-4e53837f75b2 https://medium.com/@bishr_tabbaa/when-smart-ships-divide-by-... "On 21 September 1997, the USS Yorktown halted for almost three hours during training maneuvers off the coast of Cape Charles, Virginia due to a divide-by-zero error in a database application that propagated throughout the ship’s control systems." " technician tried to digitally calibrate and reset the fuel valve by entering a 0 value for one of the valve’s component properties into the SMCS Remote Database Manager (RDM)"
- astolarz 1y agoBad bot
- jawns 1y agoIs there any legal exposure possible? Like, a legitimate crawler suing you and alleging that you broke something of theirs?
- bilekas 1y agoPlease, just as a conversational piece, walk me through the potentials you might think there are ? I'll play the side of the defender and you can play the "bot"/bot deployer.
- echoangle 1y agoWell creating a bot is not per se illegal, so assuming the maliciousness-detector on the server isn’t perfect, it could serve the zip bomb to a legitimate bot. And I don’t think it’s crazy that serving zip bombs with the stated intent to sabotage the client would be illegal. But I’m not a lawyer, of course.
- bilekas 1y agoDisclosure, I'm not a lawyer either. This is all hypothetical high level discussion here. > it could serve the zip bomb to a legitimate bot. Can you define the difference between a legitimate bot, and a non legitimate bot for me ? The OP didn't mention it, but if we can assume they have SOME form of robots.txt (safe assumtion given their history), would those bots who ignored the robots be considered legitimate/non-legitimate ? Almost final question, and I know we're not lawyers here, but is there any precedent in case law or anywhere, which defines a 'bad bot' in the eyes of the law ? Final final question, as a bot, do you believe you have a right or a privilege to scrape a website ?
- echoangle 1y ago> Can you define the difference between a legitimate bot, and a non legitimate bot for me ? Well by default every bot is legitimate, an illegitimate bot might be one that’s probing for security vulnerabilities (but I’m not even sure if that’s illegal if you don’t damage the server as a side effect, ie if you only try to determine the Wordpress or SSHD version running on the server for example). > The OP didn't mention it, but if we can assume they have SOME form of robots.txt (safe assumtion given their history), would those bots who ignored the robots be considered legitimate/non-legitimate ? robots.txt isn’t legally binding so I don’t think ignoring it makes a bot illegitimate. > Almost final question, and I know we're not lawyers here, but is there any precedent in case law or anywhere, which defines a 'bad bot' in the eyes of the law ? There might be but I don’t know any. > Final final question, as a bot, do you believe you have a right or a privilege to scrape a website ? Well I’m not a bot but I think I have the right to build bots to scrape websites (and not get served malicious content designed to sabotage my computer). You can decline service and just serve error pages of course if you don’t like my bot.
- mahi_novice 1y agoDo you mind sharing your specs of your digital ocean droplet? I'm trying to setup one with less cost.
- foxfired 1y agoThe blog runs on a $6 digital ocean droplet. It's 1GB RAM and 25GB storage. There is a link at the end of the article on how it handles typical HN traffic. Currently at 5% CPU.
- mahi_novice 1y agoThanks for sharing!
- bilekas 1y ago> At my old employer, a bot discovered a wordpress vulnerability and inserted a malicious script into our server I know it's slightly off topic, but it's just so amusing (edit: reassuring) to know I'm not the only one who, after 1 hour of setting up Wordpress there's a PHP shell magically deployed on my server.
- ianlevesque 1y agoYes, never self host Wordpress if you value your sanity. Even if it’s not the first hour it will eventually happen when you forget a patch.
- sunaookami 1y agoHosting WordPress myself for 13 years now and have no problem :) Just follow standard security practices and don't install gazillion plugins.
- carlosjobim 1y agoThere's a lot of essential functionality missing from WordPress, meaning you have to install plugins. Depending on what you need to do. But it's such a bad platform that there really isn't any reason for anybody to use WordPress for anything. No matter your use case, there will be a better alternative to WordPress.
- aaronbaugher 1y agoCan you recommend an alternative for a non-technical organization, where there's someone who needs to be able to edit pages and upload documents on a regular basis, so they need as user-friendly an interface as possible for that? Especially when they don't have a budget for it, and you're helping them out as a favor? It's so easy to spin up Wordpress for them, but I'm not a fan either. I've tried Drupal in the past for such situations, but it was too complicated for them. That was years ago, so maybe it's better now.
- Scoundreller 1y agoAttacked Over Tor [2017] https://www.hackerfactor.com/blog/index.php?/archives/762-Attacked-Over-Tor.html https://www.hackerfactor.com/blog/index.php?/archives/762-At...
- kazinator 1y agoI deployed this, instead of my usual honeypot script. It's not working very well. In the web server log, I can see that the bots are not downloading the whole ten megabyte poison pill. They are cutting off at various lengths. I haven't seen anything fetch more than around 1.5 Mb of it so far. Or is it working? Are they decoding it on the fly as a stream, and then crashing? E.g. if something is recorded as having read 1.5 Mb, could it have decoded it to 1.5 Gb in RAM, on the fly, and crashed? There is no way to tell.
- MoonGhost 1y agoTry content labyrinth. I.e. infinitely generated content with a bunch of references to other generated pages. It may help against simple wget and till bots adapt. PS: I'm on the bots side, but don't mind helping.
- palijer 1y agoThis doesn't work if you pay bandwidth and CPU usage for your servers though.
- MoonGhost 1y agoThat will be your contribution. If others join scrapping will become very pricey. Till bots become smarter. But then they will not download much of generated crap. Which makes it cheaper for you. Anyway, from bots perspective labyrinths aren't the main problem. Internet is being flooded with quality LLM-generated content.
- Twirrim 1y agoThe labyrinth doesn't have to be fast, and things like iocaine (https://iocaine.madhouse-project.org/ https://iocaine.madhouse-project.org/) don't use much CPU if you don't go and give them something like the Complete Works of Ahakespeare as input (Mine is using Moby Dick), and can easily be constrained with cgroups if you're concerned about resource usage. I've noticed that LLM scrapers tend to be incredibly patient. They'll wait for minutes for even small amounts of text.
- cynicalsecurity 1y agoThis topic comes up from time to time and I'm surprised no one yet mentioned the usual fearmongering rhetoric of zip bombs being potentially illegal. I'm not a lawyer, but I'm yet to see a real life court case of a bot owner suing a company or an individual for responding to his malicious request with a zip bomb. The usual spiel goes like this: responding to his malicious request with a malicious response makes you a cybercriminal and allows him (the real cybercriminal) to sue you. Again, except of cheap talk I've never heard of a single court case like this. But I can easily imagine them trying to blackmail someone with such cheap threats. I cannot imagine a big company like Microsoft or Apple using zip bombs, but I fail to see why zip bombs would be considered bad in any way. Anyone with an experience of dealing with malicious bots knows the frustration and the amount of time and money they steal from businesses or individuals.
- os2warpman 1y agoAnyone can sue anyone else for any reason. This is what trips me up: >On my server, I've added a middleware that checks if the current request is malicious or not. There's a lot of trust placed in: >if (ipIsBlackListed() || isMalicious()) { Can someone assigned a previously blacklisted IP or someone who uses a tool to archive the website that mimics a bot be served malware? Is the middleware good enough or "good enough so far"? Close enough to 100% of my internet traffic flows through a VPN. I have been blacklisted by various services upon connecting to a VPN or switching servers on multiple occasions.
- immibis 1y agoYes. A user has to manually unpack a zip bomb, though. They have to open the file and see "uncompressed size: 999999999999999999999999999" and still try to uncompress it, at which point it's their fault when it fills up their drive and fails. So I don't think there's any ethical dilemma there.
- wing-_-nuts 1y agoFor some reason I was under the impression that browsers had the ability to transparently decompress certain archive formats? I may be thinking of less and gzip though
- marcusb 1y agoZip bombs are fun. I discovered a vulnerability in a security product once where it wouldn’t properly scan a file for malware if the file was or contained a zip archive greater than a certain size. The practical effect of this was you could place a zip bomb in an office xml document and this product would pass the ooxml file through even if it contained easily identifiable malware.
- secfirstmd 1y agoEh I got news for ya. The file size problem is still an issue for many big name EDRs.
- marcusb 1y agoUndoubtedly. If you go poking around most any security product (the product I was referring to was not in the EDR space,) you'll see these sorts of issues all over the place.
- j16sdiz 1y agoIt have to be the way it is. Scanning them are resources intensive. The choice are (1) skip scanning them; (2) treat them as malware; (3) scan them and be DoS'ed. (deferring the decision to human iss effectively DoS'ing your IT support team)
- deleted 1y ago[deleted]
- crazygringo 1y ago> For the most part, when they do, I never hear from them again. Why? Well, that's because they crash right after ingesting the file. I would have figured the process/server would restart, and restart with your specific URL since that was the last one not completed. What makes the bots avoid this site in the future? Are they really smart enough to hard-code a rule to check for crashes and avoid those sites in the future?
- fdr 1y agoSeems like an exponential backoff rule would do the job: I'm sure crashes happen for all sorts of reasons, some of which are bugs in the bot, even on non-adversarial input.
- monster_truck 1y agoI do something similar using a script I've cobbled together over the years. Once a year I'll check the 404 logs and add the most popular paths trying to exploit something (ie ancient phpmyadmin vulns) to the shitlist. Requesting 3 of those URLs adds that host to a greylist that only accepts requests to a very limited set of legitimate paths.
- jeroenhd 1y agoThese days, almost all browsers accept zstd and brotli, so these bombs can be even more effective today! [This](https://news.ycombinator.com/item?id=23496794 https://news.ycombinator.com/item?id=23496794) old comment showed an impressive 1.2M:1 compression ratio and [zstd seems to be doing even better](https://github.com/netty/netty/issues/14004 https://github.com/netty/netty/issues/14004). Though, bots may not support modern compression standards. Then again, that may be a good way to block bots: every modern browser supports zstd, so just force that on non-whitelisted browser agents and you automatically confuse scrapers.
- kevin_thibedeau 1y agoIf you nest the gzip inside another gzip it gets even smaller since the blocks of compressed '0' data are themselves low entropy in the first generation gzip. Nested zst reduces the 10G file to 99 bytes.
- galangalalgol 1y agoCan you hand edit to create recursive file structures to make it infinite? I used to use debug in dos to make what appeared to be gigantic floppy discs by editing the fat
- Cloudef 1y agoWouldnt that defeat the attack though as you arent serving the large content anymore
- 1y ago
- fracus 1y agoI'm curious why a 10GB file of all zeroes would compress only to 10MB. I mean theoretically you could compress it to one byte. I suppose the compression happens on a stream of data instead of analyzing the whole, but I'd assume it would still do better than 10MB.
- dagi3d 1y agoI get your point(and have no idea why it isn't compressed more), but is the theoretical value of 1 byte correct? With just one single byte, how does it know how big should the file be after being decompressed?
- kulahan 1y agoIt’s a zip bomb, so does the creator care? I just mean from a practical standpoint - overflows and crashes would be a fine result.
- hxtk 1y agoIn general, this theoretical problem is called the Kolmogorov Complexity of a string: the size of the smallest program that outputs a the input string, for some definition of "program", e.g., an initial input tape for a given universal turing machine. Unfortunately, Kolmogorov Complexity in general is incomputable, because of the halting problem. But a gzip decompressor is not turing-complete, and there are no gzip streams that will expand to infinitely large outputs, so it is theoretically possible to find the pseudo-Kolmogorov-Complexity of a string for a given decompressor program by the following algorithm: Let file.bin be a file containing the input byte sequence. 1. BOUNDS=$(gzip --best -c file.bin | wc -c) 2. LENGTH=1 3. If LENGTH==BOUNDS, run `gzip --best -o test.bin.gz file.bin` and HALT. 4. Generate a file `test.bin.gz` LENGTH bytes long containing all zero bits. 5. Run `gunzip -k test.bin.gz`. 6. If `test.bin` equals `file.bin`, halt. 7. If `test.bin.gz` contains only 1 bits, increment LENGTH and GOTO 3. 8. Replace test.bin.gz with its lexicographic successor by interpreting it as a LENGTH-byte unsigned integer and incrementing it by 1. 9. GOTO 5. test.bin.gz contains your minimal gzip encoding. There are "stronger" compressors for popular compression libraries like zlib that outperform the "best" options available, but none of them are this exhaustive because you can surely see how the problem rapidly becomes intractable. For the purposes of generating an efficient zip bomb, though, it doesn't really matter what the exact contents of the output file are. If your goal is simply to get the best compression ratio, you could enumerate all possible files with that algorithm (up to the bounds established by compressing all zeroes to reach your target decompressed size, which makes a good starting point) and then just check for a decompressed length that meets or exceeds the target size. I think I'll do that. I'll leave it running for a couple days and see if I can generate a neat zip bomb that beats compressing a stream of zeroes. I'm expecting the answer is "no, the search space is far too large."
- manmal 1y ago> Before I tell you how to create a zip bomb, I do have to warn you that you can potentially crash and destroy your own device Surely, the device does crash but it isn’t destroyed?
- cantrecallmypwd 1y agoWouldn't it be cheaper to use Cloudflare than task a human to obsessively watch webserver logs on a box lacking proper filtering?
- gkbrk 1y agoIt's also cheaper to search Google Images for "Eiffel tower" than booking a flight to Paris and going there, but a lot of people enjoy doing the latter.
- charcircuit 1y agoMany people would be better off sticking with the former than realizing what Paris actually is and being disappointed. https://en.wikipedia.org/wiki/Paris_syndrome https://en.wikipedia.org/wiki/Paris_syndrome
- Mashimo 1y agoI had this in mind when visiting Paris and was pleasantly surprised. Lovely and beautiful city. And to heck with cloudflare :S We don't need 3 companies controlling every part of the internet.
- tga_d 1y agoThere was an incident a little while back where some Tor Project anti-censorship infrastructure was run on the same site as a blog post about zip bombs.[0] One of the zip files got crawled by Google, and added to their list of malicious domains, which broke some pretty important parts of Tor's Snowflake tool. Took a couple weeks to get it sorted out.[1] [0] https://www.bamsoftware.com/hacks/zipbomb/ https://www.bamsoftware.com/hacks/zipbomb/ [1] https://www.bamsoftware.com/hacks/zipbomb/#safebrowsing https://www.bamsoftware.com/hacks/zipbomb/#safebrowsing
- vivzkestrel 1y ago"But when I detect that they are either trying to inject malicious attacks, or are probing for a response" how are you detecting this? mind sharing some pseudocode?
- seanhunter 1y agoOnce upon a time around 2001 or so I used to have a static line at home and host some stuff on my home linux box. A windows NT update had meant a lot of them had enabled this optimistic encryption thing where windows boxes would try to connect to a certain port and negotiate an s/wan before doing TCP traffic. I was used to seeing this traffic a lot on my firewall so no big deal. However there was one machine in particular that was really obnoxious. It would try to connect every few seconds and would just not quit. I tried to contact the admin of the box (yeah that’s what people used to do) and got nowhere. Eventually I sent a message saying “hey I see your machine trying to connect every few seconds on port <whatever it is>. I’m just sending a heads up that we’re starting a new service on that port and I want to make sure it doesn’t cause you any problems.” Of course I didn’t hear back. Then I set up a server on that port that basically read from /dev/urandom, set TCP_NODELAY and a few other flags and pushed out random gibberish as fast as possible. I figured the clients of this service might not want their strings of randomness to be null-terminated so I thoughtfully removed any nulls that might otherwise naturally occur. The misconfigured NT box connected, drank 5 seconds or so worth of randomness, then disappeared. Then 5 minutes later, reappeared, connected, took its buffer overflow medicine and disappeared again. And this pattern then continued for a few weeks until the box disappeared from the internet completely. I like to imagine that some admin was just sitting there scratching his head wondering why his NT box kept rebooting.
- mkwarman 1y agoI enjoyed reading this, thank you for sharing. When you say you tried to contact the admin of the box and that this was common back then, how would you typically find the contact info for an arbitrary client's admin?
- geocrasher 1y ago15+ years ago I fought piracy at a company with very well known training materials for a prestigious certification. I'd distribute zip bombs marked as training material filenames. That was fun.
- fareesh 1y agoIs there a list of popular attack vector urls located somewhere? I want to just auto-ban anyone sniffing for .env or ../../../../ etc. Rather not write it myself
- kqr 1y agoIt would be a fairly short Perl script to read the access logs and curl a HEAD request to all URLs accessed, printing only those with 200 OK responses. Here's a start hacked together and tested on my phone: perl -lnE 'if (/GET ([^ ]+)/ and $p=$1) { $s=qx(curl -sI https://BASE_URL/$p | head -n 1); unless ($s =~ /200|302/) { say $p } }'
- vander_elst 1y agoAlso interested in this. For now I've left a server up for a couple of weeks, went through the logs and set up fail2ban for the most common offenders. Once a month or so I keep checking for offenders but the first iteration already blocked many of them.
- deleted 1y ago[deleted]
- BehindTheMath 1y agoCheck out Modsecurity WAF and CoreRuleSet.
- efilife 1y agocheck out the lists in this repo https://github.com/danielmiessler/SecLists/blob/master/Discovery/Web-Content/big.txt https://github.com/danielmiessler/SecLists/blob/master/Disco... I combined a few of the most interesting lists from here into one and never miss an attack now
- eru 1y agoSee https://research.swtch.com/zip https://research.swtch.com/zip for how to make an infinite zip bomb: ie a zip file that unzips to itself, so you can keep unzipping forever without ever hitting bottom.
- guardian5x 1y agoI guess it goes without saying, that the first thing should be to follow security best practices. Patch vulnerabilities fast etc., before doing things like that. Then maybe his first website wouldn't have compromised either.
- Ey7NFZ3P0nzAe 1y agoIf anyone is interested in writing a guide to set this up with crowdsec or fail2ban I'm all ears
- foundzen 1y agoIt is surprising that it works (I haven't tried it). `Content-Length` had one goal - to ensure data integrity by comparing the response size with this header value. I expect http client to deal with this out of the box, whether gzip or not. Is it not the case? If yes, that changes everything, a lot of servers need priority updates.
- Aachen 1y agoYou don't need to set a content length header, it'll take the page as finished when you close the connection
- nottorp 1y agoBut what about the bots written in Rust? Will that get rid of them too?
- dspillett 1y agoRust born processes are memory-safe in terms of avoiding corruption of their heaps & stacks by C-like problems like rogue pointers and use-after-free, but they are still subject to OOM conditions, or running out of other storage, so can easily be killed by a zip-bomb if not coded in an appropriately defensive manner.
- welder 1y agoI like a similar trick, sending very large files hosted on external servers to malicious visitors using proxies. Usually those proxies charge by bandwidth, so it increases their costs.
- JodieBenitez 1y agoThe same, for Caddy: https://www.dustri.org/b/serving-a-gzip-bomb-with-caddy.html https://www.dustri.org/b/serving-a-gzip-bomb-with-caddy.html 10T is probably overkill though.
- deleted 1y ago[deleted]
- deleted 1y ago[deleted]
- b2ccb2 1y agoHilarious because the author, and the OP author, are literally zipping `/dev/null`. While they realize that it "doesn't take disk space nor ram", I feel like the coin didn't drop for them. Think about it: $ dd if=/dev/zero bs=1 count=10M | gzip -9 > 10M.gzip $ ls -sh 10M.gzip 12K 10M.gzip Other than that, why serve gzip anyway? I would not set the Content-Length Header and throttle the connection and set the MIME type to something random, hell just octet-stream, and redirect to '/dev/random'. I don't get the 'zip bomb' concept, all you are doing is compressing zeros. Why not compress '/dev/random'? You'll get a much larger file, and if the bot receives it, it'll have a lot more CPU cycles to churn. Even the OP article states that after creating the '10GB.gzip' that 'The resulting file is 10MB in this case.'. Is it because it sounds big? Here is how you don't waste time with 'zip bombs': $ time dd if=/dev/zero bs=1 count=10M | gzip -9 > 10M.gzip 10485760+0 records in 10485760+0 records out 10485760 bytes (10 MB, 10 MiB) copied, 9.46271 s, 1.1 MB/s real 0m9.467s user 0m2.417s sys 0m14.887s $ ls -sh 10M.gzip 12K 10M.gzip $ time dd if=/dev/random bs=1 count=10M | gzip -9 > 10M.gzip 10485760+0 records in 10485760+0 records out 10485760 bytes (10 MB, 10 MiB) copied, 12.5784 s, 834 kB/s real 0m12.584s user 0m3.190s sys 0m18.021s $ ls -sh 10M.gzip 11M 10M.gzip
- onethumb 1y agoThe whole point is for it to cost less (ie, smaller size) for the sender and cost more (ie, larger size) for the receiver. The compression ratio is the whole point... if you can send something small for next to no $$ which causes the receiver to crash due to RAM, storage, compute, etc constraints, you win.
- PeterStuer 1y ago"On my server, I've added a middleware that checks if the current request is malicious or not" How accurate is that middleware? Obviously there are false negatives as you supplement with other heuristics. What about false positives? Just collateral damage?
- thrwyep 1y agoI thought he maintains his own list of offenders
- PeterStuer 1y agoThe code shows both the 'middleware' and the custom list can put you in the naughty box
- deleted 1y ago[deleted]
- mightyrabbit99 1y agoOP: Hi guys this is how I fend off hackers! Hackers: Note taken.
- gherard5555 1y agoThere is a similar thing for ssh servers, called endlessh (https://github.com/skeeto/endlessh https://github.com/skeeto/endlessh). In the ssh protocol the client must wait for the server to send back a banner when it first connects, but there is no limit for the size of it ! So this program will send an infinite banner very ... very slowly; and make the crawler/script kiddie script hang out indefinitely or just crash.
- dspillett 1y agoAs an aside, there are a lot of people out there standing up massive microservice implementations¹ for relatively small sites/apps, which need to have this part printed, wrapped around a brick, and lobbed at their heads: > A well-optimized, lightweight setup beats expensive infrastructure. With proper caching, a $6/month server can withstand tens of thousands of hits — no need for Kubernetes. ---- [1] Though doing this in order to play/learn/practise is, of course, understandable.
- InDubioProRubio 1y agoIf one wanted to create the ICE of cyberspace in cyberpunk, capable to destroy the device ...
- goodboyjojo 1y agothis was a cool read.very interesting stuff.
- OutOfHere 1y agoServing a zip bomb is pretty illegal. The bot will restart its process anyway, and carry on as if nothing happened.
- VladVladikoff 1y agoIsMalicious() doing some real heavy lifting in that pseudo code. Would love to see a bit more under THAT hood.
- seethishat 1y agoIt's probably watching for connections to files listed in robots.txt that should not be crawled, etc. Once a client tries to do that thing (which it was told not to do), then it gets tagged malicious and fed the zip file.
- jofla_net 1y agoI know ive been on THAT list before. Heaven forbid i dont have chrome or keep it up to date, shame on me!
- foxfired 1y agoLong story short, I use memcached to track ips, user agent, and the use of POST method. The requests per minute, request payload, and past behavior will make isMalicious() return true.
- monus 1y agoThe hard part is the content of isMalicious() function. The bots can crash but they’d be quick to restart anyway.
- jcynix 1y agoAs I don't use PHP in my server, but get a lot of requests for various PHP related stuff, I added a rule to serve a Linux kernel encrypted with a "passphrase" derived from /dev/urandom as a reply for these requests. A zip bomb might be a worse reply ... For all those "eagerly" fishing for content AI bots I ponder if I should set up a Markov chain to generate semi-legible text in the style of the classic https://en.wikipedia.org/wiki/Mark_V._Shaney https://en.wikipedia.org/wiki/Mark_V._Shaney ...
- tonyhart7 1y agook but where I put this?? at the files directory???
- marginalia_nu 1y agoI can't imagine using anything other than a stream interface when dealing with web requests in a crawler. You need that to protect against not only these types of shenanigans, but also large or slow responses.
- geek_at 1y agoThis post is suspiciously similar to my post from 2017 "How to defend your website with ZIP bombs" https://blog.haschek.at/2017/how-to-defend-your-website-with-zip-bombs.html https://blog.haschek.at/2017/how-to-defend-your-website-with...
- speerer 1y agoSame concept, but I found yours more informative. Quite different overall.
- perdomon 1y agoCan someone explain why mods change post titles? What value does it provide in their mind?