3 ms·
The customer had set an extremely large buffer size and nobody thought to mention that? Perhaps the individual reporting the problem was different and was unawa
by TwoBit 6y ago
The customer had set an extremely large buffer size and nobody thought to mention that? Perhaps the individual reporting the problem was different and was unaware of that unusual change.
- atombender 6y agoTo be fair, there are a lot of sysctl settings. To be sure, it's one of the first places I would look for networking weirdness, but it's also often hard to tell what impact those settings have on anything.
- bluedino 6y agoAnyone that messes with them, better know what they are doing. 99% of the time they read about or heard about the setting somewhere and are blindly following along. I see this a lot with "performance improvements" when they should be looking other places like their web server configuration. Like why would you tweak those on a low end wordpress server?!
- kevin_nisbet 6y agoI spend a decent amount of time investigating trouble reports, and in my experience it's quite uncommon to even get as much information as was provided in what google showed. It's also fairly uncommon to get any of these sorts of rare configurations in trouble reports, and usually takes some probing. When I was a new engineer working in telco, one of the longest investigations I worked was when connectivity broke between one of our regional roaming partners and 1/3 of our nodes (I'm summarizing to try and keep the story brief). We called them and asked if they changed anything, reviewed the configuration and secrets used on the tunnels, etc. And were working with the vendor to go through any problems with the implementation. Saturday morning and probably 20 hours of investigation later, a new engineer at the regional partner see's there is a work order for changes to the connectivity to our nodes (we were adding some new ones) that was supposed to be executed that week. A typo in the change overwrote the secrets used by an existing tunnel instead of creating a new secret for the new peer. The person we were working with to investigate, was the person who implemented that change and told us several times nothing changed. He was also the one we worked with and read through all the secrets for typos or issues and didn't notice anything. Saturday morning he get's into the office, is shown the work order, and goes, oh yea, I did that at exactly the time the tunnel went down. Fix of typo'd secret later and everything comes right back up. So just in my experience, I find it quite plausible that buffer size was not mentioned. And even besides this story, I know I've personally missed connecting causes with potential effects when investigating a problem, it's very easy to dismiss some setting, like the buffer size, as being connected specifically to DNS behaviours, especially if they are not noticed together or with a strong change management system that helps connect the timelines together.
- GABeech 6y agoThere are many things here that concern me from a system view. 1. They need guaranteed delivery, but chose to use UDP 2. They jacked up the default rmem buffer to ~2GB which is insane. Also, applies to all sockets not just UDP, so I wouldn't be surprised if they where also running into issues with memory pressure especially under load 3. Support didn't seem to let them know that's a pretty unconventional configuration That was an interesting debugging story, and catching a bug like this is always good IMO. But, there is just so much WTF in this setup.
- bluedino 6y agoI would have told the customer to reduce the number by 1000 or whatever and closed the case.
- livueta 6y agoI used to do work kind of like this stuff on enterprise storage arrays that ran a modified BSD. We didn't lock the system down much, so customers could go and set whatever system guts stuff they wanted. We had a tool on-system that would basically tar up all the system configs and phone home with them when the customer hit a button. You'd better believe that one of my first steps investigating anything was diffing the crap out of any relevant configs against a clean base version. As another commenter mentioned, this was the result of customers never actually mentioning their weird sysctl tuning in the original issue description. It's not like they're trying to screw you over or anything - there's just an awful lot of config options in an entire system that does anything interesting, and in the case of a big enterprise appliance, it's likely that dozens of people have had admin on it at one point or another.