21 ms·
A Google Cloud support engineer solves a tough DNS case
- itsmemattchung 6y ago> they use raw sockets! Raw sockets are different than normal sockets: they bypass iptables, and they are not buffered! Can someone elaborate on the above statement from the article? Does this imply that raw sockets have unbounded buffer?
- jeffbee 6y agoClearly not, but raw is raw. For the purposes of this article, it's enough to say that raw sockets, being raw, don't traverse the flawed block of code, which is in net/ipv4/udp.c
- foota 6y agoProbably the opposite, they send without waiting for any bytes to build up.
- user5994461 6y agoIn short, raw sockets can push bytes into the network card. Whereas the commonly used socket functions recv/send construct the required headers for TCP, UDP and whatnot, they handle encapsulation/buffering/connection/etc so they're easy to use for developers, just read and write application data. By nature raw sockets skip TCP/UDP libraries and a good chunk of the network kernel code. Including the place where the bug was located.
- rantwasp 6y agothat’s a pretty good and detailed explanation. 2 things: 1) i hope they have a runbook for situations like this (ie the support engineer does not have to figure all this on the fly) 2) the customer should have provides more details and maybe should have thought of the tweaks they made (classic solution is to compare 2 instances - one works one does not)
- donalhunt 6y agoYou can't runbook these sort of issues... You can ensure you have the knowledge and tools to root cause the problem.
- rantwasp 6y agoyou most definitely can and you should. there are steps that you can do to gather the info and the linked example shows basic things to try. when you exhaust the run-book is when you start digging
- donalhunt 6y agoI agree with the point around information gathering. Good technical teams have guides for endusers to collect relevant information. In this case the engineering team may have shared information with the TSCs to help narrow down the root cause. This also highlights why support agents shouldn't be 100 engaged with customers. They need time to review and amend runbooks, consult with engineering teams, etc.
- amessina1 6y ago1) you cannot have a runbook for everything, and even if you have a runbook you the best you could have found in this case is that something weird was happening in the VM. The setting had an insanely big value but it was accepted by the kernel, so you would assume it was a valid one. 2) the customer provided a huge amount of details, but it is usually very hard to explain what did you change from the base image. Most customers might not be willing to provide the full spec of their running system, as they might contain information they don't want to disclose. It is easier when the issue starts right after a change has been made, but this was not the case.
- rantwasp 6y agoi’m not saying a runbook will catch everything. but it will give you a chance to solve the problem quicker.
- rmac 6y ago" I find that 2147481343 seems to do the trick. This number doesn't make any sense to me. I suggest the customer try this number. The customer replies back: it works with google.com, but it does not work with other domains." My extrapolation: I find a potential fix: don't test it, don't understand it, and send it to customer. How is this acceptable?
- Diederich 6y agoDid you read the next sentence? "I continue my investigation."
- SpicyLemonZest 6y agoTechnically-inclined customers generally appreciate being actively involved in the debugging process, rather than the alternative of being stonewalled until a solution everyone's 100% confident in can be found.
- floatingatoll 6y agoYour extrapolation is incorrect. The following description matches the factual components of what you see as unacceptable, but includes the missing context from the post that makes the engineer's actions acceptable rather than unacceptable: It was sent to the customer to collect experimental results while investigation continued in parallel. Per the post, production mitigations had already been put into place, and the customer was knowingly participating in an investigation of an unknown issue. The customer provided the requested experimental results.
- helipad 6y agoI don't think they were pitching it as a solution, instead a "try this and see if it works", which helps continue the debugging. It can also help because it shows the customer you're making progress.
- floatingatoll 6y agoThe LKML message described in the post is here: https://lkml.org/lkml/2019/12/19/482 https://lkml.org/lkml/2019/12/19/482
- mav3rick 6y agoSomething I'd like to add here the actual fix is - + if (rmem > (size + (unsigned int)sk->sk_rcvbuf)) However in reality this would have worked too - + if (rmem > (unsigned int)(size + sk->sk_rcvbuf)) (The bit pattern of the result remains the same and it's still casted as unsigned int during the comparison) However, signed integer overflow is undefined behavior in C and unsigned integer overflow isn't. Hence, the submitted patch is the correct solution
- floatingatoll 6y agoThose are not safe to treat as equivalent, even if it might work in theory. You should always cast as narrowly as possible, and when you see code doing otherwise, look very carefully for bugs. If A + (cast)B is a correct form, then (cast)(A + B) is generally an inappropriate form. As you note, it’s possible it will happen to work, but it’s not good form.
- mav3rick 6y agoYes I qualified it at the end why the former is the correct solution and not the latter.
- floatingatoll 6y agoI figured that "in reality this would have worked too" is a easy phrase to misinterpret as "this would have worked" (assuming they miss the context a paragraph later), and so the reply helps ensure that others do not misread it as I initially did.
- mav3rick 6y ago
- jldugger 6y ago> After another spin around the world the case comes back to our team. So basically, follow the sun doesn't work for hard problems? Can you really say the people are working on this 24/7 if progress is only made in one time zone?
- donalhunt 6y agoI suspect the author was in the same timezone as the customer. Not all customers will have follow the sun coverage. The customer had also mitigated the problem for the time being. It does show that embracing good change management techniques is important. The VM in question clearly had runtime configs updated manually or other VMs would have exhibited the same problem. Ideally orgs should be able to identify when the change was introduced and the rationale / tests that were used.
- dtech 6y agoBy the time his team had it, numerous things had been investigated and tried, and the customer already had a workaround. Presumably these things all happened in the other time zones. By the time it stopped transferring, it was already diagnosed as a hard problem and the time pressure was off.
- amessina1 6y agoFollow the sun has a lot of overhead: imagine having to dump the engineers' thoughts and current hypothesis and load them in the brain of the next oncallers. Also it only really works if also the customer is active 24/7, as often you might need the customer to perform some action on their systems. Once the time pressure is off you might get better results dedicating the engineer who is best suited to work on the case (both from the point of view of the timezone and skill set) and give them the time needed to troubleshoot the issue.
- xvector 6y agoIs Google Cloud support this good in general, or only for certain tiers or plans?
- jcroll 6y agoThe only support for basic accounts is for billing
- dilyevsky 6y agoEven on platinum plans they are terrible in my experience. Endless pinballing of tickets and trying to blame issue on the customer. We had two day outage several times in my previous gig because google support refused to acknowledge the problems
- bpodgursky 6y agoThe front-line support is the same as anywhere else. But Google Cloud has really really good second and third line support, if the first tier can't figure it out. And in many cases, it'll get escalated directly to the implementing engineers. In my experience, Google Cloud is better than most organizations about escalating hard issues up to the chain. Admittedly, this happened at a company with substantial spend, and I can't say one way or another whether a smaller player would get the same quality of support.
- dilyevsky 6y agoHow much is “substantial” for gcp to get decent support? Tens of millions/y certainly wasn’t cutting it...
- bpodgursky 6y agoIf you weren't getting substantial support at tens of millions/y, your finance team did something deeply wrong when negotiating the contract.
- 6y ago
- econcon 6y agoGoogle support is basically non existent for customers even spending like 5-10K a month on their platform. Even questions you raise on their Reddit sub go unnoticed. They don't have bunch of people actively trying to solve their customers problem. That said the only time my problems were actually listed to were from BigQuery team. Other than this, I don't think it's possible to get any explanation on a feature from any other product team at Google cloud.
- doh 6y agoWe spent orders of magnitude of that and they still treated us poorly. The silver lining is that they at least don't discriminate based on the spend.
- Axsuul 6y agoDo you have a support plan?
- ellard 6y agoSupport plans make sense for products/software that are sold and expected to operate independently (cars as always are a good example). Heck, even those products generally come with some sort of free limited time support which come in the form of warranties. Services requiring a separate support plan is... not intuitive.
- kelnos 6y agoWhy should you need a support plan for a product you're paying for? "Ok, you can pay us $X/mo for the service, but if something goes wrong, we won't help you unless you also pay an additional $Y/mo." It's absolute garbage that this is where the industry is.
- TwoBit 6y agoThe customer had set an extremely large buffer size and nobody thought to mention that? Perhaps the individual reporting the problem was different and was unaware of that unusual change.
- atombender 6y agoTo be fair, there are a lot of sysctl settings. To be sure, it's one of the first places I would look for networking weirdness, but it's also often hard to tell what impact those settings have on anything.
- bluedino 6y agoAnyone that messes with them, better know what they are doing. 99% of the time they read about or heard about the setting somewhere and are blindly following along. I see this a lot with "performance improvements" when they should be looking other places like their web server configuration. Like why would you tweak those on a low end wordpress server?!
- kevin_nisbet 6y agoI spend a decent amount of time investigating trouble reports, and in my experience it's quite uncommon to even get as much information as was provided in what google showed. It's also fairly uncommon to get any of these sorts of rare configurations in trouble reports, and usually takes some probing. When I was a new engineer working in telco, one of the longest investigations I worked was when connectivity broke between one of our regional roaming partners and 1/3 of our nodes (I'm summarizing to try and keep the story brief). We called them and asked if they changed anything, reviewed the configuration and secrets used on the tunnels, etc. And were working with the vendor to go through any problems with the implementation. Saturday morning and probably 20 hours of investigation later, a new engineer at the regional partner see's there is a work order for changes to the connectivity to our nodes (we were adding some new ones) that was supposed to be executed that week. A typo in the change overwrote the secrets used by an existing tunnel instead of creating a new secret for the new peer. The person we were working with to investigate, was the person who implemented that change and told us several times nothing changed. He was also the one we worked with and read through all the secrets for typos or issues and didn't notice anything. Saturday morning he get's into the office, is shown the work order, and goes, oh yea, I did that at exactly the time the tunnel went down. Fix of typo'd secret later and everything comes right back up. So just in my experience, I find it quite plausible that buffer size was not mentioned. And even besides this story, I know I've personally missed connecting causes with potential effects when investigating a problem, it's very easy to dismiss some setting, like the buffer size, as being connected specifically to DNS behaviours, especially if they are not noticed together or with a strong change management system that helps connect the timelines together.
- luhn 6y agoThis is a fun debugging story, but is a great example why servers should be cattle not pets. Having trouble with a VM? Blow it up and get a fresh one. Still having trouble? The provisioning steps are codified, you can walk through them and find the one that causes the issue.
- daenz 6y agoThe end result was a kernel patch to LKML, so I for one am happy they solved this problem at the source.
- atombender 6y agoIf a freshly provisioned VM doesn't have the same issue, then they're not using automated configuration management (Puppet, Chef, or similar) for these settings, and then they have a more serious problem, as nothing in their runtime environment is "codified" or predictable.
- throwaway2048 6y agoplenty of issues are not deterministic, even with 100% of everything managed by configuration management software.
- atombender 6y agoBut /etc/sysctl.conf is deterministic.
- jeffbee 6y agoWhat is a reason why that file would correspond to the actual sysctls in effect?
- atombender 6y agoIf you use an automated configuration management system such as Puppet, you don't ever run sysctl manually in a shell. Instead, everything is controlled by the configuration management system. sysctl is a bit problematic in terms of exhaustiveness. That is, how do you ensure that the kernel only has its original values plus whatever you put in sysctl.conf, and nobody actually ran sysctl manually at some point? But it's possible to do.
- ser_tyrion 6y ago> they use raw sockets! Raw sockets are different than normal sockets: they bypass iptables But this bugreport says raw sockets would be filtered by the OUTPUT chain of iptables: https://bugzilla.redhat.com/show_bug.cgi?id=1269914#c4 https://bugzilla.redhat.com/show_bug.cgi?id=1269914#c4 Is that accurate across distros? It does make sense for some socket types, like device sockets, to not be routed through iptables.
- barbegal 6y agoI think that bug report is misleading. Raw sockets do bypass iptables but they still go through ebtables. They hook in at the ebtables NAT OUTPUT chain. See the diagram here https://erlerobotics.gitbooks.io/erle-robotics-introduction-to-linux-networking/security/introduction_to_iptables.html https://erlerobotics.gitbooks.io/erle-robotics-introduction-...
- ser_tyrion 6y agoThanks for that great link
- schoolornot 6y agoI've been supporting AWS environments almost from the beginning but can't ever remember a case where I was asked or even considered offering Support a copy of a VM's storage volume. Is this common on Google/Azure/etc.?
- CSDude 6y agoIt reminds me of my friend's hosting company that failed. They got a big customer and created a VM and the customer asked to fix a problem that involved getting a shell in the VM. Friend does it and the customer is gone next day. This aside, even though we had too many support cases so far with AWS, and having highest support level, they mostly cannot access user data, just the metadata. We had a major problem with RDS once, and they specifically requested to load that snapshot to an internal instance to reproduce. It can happen in AWS, but not very common, in my experience.
- DaiPlusPlus 6y ago> It reminds me of my friend's hosting company that failed. They got a big customer and created a VM and the customer asked to fix a problem that involved getting a shell in the VM. Friend does it and the customer is gone next day. The customer asked the support people to access a shell on their VM and they then quit because...? - or the support people accessed a shell without the customer’s express permission?
- kelnos 6y agoI think it's because the support people even had the possibility of that access at all. And I agree; support people shouldn't have that kind of access, with or without consent.
- gvjddbnvdrbv 6y agoIt is a fairly common method to test to see if they have access.
- tallanvor 6y ago
- jtchang 6y agoif (rmem > (size + sk->sk_rcvbuf)) goto uncharge_drop; What is rmem in this case? I'm a bit confused as to why it is written that way. This drops the packet right when it overflows the buffer?
- jeffbee 6y agoIt's not very literate, is it? rmem is initially the sk_backlog.rmem_alloc field of struct sock. There is no comment in net/sock.h what this field might mean. People who modify this function just have to guess. I also appreciate that this function adds |size| to rmem_alloc, tests for limits, then later it subtracts |truesize| from rmem_alloc. This happens to seem correct, but it's just asking for someone to accidentally screw up the accounting in a later change. Reading this function only reinforces my view of Linux code quality.
- jeffbee 6y agoAnother interesting thing is how stewardship of this logic and the comment right above it have diverged. Original: /* we drop only if the receive buf is full and the receive * queue contains some other skb */ rmem = atomic_add_return(size, &sk->sk_rmem_alloc); if ((rmem > sk->sk_rcvbuf) && (rmem > size)) goto uncharge_drop; Later: /* we drop only if the receive buf is full and the receive * queue contains some other skb */ rmem = atomic_add_return(size, &sk->sk_rmem_alloc); if (rmem > (size + sk->sk_rcvbuf)) goto uncharge_drop; How does the comment correspond to each block?
- crdrost 6y agoBasically rmem - size was the number of bytes that were consumed before this current packet. In the line before, we have thread-atomically added size to rmem immediately prior and read the current value in one single step. Call this "staking our claim" to part of the buffer, and the meaning of "uncharge" here is discharging this claim by atomically decrementing the counter, before dropping the packet. Probably this bug would not have happened if this comparison were written as `rmem - size > sk->sk_rcvbuf`? So it is saying that the simplest sanity check is "if the buffer was already full before we staked our claim we should drop this packet immediately." As the "goto" indicates, there are then a bunch more checks on other circumstances where we should also drop the packet. Due to the quirks of multi-threading it is of course possible that some packets get unnecessarily dropped between when we stake the claim and when we discharge it, which the code just accepts -- the thinking is presumably "yeah if the buffer is full a lot of packets are gonna get dropped and that's just life -- it's much less important that we dropped some extra packets when we were already dropping packets, and much more important that we don't mismanage the buffer's memory when it's nearly full." A comment suggests that part of the reason for this awkward phrasing is that it is possible for rmem = size, in other words the buffer was empty when we staked our claim--and in this case we don't want to drop this packet even if it would overflow a small max-buffer-size. I think the idea there is "we already have the socket buffer allocated, obviously this thing fits in memory, so let's just handle it if the queue is empty rather than dropping every single packet that is larger than the queue size."
- hidiegomariani 6y agoamazing. I wish I knew where to go learn the basics to navigate that many layers of knowledge (kernel/os/network)..
- auspex 6y agoThese books should help! Computer Networks and Internet Internet working with TCP/IP Volume III The Linux Programming interface
- mcguire 6y agoAnd... Computer Architecture: A Quantitative Approach Operating System Concepts Computer Networking: A Top-Down Approach W. Richard Stevens (Somewhat out of date, but I haven't seen anything to beat them.) Doesn't look horrible: http://intronetworks.cs.luc.edu/ http://intronetworks.cs.luc.edu/
- srnvs123 6y agoEvery time I run into bugs, and feel like I'm doing everything right, but it's just behaving unreasonably for some reason, but then find the issue and feel stupid. This one is one of those
- joshsyn 6y agoLong-term solution - Use rust
- hi41 6y agoCan you expand on your answer. How does Rust resolve this issue? Does it solve by disallowing type casts thereby preventing overflow when the number becomes negative?
- steveklabnik 6y agoI think this person is trolling. However, for the purposes of discussion, two things here: So, in Rust, overflow panics in debug builds, but does wrap around in release builds. So, it is possible this bug would have been caught in testing, but if it wasn't, it still would have slipped into production. However, that being said, Rust does not do implicit casting between numeric types. So it's very likely that this code would not have compiled in the first place, though I haven't examined it super closely. At that time, the person would have had to cast it, and so the end result would have been roughly the same.
- cesarb 6y agoAFAIK, Rust can panic on overflow even in release builds if you want it to, at a somewhat heavy performance cost (which is why this is not enabled by default in release builds). In this case, it would convert the issue from "some packets are unexpectedly being discarded" into an immediate crash within the kernel.
- steveklabnik 6y agoYes, there is a flag to change the default behavior. I am not sure how many folks actually use this flag.
- MrMorden 6y agoIf you want flags, gcc has -Wall -Wextra -Werror which seems like it would have caught this bug. (Of course, if you weren't using -Wall -Wextra from the beginning you'll have a lot of catching up to do before you can build with -Werror.)
- pbhjpbhj 6y agoSeems like a good place to mention this: I once was troubleshooting an Outlook issue where email stopped working after some time, seemingly at random. Turned out that Outlook picked up the IP address for the mail server backwards - so instead of WW.XX.YY.ZZ Outlook tried ZZ.YY.XX.WW. Found that by using sysinternals network tools, confirmed with Wireshark. Thunderbird worked, ping to the domain name (of the mail server) worked, Outlook worked on-and-off, ... As it was MS Windows and I'm a Linux native I didn't really know how to investigate further - I guess I couldn't without Outlook source. Luckily setting a hosts entry fixed it. I only found one other post online with the same issue, and they didn't have a solution. Presumably it was something like ISP automated rDNS entries getting parsed .. but honestly I don't know. Still curious ... Would have loved to have found work investigating such things, as they said in the post, it's fun!
- maccam94 6y agoIt sounds like maybe your mail server had a misconfigured reverse DNS entry? Those DNS PTR records look like ZZ.YY.XX.WW.in-addr.arpa. I'm aware that SMTP in particular has a dependency on rDNS, but I'm not sure about the details. https://en.wikipedia.org/wiki/Reverse_DNS_lookup#Uses https://en.wikipedia.org/wiki/Reverse_DNS_lookup#Uses
- erikig 6y ago“...This means that the case will Follow the Sun by default, to provide 24/7 support” I love the concept of “Follow the Sun” to describe 24/7 support - I don’t think I’ve heard it described that way. I wonder how much we’d have to spend to get that tier of service?
- amessina1 6y agoFor Google Cloud, 250$ per month per user: https://cloud.google.com/support https://cloud.google.com/support
- erikig 6y agoThank you.
- CobrastanJorji 6y ago"Follow the Sun" is subtly different from 24/7. 24/7 can mean follow the sun, but it can also mean "we are prepared to page someone at 2 AM and wake them up." Follow the sun means "there is an engineer in China, India, France, Boston, and San Francisco, and at least one of them is always at their desk and ready to take work." The difference for users can be fairly small, but as someone who used to carry a pager for Amazon, the difference is really huge for the support person.
- drudru11 6y agoI know that this phrase was used way back in the early 90s. I suspect it was used before then as well.
- yalogin 6y agoThat brings up another point. Should the kernel standardize on unsigned scalars completely? How many legitimate use cases are there to use signed scalars in the kernel?
- jeffbee 6y agoUsing unsigned does not generally fix overflow flaws. It just moves the threshold.
- yalogin 6y agoSure. I was not suggesting it that it will eliminate overflows but would eliminate one source of them. Also mostly because there are probably few use cases that warrant signed values.
- saagarjha 6y agoUsing unsigned numbers doesn't really fix anything here, because the Linux kernel defines overflow of signed numbers. In both cases, you have generally surprising behavior when the number gets large enough: changing the type doesn't help; it just hides the issue in one of the cases.
- cesarb 6y agoThe Linux kernel internally uses the common C convention in which negative numbers are errors (-ENOMEM and so on) while positive numbers (and zero) are successful values. (Some parts of the kernel use a related convention in which values within the "last page" of the address space, that is, -4096 to -1 inclusive, are errors, while other values are valid memory addresses; there are macros to convert between both conventions.)
- bluedino 6y agoI'd like to hear Linus's thoughts on this particular issue.
- supernova87a 6y agoI'm not an expert with AWS or Google Cloud, so I'm interested in knowing: What "level" of customer or SLA do you have to be to get a certain quantity or guarantee of support and troubleshooting? Or is it that if even a free-tier customer points out something that is fundamentally a problem, it will receive attention by certain solutions engineers? Are there $ spending, 20 x (c3.4x.large), or I-pay-you-for-certain-uptime/troubleshooting levels that get you certain response levels? Do certain problems get resolved with "well, you just have to live with that behavior, we're not fixing that". Do you get to call them or chat live? Or is it all via tickets?
- andyljones 6y agoHere you go: https://cloud.google.com/support#support-plans https://cloud.google.com/support#support-plans $250/month/dev is the minimal for phone calls on technical issues, $150k + 4% of GCP spend for 'come running' support. There're more details here https://cloud.google.com/support/docs/procedures#additional_services_for_gold_platinum_production_and_enterprise_support_customers https://cloud.google.com/support/docs/procedures#additional_... though they use the old names for the support tiers.
- fenesiistvan 6y agoWhat "dev" or "user" means here ...for example if I run a server serving some HTTP API.
- londons_explore 6y agoI believe it means "person who wants the ability to contact cloud support". For small to medium sized businesses, that number is probably 1.
- ehsankia 6y agoIt seems like the blog post talks about a written case report, which the 100$ tier has access to, albeit with 4 hour first response instead of 1 hour. So it is possible that you could get your case escalated to such an in depth debugging with that tier?
- peterwwillis 6y agoNote: You shouldn't use int, unsigned int, char, short, long. Use int16_t, uint16_t, uint8_t, etc (or their _fast equivalents) from stdint.h. The former's sizes change based on platform, cpu, and compiler; the latter are fixed-width (or flexible, where _fast may use a larger size if it's faster). I started brushing up on my C recently and have been collecting these little nuggets: https://gist.github.com/peterwwillis/53cd9d34d8755784e48379047af9f358 https://gist.github.com/peterwwillis/53cd9d34d8755784e483790...
- kccqzy 6y agoNot when you are writing the Linux kernel, when you know exactly which sizes the integers are. Or even when you are writing low-level code on a known platform (e.g. an LP64 platform).
- peterwwillis 6y agoAll the typical sizes depend on the compiler and the flags you provide. If you provide the wrong flags, the sizes of each type may change, but you won't see any warnings about it because it's expected behavior, and now you've got different binaries with different behavior. Or you could use fixed-width types and if the wrong flags get passed (no C99 support) your code just doesn't compile. I just checked the Linux kernel style guide, and they explicitly suggest you can use fixed-width types from C99 when it makes sense. I get that they want a balance for their project, but for general programming, it's just safer to be explicit. https://www.kernel.org/doc/html/v5.1/process/coding-style.html#typedefs https://www.kernel.org/doc/html/v5.1/process/coding-style.ht...
- deleted 6y ago[deleted]
- tobykimmel 6y agoThe big news here is Google support engineer solves ANY problem. I can never get through to them. That’s the downside to a fully automated support system, no humans.
- bauerd 6y agoThere are lots of support engineers around, you just need to pay Google to use up their time.
- tobykimmel 6y agoThat’s not true for all services. Google Voice for example has become unusable, and this is confirmed by thousands of support forum posts. There’s no way to reach a human. All the support forum posts are eventually closed with no resolution. Google should just shut down Google Voice instead of keeping it around for free, but broken and with no support.
- andred14 6y agoReally enjoyed this :) As our systems grow in size and complexity we will inevitably encounter limits (and resulting problems) that previously were not approached. For the future (interstellar space travel etc.) all this will need to be recreated for greater scale
- mcguire 6y ago"When sk_rcvbuf gets close to 2^31, adding the size of the packet can cause an integer overflow. And since it’s an int it becomes a negative number, therefore the condition is true when it should be false (for more, also check out this discussion of signed magnitude representation)." And this is why you don't generally use signed numbers in systems code, unless you specifically need negative numbers. And why you gradually develop a paranoia about the sizes of numbers.
- FreeFull 6y agoI'm not sure how using an unsigned number would help, given that when it overflows you're still going to have some code do unexpected stuff anyway.
- mcguire 6y agoFor one thing, it's a clue you need to step back and think, "What happens when this overflows?" rather than "Oh, it's just a number." For another, that's why you get paranoid. (For a third, I strongly recommend something like Frama-C with the Weakest-Precondition module---it's very good at finding issues like these.)
- sleepydog 6y agoI'm not convinced that it would have been any more obvious to the person who made the error that the variable could overflow if it were unsigned. It's also much easier, IMO, to accidentally underflow an unsigned integer; it's so much more common to work with 0 than it is to work with +/- 2 billion.
- karagenit 6y agoI'm being pedantic, but technically wrapping from 0 to MAX_INT is still considered overflow. Underflow refers to decimal truncation e.g. by integer division.
- koheripbal 6y ago
- pistolpeteDK 6y agodidn’t read the story - but a tough case deserves an upvote!
- slim 6y agoHe forgot mention that the kernel is Linux. It's almost like linux has become the standard OS. Linux is the new windows.
- ryanmarsh 6y agoOh look another C overflow bug.
- saagarjha 6y agoA slightly different case, as the Linux kernel uses a nonstandard (-fwrapv) C where overflow is defined to wrap rather than be undefined.
- truthwhisperer 6y agowhy can't you reach Google just over the phone. Support is very bad
- deleted 6y ago[deleted]
- biohax2015 6y agoI envy people who find things like this fun.
- amrx101 6y agoOh boy. It was like a thriller. Enjoyed it :)
- xyproto 6y agoGreat troubleshooting. TLDR; int overflow in C
- happppy 6y ago#boycottGoogle
- rishabhd 6y agoReminds me of this classic haiku - It's not DNS. There's no way it's DNS. It was DNS.
- ninj4fly 6y agoNice reading. Just ignore the metadata server thing. The guy probably means a resolver.
- HugoDaniel 6y agoThat is why google only hires the best
- jiveturkey 6y agolol. google hires vast swathes of the unwashed
- bogomipz 6y agoI enjoyed reading this but wouldn't have running either "netstat -s" or "ss -s" to show protocol statistic have shown either receive buffer errors/receive packet errors statistics? It seems like this basic tool was noticeably absent from the early troubleshooting steps and other standard troubleshooting tools used. I understand the importance of the ultimate fix but wouldn't seeing an incrementing error counter for UDP have shortened some of the troubleshooting done to identify and resolve the customer's immediate issue?
- jiveturkey 6y agoeasily caught by SAST. not even that. standard compiler warning should catch this.