11 ms·
How a 20-year-old kernel feature helped USDS improve VA’s network
- ceworthington 9y agoSome people might not have realized USDS is still around since it was best known for the Healthcare.gov rescue under Obama. But it's still here, and still hiring people to work on problems like this www.usds.gov/join
- coleca 9y agoI am glad to see it's still around. I was worried that it would be cut with the change in administrations since it reports into the White House. The USDS is a shining example of civil service and the best of government.
- murtazab 9y agoHas work there changed with the new administration? Wondering how much USDS tech goals change with different administrations.
- Matt_Cutts 9y agoHey there, I'm the USDS Administrator. Most of our projects are the same, like making sure that veterans can get their health benefits. We've found partners in the White House who really want to make government work better for the American people. An example is Chris Liddell, who was CFO for Microsoft. Every administration is going to be different and have different interests, but we've been able to find common ground on projects, and those projects are helping people who need it.
- ceworthington 9y agoOur mission has remained very consistent: use design and technology best practices to improve government services. The new administration has different policy priorities, but it has been remarkable to see the bipartisan support for our mission. Our work has remained largely the same. The major difference is fewer technologists are raising their hands for public service now, which constrains our ability to improve things.
- dragonwriter 9y ago> The major difference is fewer technologists are raising their hands for public service now I suspect you are making a false, or at least unwarranted, generalization from “federal executive branch” to “public” here.
- ceworthington 9y agoGood point! Lots of ways to serve.
- F00Fbug 9y agoHmm.... I apply every 6 months or so and get the thumbs down. Not sure what they're looking for. I've got 30 years of every kind of experience (dev, DBA, network, security, product mgmt, analytics/data science, business mgmt, and more) with good credentials and they never bite. I wish I knew more what the ideal profile was; I'd love to help out!
- steven777400 9y agoI wonder if the environment of experience is significant? USDS positions itself like a startup (even their page has a section on "dress code" which mentions being like "any other startup"). Someone whose experience is primarily enterprise or BigCo might be less appealing. It would be interesting to see a roster of current USDS FTEs and their backgrounds (I didn't see a "Who's Who" on their page, but didn't look extensively).
- noir_lord 9y agoI think that startup mentality might bite them in the arse. I saw "React on Ruby" and winced. There is nothing wrong with that platform as a "We are in a market where things will change radically in two years" but for the VA? Where things might change once a decade, that's a recipe for pain. Look at where the Web was 5 years ago (hell React didn't exist) never mind 10. Angular is 7 years old, KnockoutJS is 7, jQuery is the grandaddy at 11, React is 4. Not a criticism (they are clearly doing important impactful work) more a concern. If someone said to me "You will have to support this for at least 10 years" the choices I made would be extremely conservative.
- mquander 9y agoThere's nothing that would prohibit maintaining a Backbone or Knockout app (or just one written with a bunch of non-spaghetti-code jQuery) today, and it's hard to say that any other choice for writing a piece of software with a GUI would have fared better. Why do you think that using React will have a worse result than that? I think that of the kinds of tools people are using to make web applications in 2017, React and Rails are probably in the more conservative, most likely to be maintainable in 10 years category. (I wouldn't believe this about Rails except it's been so popular for the past 10 years.)
- dkhenry 9y agoIts amazing to see how "solving" the problems can often not solve the problem. Immediately when faced with a error that happened after five minutes I might just put a sleep(301) in the startup script, but that totally would have masked the issue for others. Also amazing foresight by the kernel team to think ahead and make this wrap explicit.
- askldjd 9y agoAuthor here. Completely agreed. My jaws dropped when I saw the INITIAL_JIFFIES. The kernel developers really saved our butt. I could not imagine debugging this problem if INITIAL_JIFFIES was randomized. It may takes days/weeks/months for this bug to appear.
- lfowles 9y agoSimilarly, Unreal Engine 4 offsets platform time (a double) by some large value so if it's stored in a float, accuracy errors will be exposed almost immediately. Looking it up, the offset starts out large enough that the epsilon is two seconds.
- vageli 9y agoDo you have a link with more info? I'd love to read more about this.
- lfowles 9y agoSorry, no. It's not something documented other than a cryptic comment in the source code ( FPlatformTime::Seconds() ) assuming some knowledge of floating point number gotchas. Edit: Here's a more detailed post o made about the specific gotcha if you're interested: https://community.gamedev.tv/t/why-is-fplatformtime-seconds-already-past-6-months/9701 https://community.gamedev.tv/t/why-is-fplatformtime-seconds-...
- aaronchall 9y agoHere's the patch that applies this: https://www.kernel.org/pub/linux/kernel/people/akpm/patches/2.5/2.5.62/2.5.62-mm2/broken-out/initial-jiffies.patch https://www.kernel.org/pub/linux/kernel/people/akpm/patches/...
- jyz 9y agoI had the pleasure of meeting and working with many amazing USDS engineers. Lots of talents, many are truly dedicated to the higher purpose and truly believe in the mission of serving our country. It's a shame that because of the current administration, people are less and less interested in the government.
- askldjd 9y agoThe government's current IT infrastructure crisis is not caused by any one administrations. The root cause goes back decades. Things like "Improving Veterans' lives so they don't have to wait 5-10 years for an Appeals decision" shouldn't be political. I can honestly say that the projects I've been involved with in USDS are the most impactful and meaningful projects I've worked on in my entire life.
- heywire 9y agoI just wanted to say that I love these type of postmortem stories. Thanks for sharing!
- jacquesm 9y agoI love bugs like these. Make it crash is often the hardest part of solving any bug and without this you'd have never known. There is one nasty bit to this story though: the NSOC was running outdated firmware on their Cisco's and wouldn't have known about it if an outside party had not alerted them to this fact. That's pretty sloppy on their end.
- deleted 9y ago[deleted]
- askldjd 9y agoThanks. I think the fact Cisco routers fail to route TCP packets bothers me even more. /you had one job
- fartbagxp 9y agoBut it did route those TCP packets over. For exactly 5 minutes. It just means you need to route everything faster, and then kill your connection, and restart it. :)
- UnoriginalGuy 9y agoI really hope USDS can help introduce interdepartmental digital transfers within the federal government. To give an example of how frustrating it can be... I went through the visa -> green card -> citizenship process. No two departments talk to one another, and when they do they seemingly transmit information on paper which is then transcribed by hand introducing errors/typos. For example USCIS does not talk to the SSA digitally at all. I filled in a single form which was used for both my Visa/Green Card and to apply for a Social Security Card on my behalf, my name was spelt correctly on the visa, but got typo-ed during entry into the SSA's system (then the emphasis placed on me to prove their error, even though other US government departments don't have the typo, including any official ID I hold or naturalisation certificate). Additionally when you earn citizenship the USCIS won't tell anyone. You have to get your piece of paper and physically go tell each department one by one about the change, otherwise nothing will happen. Why doesn't the federal government just have a big database? Or failing that, why does one department not electronically transfer records to another department? Why are people still hand re-entering information already held digitally?
- nacin 9y agoThis sounds incredibly frustrating! Almost all of my USDS projects have involved moving data across agencies, including quite a few with USCIS. It's definitely one of our most common challenges, and USDS is often called into help because we are uniquely positioned to work across departments. As you probably realized when going through the process, USCIS has historically been 98% paper [1] (I've been to the underground limestone cave where they store a lot of it), and we've been working hard to help them modernize the entire agency, including an online application for naturalization [2] and the corresponding backend processing systems. There are a lot of reasons why it's so hard to get agencies to work together, but making it better starts with modernizing individual systems and processes, especially when we're starting with paper that, obviously, can't be transferred seamlessly. For the most part, USCIS doesn't store your entire case file digitally (yet), and even the metadata is stored in a bunch of different systems, which includes (of course) an actual mainframe. (I've seen the mainframe too; it has pretty sweet green LED strips, and not much else going for it.) I'm actually not familiar with how USCIS triggers SSA cards, but I'm going to ask. As it happens, USCIS and SSA do interact electronically in some situations. USCIS runs the E-Verify program, which talks with SSA, as described in this dense but refreshingly public privacy document: https://www.dhs.gov/sites/default/files/publications/privacy_pia_uscis_vis.pdf. https://www.dhs.gov/sites/default/files/publications/privacy.... We've also helped USCIS introduce and improve data exchanges with State, including an early engagement on modernizing the immigrant visa process [3], which you went through, and work on refugee admissions [4]. One of the things I love most about my time at USDS is how many civil servants have embraced and championed best practices for building digital services to best serve the American people. The former director of USCIS, in particular, intuitively understood how technology can improve the immigration process and continued to push us and the agency until their final day in office. As USCIS makes more benefits applications available online, they'll have more data in a readily accessible digital format. They'll be able to streamline the user experience as you progress through the process over the years, and by the end of it, they shouldn't need to ask you a whole lot. It's been great to see human-centered design being championed over and over. Congratulations on becoming a citizen! Want to help us continue to improve the immigration system? We could use the help: https://www.usds.gov/join https://www.usds.gov/join [1] https://medium.com/the-u-s-digital-service/technology-is-helping-to-modernize-our-immigration-system-here-s-how-96162d615b6a https://medium.com/the-u-s-digital-service/technology-is-hel... [2] https://my.uscis.gov/exploremyoptions/us_citizen_through_naturalization https://my.uscis.gov/exploremyoptions/us_citizen_through_nat... [3] https://obamawhitehouse.archives.gov/blog/2015/07/15/bringing-our-immigration-system-digital-age https://obamawhitehouse.archives.gov/blog/2015/07/15/bringin... [4] https://www.usds.gov/report-to-congress/2016/refugee-admissions/ https://www.usds.gov/report-to-congress/2016/refugee-admissi...
- kevin_nisbet 9y agoWhile I found this article very interesting, I feel like something is missing here. So linking this issues to a Cisco bug is very interesting, that dropping connections would cause the application to lock up / crash, while all the connections to the database were dead. My question is why would the application lock up and the servers would crash? I don't see it very often, but when striving for high availability and strong resiliency (which isn't reasonable for everyone), issues need to be looked at in great detail. So I would be trying to look at the second side of the story, which is why was there crashes encountered under these circumstances, and are there other plausible triggers that could cause a similar set of circumstances. Disabling timestamps does avoid the Cisco bug, but a similar set of triggers could be encountered anytime the VPN connection dropped, or if the firewall failed over without the state tables in sync, or any number of other network conditions. And don't take me wrong, I don't know if the OP did this, but based on the article, I would lean towards disabling timestamps as a workaround, and this might still be an indicator that something in the app isn't behaving correctly when the database is unavailable.
- askldjd 9y agoYou are dead on. We do have a bug where we are not recovering the Oracle connectivity correctly. It is on our radar to address the issue. https://github.com/department-of-veterans-affairs/caseflow-monitor/issues/15 https://github.com/department-of-veterans-affairs/caseflow-m... However, There is actually another 50% of the story that I never posted. VACOLS is a really old Oracle DB (from the 80s) that is out of our control. Somehow, it has a "feature" where you can only make one TCP connection to it every 2-3 second. So if we lose connection to the database, it will take many seconds to recover. At that point, our ELB health-check would've fired and restarted our EC2 instances. This is why recoverability of the database connection is not an immediate priority. Here's how we preallocate the VACOLS connection pool to workaround this throttling feature. https://github.com/department-of-veterans-affairs/caseflow/blob/master/config/initializers/warmup_vacols.rb https://github.com/department-of-veterans-affairs/caseflow/b... The infrastructure we operate in are very challenging (and interesting) because of legacy systems. That's why common sense engineering often may not apply in USDS.
- 9y ago
- rlucas 9y agoReminds me of trying to debug long-lived SSH tunnels which would fail every 2:11:15 hours. Right down to the hardcoded value in net.ipv4.tcp
- js2 9y agoI probably would've started at the TCP layer only because I've been bitten at that layer many times and it always has these sorts of strange symptoms. Some examples: 1) Connections hanging over a frame relay network that one day started dropping packets over a certain size. Work-around was adjusting the MTU until I was able to convince the frame relay network operator that something was broken in their network. Initially it was confusing because an interactive telnet session over the network would work fine till you did something like "ls -l" or tried to read a man page which generated enough text to send a full size packet, then the connection would hang. 2) Unable to reach a Verizon e-mail paging gateway but only when connecting from a Linux box. An OS X box on the same network as the Linux box could reach the gateway fine. Turned out Verizon had a firewall rejecting connections where the ECN bit was set. Linux was setting ECN, OS X was not. 3) Solaris box A could initiate a connection to box B, but not the other way. After A talked to B, B could then talk to A, but only for a short period. Someone had deleted A's own MAC from A's ARP table, so A wasn't replying to ARP requests for itself. But if A connected to B, B would keep A's MAC in its own table till it timed out after which B couldn't initiate connection to A any more. 4) All manner of misconfigurations over the years where you learn to recognize the symptoms: misconfigured netmask size; misconfigured duplex; duplicate IP address on same network. You rarely see these any more. 5) The infamous 500-mile e-mail. :-) 6) And my favorite - https://www.pagerduty.com/blog/the-discovery-of-apache-zookeepers-poison-packet/ https://www.pagerduty.com/blog/the-discovery-of-apache-zooke...
- brazzledazzle 9y agoI agree. I'd probably look at it from TCP layer shortly after initial failures to diagnose if not from he start. Especially when dealing with communication between a cloud provider and on-prem gear and infrastructure. However, it's tempting to exhaust all other avenues depending on how likely the on-prem ops folks are to punt the issue.
- askldjd 9y agoI actually did look at the TCP layer early on. However, I didn't pay close attention to the TS Val. From the packet dumps, it just appeared that the TCP window had stopped sliding. I couldn't conclude that NSOC's router was at fault. Getting NSOC on-board is a big deal. After all, they deal with the entire VA network with 100,000+ employees. If you think about it from their perspective, why is USDS' TCP connections so special?
- ChicagoBoy11 9y agoIt's like a 2017 version of the 500 mile email bug https://www.ibiblio.org/harris/500milemail.html https://www.ibiblio.org/harris/500milemail.html
- sydney6 9y agoWould this (disabling TCP Timestamps) affect TCP Performance with other OSes in regard of their respective TCP Window Auto Scaling Implementations? I believe Linux uses DRS (1) and doesn't necessarily depend on TCP Option TS for TCP Window Auto Scaling and FreeBSD has got this (2) commit ~ 2 Months ago. (1) http://public.lanl.gov/radiant/pubs.html#DRS http://public.lanl.gov/radiant/pubs.html#DRS (2) https://svnweb.freebsd.org/base?view=revision&revision=316676 https://svnweb.freebsd.org/base?view=revision&revision=31667...
- askldjd 9y agoYup, it would. Disabling the TS option was just a stopgap measure to make our deployments stable for the time being.
- sydney6 9y agoOf course, setting priorities.. I was just wondering how different OSes would behave under these circumstances. For instance AWS S3 also doesn't support TCP Timestamps and this had a rather big impact on e.g. FreeBSDs TCP Performance until recently.
- firebones 9y agoThe Jiffies root cause leads to an interesting idea: an Glossary of Magic Constants where all kinds of important constants, limits, and overflows are tracked to aid in debugging. You could imagine a search engine where "tcp connection drops after 5 minutes" lists every piece of software and firmware with 5 minute and 300 second constants.
- gwern 9y agoSo OEIS for programming? Largely seems covered by Google: punch in an oddly specific number and someone will probably have discussed it on Stack Exchange.
- equalunique 9y agoA team at Veterans Affairs is my customer. They have been tasked to integrate with some AWS hosted intranet system. Other than points of integration, we know very little about it. This article seems to be a big clue.