3 ms·
> The article showed a case that the other peer just quit alltogether. What kernel thinks about CLOSE-WAIT socket is irrelevant. The connection is _dead_. You c
by dap 10y ago
> The article showed a case that the other peer just quit alltogether. What kernel thinks about CLOSE-WAIT socket is irrelevant. The connection is _dead_. You can send data over it, you can attempt to read, you can try to close it gracefully, no difference. The other end refused to cooperate.
Remember that it's possible for an application to send FIN without closing the socket. In the article's case, the connection was dead, but from the TCP stack's perspective, it has no way to tell that case from the case where the connection is still alive on the other end.
While weird, and I don't think it's a good idea, it's possible for host A to send FIN over a connection to host B, causing host B's socket to enter CLOSE_WAIT, but to have the connections exist in this state for an extended period (many minutes) while host B continues to send data to host A. A totally plausible interaction is an HTTP client that connects to the server, sends the request headers, sends FIN (causing the server to wind up in CLOSE_WAIT), and then the server spends several minutes sending the requested data. Even if the socket were idle, the connection may be perfectly healthy.
Plus, even if the kernel cleaned up the socket, the application has still leaked a file descriptor (also a limited resource).
> Basically, any application has 60 seconds to flush the data and call the close() after receiving FIN. After that, it all gets weird. So why not enforce this 60 seconds and just move the socket off CLOSE_WAIT to some next stage?
For the reasons I mentioned, as long as the socket is open in the application on host B, I don't think the kernel can conclude that the program is done with it.
That's okay, because this isn't all that hard to handle: if host A has actually closed its end of the socket (rather than just sending FIN), then as the post describes host A's socket will typically close about 60 seconds after receiving the ACK for its FIN. In that case, if host B continues sending data, it will eventually get an RST and socket operations will fail with ECONNRESET. No connection will be leaked.
As I understand it, the problem only happens when host B stops sending data (because the socket is leaked), in which case it will never realize that host A is gone. This is not very different than the case where the socket is still ESTABLISHED and host A has panicked, power-cycled, or failed in some other non-graceful way. A likely solution to both of these would be to use TCP keep-alive or application-level keep-alive, which will cause the same ECONNRESET behavior, and no connection will be leaked.
I agree that this case is very non-obvious, and I ran into a very similar problem recently that was painful to debug. But I don't think it's safe for the kernel to attempt to solve this, there would still be a leak even if it did, and the same mechanism that the application can use to identify failed connections in the ESTABLISHED state can likely be used to identify this case as well.
- pedasmith 10y agoThis sounds like a half-closed socket. I helped with WinRT socket API for Windows; we had to decide whether to support that or not. In the end, we couldn't come up with a truly legitimate case where an app really needs to do this. Even with no realistic use cases and no efficiency gains, we still got complaints from a small number of developers who wanted the flexibility of closing a socket one way but not the other.
- wahern 10y agoHuh? No realistic use cases? A half-closed socket is how you emulate EOF. It's like saying that EOF has no realistic use case; that _every_ protocol should implement a custom channeling and signaling mechanism at the application layer in order to get bidirectional streams, even though they're using TCP sockets. Now, many of the most popular protocols do that, but they do that because they're intended to be universally deployed and must deal with all the random pathological cases out in the wild. But if you don't have to worry about broken software, such as when you know there won't be broken proxies or routers in your path, then you can dramatically simplify many purpose-built protocols. As a practical matter most people can safely make that assumption. Most of the time it's cheaper and easier (in the grand scheme of things) to make the broken software shoulder the burden. Not everybody is writing software for Cloudflare or Chrome; and fortunately _most_ responsible vendors do a decent job at implementing standards correctly. I hope there's some confusion on your part or my part, and that you just didn't admit that WinRT was completely broken by design. That should be a hanging offense for anybody designing or implementing core infrastructure software. Decisions like that impose millions, if not billions of dollars in unnecessary costs on the world.
- majke 10y ago> For the reasons I mentioned, as long as the socket is open in the application on host B, I don't think the kernel can conclude that the program is done with it. Wow. Awesome explanation. Ok, so basically as long as socket is in CLOSE-WAIT the server may want to send() something. Do you happen to know what are the detailed semantics of `tcp_fin_timeout` sysctl? Will FIN-WAIT-2 progress to cleanup after 60 seconds, or after 60 seconds since last received packet? Would this be sane: - Let's allow kernel to expire CLOSE-WAIT if the application didn't send anything for more than 60 seconds.