4 ms·
This is an interesting case, but I'm confused about some of the details. > A little known fact is that it's not possible to have any packet loss or congestion
by dap 10y ago
This is an interesting case, but I'm confused about some of the details.
> A little known fact is that it's not possible to have any packet loss or congestion on the loopback interface.
This seems a bit misleading, given the two counterexamples that the article describes after this.
> If you think about it - why exactly the socket can automatically expire the FIN_WAIT state, but can't move off from CLOSE_WAIT after some grace time. This is very confusing... And it should be!
On illumos, the FIN_WAIT_2 -> TIME_WAIT transition happens only after 60 seconds if the application has closed the socket file descriptor. In that case, by definition the application has no handle with which to perform operations on the socket. The resource belongs exclusively to the kernel. If the other system disappeared forever, and there were no timeout, that socket would be held open forever.
By comparison, in CLOSE_WAIT, the application still has a handle on the socket, and it's responsible for the resource. The application can even keep sending more data in this case (as part of a graceful termination of a higher-level protocol). Or it could enable keep-alive. It's able to respond to the case where the other system has gone away, and it could break the application if the kernel closed the socket on its behalf.
I think the behavior is non-obvious, but pretty reasonable.
- majke 10y ago>> A little known fact is that it's not possible to have any packet loss or congestion on the loopback interface. > This seems a bit misleading, given the two counterexamples that the article describes after this. Ok, allow me to be more precise. The loopback _link_, the virtual cable, the virtual interface thingy has zero packet loss. If you imagine a wire that represents loopback, it will never congest, never have interference, never have any loss. The packets will _always_ reach the other end. Now, what happens after they reach the other end, that's a separate story. The article indeed showed that it's possible that the kernel network stack dropped packets because of full buffers, or the CLOSE-WAIT thing. But it's the kernel network stack on the receiving end dropping packets actively, not the loopback "wire". > ...and it could break the application if the kernel closed the socket on its behalf. The article showed a case that the other peer just quit alltogether. What kernel thinks about CLOSE-WAIT socket is irrelevant. The connection is _dead_. You can send data over it, you can attempt to read, you can try to close it gracefully, no difference. The other end refused to cooperate. I argue that keeping the socket in CLOSE-WAIT for more than, say, 60 seconds is stupid, since the other party _will_ give up. Basically, any application has 60 seconds to flush the data and call the close() after receiving FIN. After that, it all gets weird. So why not enforce this 60 seconds and just move the socket off CLOSE_WAIT to some next stage? Now, what precisely file descriptor represents is another story. Maybe you're right, maybe a file descriptor can't point to a socket in TIME-WAIT or LAST-ACK states.
- toast0 10y ago> I argue that keeping the socket in CLOSE-WAIT for more than, say, 60 seconds is stupid, since the other party _will_ give up. Although it's rare[1], it's legal and possible for peer A to send a FIN, and peer B to continue to send data on the half closed connection and for peer A to expect the data and continue to receive it. How is the kernel expected to understand the difference between that behavior and a socket leak? In this case, peer A fully closed the application socket, so if peer B sends data on the half-closed socket, it would get a RST from peer A, if peer A was still mildly cooperating or the data would simply not be acked if peer A left the network. [1] It's probably rare because people don't understand it well, and additionally some of the people who don't understand it create middle-boxes that break these types of connections in exciting ways.
- dap 10y ago> The article showed a case that the other peer just quit alltogether. What kernel thinks about CLOSE-WAIT socket is irrelevant. The connection is _dead_. You can send data over it, you can attempt to read, you can try to close it gracefully, no difference. The other end refused to cooperate. Remember that it's possible for an application to send FIN without closing the socket. In the article's case, the connection was dead, but from the TCP stack's perspective, it has no way to tell that case from the case where the connection is still alive on the other end. While weird, and I don't think it's a good idea, it's possible for host A to send FIN over a connection to host B, causing host B's socket to enter CLOSE_WAIT, but to have the connections exist in this state for an extended period (many minutes) while host B continues to send data to host A. A totally plausible interaction is an HTTP client that connects to the server, sends the request headers, sends FIN (causing the server to wind up in CLOSE_WAIT), and then the server spends several minutes sending the requested data. Even if the socket were idle, the connection may be perfectly healthy. Plus, even if the kernel cleaned up the socket, the application has still leaked a file descriptor (also a limited resource). > Basically, any application has 60 seconds to flush the data and call the close() after receiving FIN. After that, it all gets weird. So why not enforce this 60 seconds and just move the socket off CLOSE_WAIT to some next stage? For the reasons I mentioned, as long as the socket is open in the application on host B, I don't think the kernel can conclude that the program is done with it. That's okay, because this isn't all that hard to handle: if host A has actually closed its end of the socket (rather than just sending FIN), then as the post describes host A's socket will typically close about 60 seconds after receiving the ACK for its FIN. In that case, if host B continues sending data, it will eventually get an RST and socket operations will fail with ECONNRESET. No connection will be leaked. As I understand it, the problem only happens when host B stops sending data (because the socket is leaked), in which case it will never realize that host A is gone. This is not very different than the case where the socket is still ESTABLISHED and host A has panicked, power-cycled, or failed in some other non-graceful way. A likely solution to both of these would be to use TCP keep-alive or application-level keep-alive, which will cause the same ECONNRESET behavior, and no connection will be leaked. I agree that this case is very non-obvious, and I ran into a very similar problem recently that was painful to debug. But I don't think it's safe for the kernel to attempt to solve this, there would still be a leak even if it did, and the same mechanism that the application can use to identify failed connections in the ESTABLISHED state can likely be used to identify this case as well.
- kbenson 10y ago> This seems a bit misleading, given the two counterexamples that the article describes after this. If you are referring to the "Assuming the target application has some space in its buffers, packet loss over loopback is not possible." caveat, then yes, it is somewhat ambiguous what's really being implied. Maybe it's just network packet loss they are referring to in the initial statement? I'm not sure that really makes sense either. > On illumos, the FIN_WAIT_2 -> TIME_WAIT transition happens only after 60 seconds if the application has closed the socket file descriptor. ... If the other system disappeared forever, and there were no timeout, that socket would be held open forever. Isn't that exactly what the automatic closing is supposed to prevent? Couldn't deliberate or accidental connection interruptions eventually cause a DOS?
- dap 10y ago> Isn't that exactly what the automatic closing is supposed to prevent? Couldn't deliberate or accidental connection interruptions eventually cause a DOS? Yes, exactly. My point was that until the application closes the file descriptor, it's up to the application to deal with that issue. (Many public-facing web servers close idle sockets for this reason.) Only once the application has closed the socket does it become the TCP stack's responsibility, and that's why the 60-second timer exists. But that timer exists only for sockets whose file descriptors are closed.
- ambrop7 10y ago> If the other system disappeared forever, and there were no timeout, that socket would be held open forever. There is actually a more "gracefully degrading" solution, to keep those connections around potentially indefinitely or with a very long timeout, but to recycle them if the memory is needed for new connections. Seems like Linux supports this for TIME-WAIT but is not enabled by default (sys.net.ipv4.tcp_tw_recycle). Actually lwIP does such recycling and it can even kill active connections - if there's no memory for a new connection it will try to kill connections in order: TIME-WAIT, LAST-ACK, CLOSING, active connections (incl. FIN-WAIT-2).
- takeda 10y agoI would highly recommend to not use that setting it solves one problem and could introduce other, harder to debug problems. Especially if a load balancer or NAT is on the way. With this setting Linux no longer follows RFC and there is a risk that it could confuse packets from different connections. It falls back to tcp timestamp from PAWS, but often people disable timestamp as well.