Winsock tcp / ip Socket listening but connection refused, race condition?

This includes two automated unit tests, each starting a tcp / ip server, which creates a non-blocking socket, and then bind () s and listen () s in a select () loop for the client, which connects and loads some data.

The catch is that they work fine when run separately, but when run as a test suite, the second test client won't be able to connect to WSACONNREFUSED ...

UNLESS

there is Thread.Sleep () in between them from a few seconds !!!!!!

Interestingly, every 1 second the retry loop connects after any failure. So the second test loop for a while before the timeout expires after 10 minutes.

During this time, netstat -na indicates that the correct port number is LISTEN for the server socket. So what if he is listening? Why is he not accepting the connection?

There are log messages in the code that show that the NEVER selection even gets a read-ready socket (meaning ready to accept a connection when applied to a listening socket).

Obviously, the problem must be with some kind of race condition between the end of one test, which means close () and shutdown () at each end of the socket, and the start of the next.

It wouldn't be so bad if the retry logic allowed it to connect eventually after a couple of seconds. However, it looks like he is "hypnotized" and won't even try again.

However, for some strange reason, the listening socket is SPEAKING it up in LISTEN state, even if it continues to refuse connections.

So this means that Windoze O / S actually captures the SYN packet and returns an RST packet (which means "Connection refused").

The only time I saw this error was when there was a problem in the code that caused hundreds of sockets to get stuck in the TIME_WAIT state. But this is not the case. netstat only shows about a dozen sockets with 1 or 2 in TIME_WAIT at any given moment.

Please, help.

+2


a source to share


3 answers


The main problem was closing the socket, the stream was trying to read any remaining bytes. This was done as a separate thread that keeps the read end of the socket open for a fixed amount of time in milliseconds while repeatedly re-reading any data.

This logic has been replaced by a smarter reading of any data and closing properly when the read returns 0. So it closes much faster.



So it turned out to be a wrong socket closing in my own code.

Thanks for the help!

+2


a source


I run many tests like this on build machines with different Windows operating systems (XP through Windows 7) with different number of cores and I have never seen this to be a problem.

I don't believe that switching the listener to TIME_WAIT

is likely to be your problem; I've never seen it, and I regularly run client server tests on the same port where I start and stop the servers during the latency period TIME_WAIT

.

If you've started the second server before your first one closes its socket (or if the socket was in TIME_WAIT

), I expect your second server to get an error when trying bind()

.).

Personally, I think it's more likely that the problem is in the code you have that accepts connections - your test may have found a bug;)

Can we see the code between your listener and the accept loop?



Do you have a problem if you change the order of the tests?

Are the clients and the server running on the same computer, do they change things if they are not?

Etc.

I have TCP testing tools http://www.lenholgate.com/blog/2005/11/windows-tcpip-server-performance.html if you configured your test system to run test client from this link versus server example from this http://www.lenholgate.com/blog/2005/11/simple-echo-servers.html are you still seeing your problem? (That is, start my server with my client on your test system so that it runs the same way it runs your stuff, and does my stuff work?).

+2


a source


From This MSDN site :

The TIME_WAIT state determines the amount of time that must elapse before TCP can release a closed connection and reuse its resources. This interval between closure and release is known as the TIME_WAIT state or the 2MSL state. During this time, the connection can be reopened for the client and server at a much lower cost than creating a new connection. The TIME_WAIT behavior is specified in RFC 793, which requires TCP to maintain a closed connection for an interval that is at least twice the maximum segment lifetime (MSL) of the network. When a connection is released, its socket pair and the internal resources used for the socket can be used to support another connection.

Windows TCP reverts to TIME_WAIT state after closing the connection. If in TIME_WAIT state, the socket pair cannot be reused. The TIME_WAIT period is configurable by modifying the following DWORD registry entry, which represents the TIME_WAIT period in seconds.

HKEY_LOCAL_MACHINE\System\CurrentControlSet\Services\TCPIP\Parameters\TcpTimedWaitDelay

      

By default, MSL is defined as 120 seconds. In the TcpTimedWaitDelay registry value, the default value is 240 seconds, which is 2 times the maximum segment life of 120 seconds or 4 minutes. However, you can use this entry to adjust the interval. Decreasing the value of this entry allows TCP to release closed connections faster, providing more resources for new connections. However, if the value is too low, TCP can free up connection resources before the connection ends, requiring the server to use additional resources to re-establish the connection. This registry value can be set between 0 and 300 seconds.

I think the minimum you can set is 30 (try decreasing, but may not work)

A more detailed explanation can be found in the Winsock Programmer FAQ .

+1


a source







All Articles