Skip to content

Flush the handshake addendum so idle connections survive the server's handshake timeout - #572

Merged
slabko merged 1 commit into
masterfrom
fix/flush-handshake-addendum
Oct 7, 2026
Merged

slabko merged 1 commit into
masterfrom
fix/flush-handshake-addendum

Conversation

@rahulmalik87

Copy link
Copy Markdown
Contributor

Problem

Since ClickHouse added the server setting handshake_timeout_milliseconds (default 30000, ClickHouse commit 8564316756d0, 2026-04-21) the server keeps a timer running from the first Hello byte until it has read the handshake addendum that follows the server Hello (quota key, plus the chunked-protocol and parallel-replicas fields on newer revisions). See TCPHandler::receiveAddendum() and in->clearHandshakeTimeout() right after it.

Client::Impl::Handshake() writes the addendum with WireFormat::WriteString(*output_, std::string()) but does not flush, so it only reaches the server together with the first query. Any connection that is opened and then not used for 30 seconds is closed by the server:

Code: 209. DB::NetException: Timeout exceeded while reading from socket (peer: 127.0.0.1:56310, local: 127.0.0.1:9001, 30000 ms). (SOCKET_TIMEOUT)

The client sees that as a ServerException on its first Execute(), or as an SSL / socket error once the FIN has been processed. The server log for such a connection shows Authenticating user 'default' from ... followed by nothing for 30 seconds, then the SOCKET_TIMEOUT from TCPHandler.

This hits anything that connects early and queries later: connection pools, a client per replica where one replica is queried first, worker threads that connect before a load phase, and so on.

Reproduction

ClickHouse 26.10, clickhouse-cpp 2.6.0 (revision 54459, so the addendum is required):

clickhouse::Client c(opts);
std::this_thread::sleep_for(std::chrono::seconds(35));
c.Execute("SELECT 1");   // SOCKET_TIMEOUT before this change, OK after it

Fix

Flush right after writing the addendum so the handshake is complete on the wire when Handshake() returns. Verified against ClickHouse 26.10 with the snippet above, and with a stress tool (pstress) whose idle worker connections previously died on their first statement.

🤖 Generated with Claude Code

Since ClickHouse added the server setting handshake_timeout_milliseconds
(default 30000, 2026-04-21) the server keeps a timer running from the
first Hello byte until it has read the handshake addendum that follows
the server Hello (quota key, plus the chunked-protocol and
parallel-replicas fields on newer revisions).

Client::Impl::Handshake() wrote the addendum into the output buffer but
did not flush it, so it only reached the server together with the first
query. A connection that was opened and then not used for 30 seconds was
closed by the server:

  Code: 209. DB::NetException: Timeout exceeded while reading from socket
  (peer: ..., 30000 ms). (SOCKET_TIMEOUT)

which the client saw as a ServerException on its first Execute(), or as
an SSL/socket error once the FIN had been processed. The server log for
such a connection shows "Authenticating user ..." followed by nothing for
30 seconds and then the SOCKET_TIMEOUT from TCPHandler.

Flush right after writing the addendum so the handshake completes on the
wire when Handshake() returns. Verified against ClickHouse 26.10: a client
that connects, sleeps 35 seconds and then runs SELECT 1 failed before this
change and succeeds after it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@slabko
slabko merged commit c49fce8 into master Oct 7, 2026
44 checks passed
@rahulmalik87
rahulmalik87 deleted the fix/flush-handshake-addendum branch October 7, 2026 13:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants