Repository navigation
Flush the handshake addendum so idle connections survive the server's handshake timeout - #572
Merged
Merged
Conversation
Since ClickHouse added the server setting handshake_timeout_milliseconds (default 30000, 2026-04-21) the server keeps a timer running from the first Hello byte until it has read the handshake addendum that follows the server Hello (quota key, plus the chunked-protocol and parallel-replicas fields on newer revisions). Client::Impl::Handshake() wrote the addendum into the output buffer but did not flush it, so it only reached the server together with the first query. A connection that was opened and then not used for 30 seconds was closed by the server: Code: 209. DB::NetException: Timeout exceeded while reading from socket (peer: ..., 30000 ms). (SOCKET_TIMEOUT) which the client saw as a ServerException on its first Execute(), or as an SSL/socket error once the FIN had been processed. The server log for such a connection shows "Authenticating user ..." followed by nothing for 30 seconds and then the SOCKET_TIMEOUT from TCPHandler. Flush right after writing the addendum so the handshake completes on the wire when Handshake() returns. Verified against ClickHouse 26.10: a client that connects, sleeps 35 seconds and then runs SELECT 1 failed before this change and succeeds after it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
slabko
approved these changes
Oct 7, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Since ClickHouse added the server setting
handshake_timeout_milliseconds(default30000, ClickHouse commit 8564316756d0, 2026-04-21) the server keeps a timer running from the first Hello byte until it has read the handshake addendum that follows the server Hello (quota key, plus the chunked-protocol and parallel-replicas fields on newer revisions). SeeTCPHandler::receiveAddendum()andin->clearHandshakeTimeout()right after it.Client::Impl::Handshake()writes the addendum withWireFormat::WriteString(*output_, std::string())but does not flush, so it only reaches the server together with the first query. Any connection that is opened and then not used for 30 seconds is closed by the server:The client sees that as a
ServerExceptionon its firstExecute(), or as an SSL / socket error once the FIN has been processed. The server log for such a connection showsAuthenticating user 'default' from ...followed by nothing for 30 seconds, then theSOCKET_TIMEOUTfromTCPHandler.This hits anything that connects early and queries later: connection pools, a client per replica where one replica is queried first, worker threads that connect before a load phase, and so on.
Reproduction
ClickHouse 26.10, clickhouse-cpp 2.6.0 (revision 54459, so the addendum is required):
Fix
Flush right after writing the addendum so the handshake is complete on the wire when
Handshake()returns. Verified against ClickHouse 26.10 with the snippet above, and with a stress tool (pstress) whose idle worker connections previously died on their first statement.🤖 Generated with Claude Code