Summary
With the Npgsql 10 default GssEncryptionMode=Prefer, opening a connection to a server that accepts TCP but never
answers sometimes blocks far past Timeout. The Timeout fires during GSS-encryption negotiation, Npgsql logs that it
is retrying without GSS encryption and opens a second socket — and in roughly 2% of attempts that second attempt runs
with no effective timeout, so Open() blocks until the peer closes the connection. GssEncryptionMode=Disable removes it
entirely.
This looks like the same family as #6524 (the fallback path after the GSS attempt not honouring the timeout), but with a
different trigger: Kerberos libraries are present here (macOS), and the server is simply unresponsive.
Environment
- Npgsql 10.0.3 (via Npgsql.EntityFrameworkCore.PostgreSQL 10.0.3), .NET 10, macOS (arm64), Kerberos tooling present (
klist)
- Connection string:
Host=127.0.0.1;Port=<fake>;Timeout=1;Pooling=false (+ database/user)
Reproduction
A TCP listener that accepts connections and never writes a byte. Loop single NpgsqlConnection.Open() calls against it,
sequentially, each with a watchdog of a few seconds:
| Setting |
Opens |
Blocked past the watchdog |
default (GssEncryptionMode=Prefer) |
800 |
16 (2.0%) — each unblocked only when the listener closed the socket |
GssEncryptionMode=Disable |
800 |
0 |
Every non-blocked attempt failed promptly as expected (NpgsqlException wrapping TimeoutException), and latency on the
timeout path was unchanged by the setting. SslMode made no difference (Prefer 4/400 vs Disable 5/400 blocked with GSS
still at its default) — consistent with GSS negotiation happening before SSL.
What the trace shows
With Npgsql's trace logging enabled (a capturing ILoggerProvider via NpgsqlLoggingConfiguration.InitializeLogging),
trials run sequentially so every line in a window belongs to one trial:
- "Attempting to connect" → "Socket connected"
- "Negotiating GSS encryption" → the Timeout fires
- the fallback message ("… retrying without it")
- a second "Attempting to connect" → "Socket connected"
- then nothing until the listener closes, then an
IOException
In the hung trials step 4's socket never gets a timeout applied; in the other ~98% the fallback attempt times out normally.
Ruled out (each by measurement)
- exception classification (all prompt failures were the timeout)
- CPU load (timeout held at ~21 ms under load average 44)
- thread-pool starvation (0/10 with every worker parked and the pool capped)
- cold start (the hang occurred at iteration 3, not 0)
SslMode (see above)
Impact / workaround
For us it hung an end-of-run cleanup pass that must give up on an unresponsive server. Workaround: set
GssEncryptionMode=Disable on connections that must never block past Timeout.
Happy to share the reproduction harness.
Summary
With the Npgsql 10 default
GssEncryptionMode=Prefer, opening a connection to a server that accepts TCP but neveranswers sometimes blocks far past
Timeout. TheTimeoutfires during GSS-encryption negotiation, Npgsql logs that itis retrying without GSS encryption and opens a second socket — and in roughly 2% of attempts that second attempt runs
with no effective timeout, so
Open()blocks until the peer closes the connection.GssEncryptionMode=Disableremoves itentirely.
This looks like the same family as #6524 (the fallback path after the GSS attempt not honouring the timeout), but with a
different trigger: Kerberos libraries are present here (macOS), and the server is simply unresponsive.
Environment
klist)Host=127.0.0.1;Port=<fake>;Timeout=1;Pooling=false(+ database/user)Reproduction
A TCP listener that accepts connections and never writes a byte. Loop single
NpgsqlConnection.Open()calls against it,sequentially, each with a watchdog of a few seconds:
GssEncryptionMode=Prefer)GssEncryptionMode=DisableEvery non-blocked attempt failed promptly as expected (
NpgsqlExceptionwrappingTimeoutException), and latency on thetimeout path was unchanged by the setting.
SslModemade no difference (Prefer 4/400 vs Disable 5/400 blocked with GSSstill at its default) — consistent with GSS negotiation happening before SSL.
What the trace shows
With Npgsql's trace logging enabled (a capturing
ILoggerProviderviaNpgsqlLoggingConfiguration.InitializeLogging),trials run sequentially so every line in a window belongs to one trial:
IOExceptionIn the hung trials step 4's socket never gets a timeout applied; in the other ~98% the fallback attempt times out normally.
Ruled out (each by measurement)
SslMode(see above)Impact / workaround
For us it hung an end-of-run cleanup pass that must give up on an unresponsive server. Workaround: set
GssEncryptionMode=Disableon connections that must never block pastTimeout.Happy to share the reproduction harness.