Skip to content

[fix][client] Don't close a shared DNS resolver group's address resolver when a client closes - #26723

Merged
lhotari merged 1 commit into
apache:masterfrom
lhotari:lh-fix-shared-dns-resolver-close
Sep 28, 2026
Merged

lhotari merged 1 commit into
apache:masterfrom
lhotari:lh-fix-shared-dns-resolver-close

Conversation

@lhotari

@lhotari lhotari commented Sep 26, 2026

Copy link
Copy Markdown
Member

Motivation

A DnsResolverGroupImpl wraps Netty's DnsAddressResolverGroup, whose getResolver(executor) hands the same resolver to every caller on the same event loop and keeps it until the group is closed. PulsarClientImpl.close() nevertheless closed the AddressResolver it got from the group, also when the group was shared and not owned by the client:

  • clients built with PulsarClientSharedResources (the DnsResolver shared resource), and
  • the broker's own clients: PulsarService.createClientImpl gives every broker-side client the shared brokerClientSharedDnsResolverGroup and the broker's ioEventLoopGroup.

Closing one such client closes the resolver's UDP channel for every other client on that event loop. Because the closed resolver stays in the group's map, clients created later on that event loop get the closed resolver too. Their lookups keep working only while the name is in the resolver's DNS cache. After that, each lookup fails at once with:

java.net.UnknownHostException: Failed to resolve '<host>' [A(1)]
Caused by: io.netty.resolver.dns.DnsNameResolverException: [/<dns-server>:53] DefaultDnsQuestion(<host>. IN A) failed to send a query '...' via UDP (no stack trace available)
Caused by: io.netty.channel.StacklessClosedChannelException

In the broker, BrokerService.closeAndRemoveReplicationClient, which runs when a cluster is deleted, closes that cluster's replication client. From then on, until the broker restarts, the broker clients on that client's event loop can't resolve uncached hostnames. These include the broker's internal client, the replication clients of the other clusters and the namespace clients. For example, reconnecting to a remote cluster fails once the cache entry has expired. Applications that share PulsarClientSharedResources across clients and close some of them are affected in the same way.

The shared DNS resolver group was added in #24784, so this affects 4.1.2 and later 4.1.x releases, and 4.2.0 and later.

Modifications

  • PulsarClientImpl.close() closes the address resolver only when the client created the resolver group itself (dnsResolverGroupLocalInstance). A shared group's resolvers are closed when the group is closed, by PulsarClientSharedResources.close() or PulsarService.close().

Verifying this change

  • Make sure that the change passes the CI checks.

This change added tests and can be verified as follows:

  • PulsarClientImplTest.testClosingClientLeavesSharedDnsResolverOpenForOtherClients builds two clients from PulsarClientSharedResources on a single event loop, with a DNS server that never answers. It closes one client and checks that the other client's lookup still sends a query to the DNS server. Without the fix the test fails, because the closed resolver never sends the query.

Does this pull request potentially affect one of the following parts:

If the box was checked, please highlight the changes

  • Dependencies (add or upgrade a dependency)
  • The public API
  • The schema
  • The default values of configurations
  • The threading model
  • The binary protocol
  • The REST endpoints
  • The admin CLI options
  • The metrics
  • Anything that affects deployment

…er group when a client closes

A DnsResolverGroupImpl hands the same Netty resolver to every client that runs on the same event loop, and keeps
it until the group closes. PulsarClientImpl.close() closed that resolver also when the group was shared, through
PulsarClientSharedResources or the broker's brokerClientSharedDnsResolverGroup, so every other client on that
event loop, and every client created on it later, failed its DNS lookups with a closed channel.

The client now closes the address resolver only when it created the resolver group itself.

Assisted-by: Claude Code (claude-opus-5-5)
@lhotari lhotari added this to the 5.0.0 milestone Sep 26, 2026
@lhotari lhotari added release/4.2.5 release/4.0.14 triage/lhotari/important lhotari's triaging label for important issues or PRs labels Sep 26, 2026
@lhotari
lhotari merged commit 986853d into apache:master Sep 28, 2026
44 checks passed
ascentstream-bot pushed a commit to ascentstream/pulsar that referenced this pull request Oct 1, 2026
…ver when a client closes (apache#26723)

(cherry picked from commit 986853d)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

release/4.0.14 release/4.2.5 triage/lhotari/important lhotari's triaging label for important issues or PRs

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants