Skip to content

[BUG] GRPC txt lookup cancel latency #4962

Description

@2S1one

Describe the issue

Calling hatchet.runs.cancel(run_id) (bulk_cancel, which POSTs to /api/v1/stable/tenants/<tenant>/tasks/cancel) takes a stable, reproducible ~8.02 seconds to respond in a Docker Compose deployment — even when run_id doesn't correspond to any real run. A plain GET /api/v1/stable/workflow-runs/{id} on the same server returns in ~10ms. Root cause: grpc-go's DNS resolver performs an extra _grpc_config.<hostname> TXT lookup by default when Hatchet's internal gRPC client dials hatchet-engine by hostname; Docker's embedded DNS resolver (127.0.0.11) doesn't answer that (nonexistent) TXT query promptly, and grpc-go retries it (~4.0s per attempt) before giving up and proceeding.

Environment

  • SDK: Python hatchet-sdk v1.40.1
  • Engine: Self-hosted, ghcr.io/hatchet-dev/hatchet/hatchet-dashboard:latest @ sha256:d2ec11aa288a64ce5078054ca4fdcf52bd604c48f65398d4d9a293c2b7c73c15 (hatchet-api binary: github.com/hatchet-dev/hatchet/cmd/hatchet-api, built with google.golang.org/grpc v1.83.2); ghcr.io/hatchet-dev/hatchet/hatchet-engine:latest @ sha256:f0661c2b0245359be1bbb8f36b67a8f1387d5d3e30cfea717db5e7ad5d3d804e. Reproduced identically with both SERVER_MSGQUEUE_KIND=postgres and RabbitMQ — the message-queue backend is not a factor.

Expected behavior

hatchet.runs.cancel() (and any other call that proxies to the engine) should return in roughly the same time as other REST calls (~10-20ms locally), not ~8 seconds, regardless of whether Docker's embedded DNS resolver answers a _grpc_config.* TXT query quickly.

Code to Reproduce, Logs, or Screenshots

from app.hatchet_client import hatchet

hatchet.runs.cancel("00000000-0000-0000-0000-000000000000")  # nonexistent run
# -> takes ~8.02s to return, every time

Or via raw HTTP, from inside the container itself (i.e. not a Docker host<->container networking artifact):

docker exec <hatchet-dashboard container> sh -c '
  time curl -s -o /dev/null -w "%{http_code}\n" -X POST \
    -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
    -d "{\"externalIds\":[\"00000000-0000-0000-0000-000000000000\"]}" \
    http://localhost:8080/api/v1/stable/tenants/<tenant>/tasks/cancel
'
# real  0m8.01s

strace -f -tt -T attached to the hatchet-api process during one cancel call (relevant excerpt, request-handling thread):

21:07:19.525233 write(12, "...\16hatchet-engine\0\0\34..." , 43)   # A/AAAA query for "hatchet-engine" — resolves instantly
21:07:19.525726 write(12, "...\f_grpc_config\16hatche...", ...)    # TXT query for "_grpc_config.hatchet-engine"
21:07:23.526735 read(12, ...)                                       # response after ~4.001s
21:07:23.526946 write(12, "...\f_grpc_config\16hatche...", 56)      # TXT query retried
21:07:27.527809 read(12, ...)                                       # response after another ~4.001s
21:07:27.528089 write(12, "..._grpc_config.hatchet-engine.lxd...")  # third query (search-domain suffix), instant this time
21:07:27.528865 connect(12, {port 7070, <engine IP>})                # only now dials the real gRPC port

Confirmed separately inside the container: nslookup -type=TXT _grpc_config.hatchet-engine against 127.0.0.11 (Docker's embedded resolver) times out after ~5s, while getent hosts hatchet-engine (plain A-record lookup) resolves instantly.

Additional context

_grpc_config.<hostname> is grpc-go's own convention for fetching DNS-based service config via TXT record — see internal/resolver/dns/dns_resolver.go. This is a known pain point in Docker environments — see grpc/grpc-go#3572 "DNS resolution does not work on Docker".

The gRPC client Hatchet dials hatchet-engine with (pkg/client/v1/grpc-client.go, used from api/v1/server/handlers/v1/proxy/proxy.go's InternalClientFactory.NewGRPCClient(...)) is constructed with grpc.NewClient(opts.hostPort, grpcOpts...), and grpcOpts does not include grpc.WithDisableServiceConfig(). This also explains why the slowdown is specific to tasks/cancel (and presumably other calls that proxy through InternalClientFactory to freshly dial the engine by hostname) rather than e.g. GET workflow-runs, which only touches an already-pooled Postgres connection and never re-dials the engine.

Confirmed workaround: grpc-go v1.75+ (this build: v1.83.2) supports disabling the TXT lookup via the environment variable GRPC_ENABLE_TXT_SERVICE_CONFIG=false. Setting this on the hatchet-dashboard container's environment and restarting it:

Before After
POST /tasks/cancel (raw) ~8.02s (5+ trials) ~0.011–0.017s (3+ trials)
Full SDK path: external cancel() to CancelledError raised in the running task ~9s ~1s (an unrelated, intentional 1s SDK-side grace period)

Verified end-to-end with a real running task afterwards: subprocess SIGTERM'd cleanly, run status settled as CANCELLED, no behavioral regression observed.

Suggested fix: preferred — add grpc.WithDisableServiceConfig() to the dial options used when constructing Hatchet's internal gRPC client (pkg/client/v1/grpc-client.go) — this config has no possible upside in a Docker Compose / Kubernetes deployment (there is never going to be a real _grpc_config.<service> TXT record for an internal service name), and a very real downside (multi-second latency, specifically worse under Docker's embedded resolver). Alternative / minimal — default GRPC_ENABLE_TXT_SERVICE_CONFIG=false in the official docker-compose.yml quickstart's hatchet-dashboard (and any other service making this internal gRPC call) environment, and/or document it as a recommended setting for self-hosted Docker deployments.


🤖 AI Disclosure
  • I acknowledge that an LLM was used in the creation of this Issue, in accordance with Hatchet's AI_POLICY.md.
  • Details: Claude (Anthropic, Claude Code) was used throughout: writing the reproduction script, attaching strace to the live hatchet-api process and interpreting the syscall trace, verifying the grpc-go TXT-lookup behavior via web search against grpc/grpc-go source/issues, applying and re-measuring the GRPC_ENABLE_TXT_SERVICE_CONFIG=false fix, and drafting this report. A second LLM session (OpenAI Codex) was used to independently re-derive the timing from the raw strace log and locate the exact Hatchet/grpc-go source lines cited above. All timings and log excerpts above were actually executed/captured during the session, not fabricated by the model.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions