Skip to content

[self-care:hosted-health] 2026-10-07 04:05 UTC degraded #16624

Description

@mnkiefer

cao.githubnext.com backend (cao-dashboard, cao-collector) is operating normally at the application layer: zero error-status spans, stable p95 request latency (~1.5ms server-side), stable goroutine counts, and no restarts across the last 24h. However, the public HTTP health-check endpoints (/api/readiness, /api/health, /api/v1/health) all return HTTP 404, and the CAO MCP tools/list call also returns HTTP 404 despite correct default-branch authorization — a genuine endpoint-reachability gap that external monitors relying on those routes would miss.

Action: Maintainer should confirm whether the readiness/health/v1-health routes and the MCP tools/list route are intentionally retired or misconfigured in the current deployment, then redeploy or update the health-check path configuration; acceptance check is all three HTTP checks and the MCP tools/list call returning 2xx on the next hosted-health run.

Evidence

Check Outcome Detail
HTTP /api/readiness ❌ fail 404, 167ms
HTTP /api/health ❌ fail 404, 145ms
HTTP /api/v1/health ❌ fail 404, 136ms
CAO MCP tools/list ❌ fail 404 (tools_list_http_404), default-branch authorized
OpenObserve trace access ✅ pass window=15m
OpenObserve metrics access ✅ pass go_memory_allocated, go_goroutine_count readable
Aggregate telemetry (window 2026-10-07 00:05–04:05 UTC vs. prior 4h)
Metric cao-dashboard cao-collector
Server spans 143,051 8,643
Error-status spans (status_code=2) 0 0
span_status=OK (postgres queries) 4,417 —
p50 / p95 / p99 server duration (ms) 0.72 / 1.50 / 2.40 0.17 / 5,026.8 / 5,030.7†
Postgres query p95 (ms) 658.7 —
Goroutines avg (prior window) 24.48 (24.45) 26.98 (26.98)
Heap allocated avg MB (prior window) 241.6 (241.6) 1,709.6 (1,614.9)

†cao-collector p95/p99 are dominated by long-poll XREADGROUP calls (~5,025ms, ~716/hour, stable), not request failures — expected blocking-read behavior, not a regression.

Prior-window comparison: span volume, error counts, latency, and goroutine counts are stable between the two 4-hour windows; no historical baseline existed in cache (first evaluation), so trend comparison beyond these two windows is unavailable.

Bottlenecks and pressure

  • No error spans, queue backlog, or retry evidence in either window.
  • cao-collector heap-allocated bytes (counter, monotonic) grows steadily (~6MB/15min) over 24h with no observed reset — consistent with normal allocation-rate accumulation for a Go counter metric, not necessarily a leak; go_memory_used{type=stack|other} stayed small and flat (stack ≤1.2MB, other ≤14MB), showing no resident-memory pressure signal.
  • cao-dashboard heap-allocated average is flat (~241.6MB) across both windows — no pressure.
  • Goroutine counts flat for both services (23–27 range) — no leak signal.

Missing telemetry and recommended metrics or APIs

  • GC pause duration is not exported by the stable Go runtime instrumentation (go.memory.allocated, go.memory.used, go.goroutine.count); this is a known instrumentation gap, not a bug. No GC-pause pressure claim can be made.
  • go_memory_used only reports stack and other memory types — no heap type is exported, so true resident/in-use heap size cannot be distinguished from the monotonic go_memory_allocated counter. Recommend exporting the heap dimension of go.memory.used (bytes) if available, to enable resident-memory trend detection distinct from cumulative allocation.
  • No HTTP request/route-level span attributes (http_route, http_response_status_code) were populated in either service's trace data for this window, limiting per-endpoint latency/error attribution; recommend confirming HTTP server instrumentation emits http.route and http.response.status_code on server spans.
  • CAO MCP data-availability tools (cao_catalog, cao_query) were not exercised because the MCP tools/list call itself returned 404; this run's CAO data-availability evidence is unavailable until that route is restored.

Actions

  • Restore or reconfigure the three HTTP health-check routes and the CAO MCP tools/list route at the expected paths.
  • Re-run hosted-health evaluation after the fix to confirm HTTP/MCP checks pass; application-layer telemetry (traces/metrics) shows no corresponding degradation requiring remediation.

cao: githubnext/gh-aw-cao, correlation: (a href="https://github.com/githubnext/gh-aw-cao/actions/runs/37568961540")37568961540-1535(/a)

Generated by SelfCare / Hosted Health · copilot · auto · 180.2 AIC · ⌖ 9.54 AIC · ⊞ 13.3K · ◷

  • expires on Oct 10, 2026, 4:17 AM UTC

Activity

  1. mnkiefer commented on Oct 7, 2026

    @mnkiefer
    ContributorAuthor

    This issue is being closed as outdated. A newer issue has been created: #16633

    View newer issue


    This action was performed automatically by the SelfCare / Hosted Health workflow.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions