cao.githubnext.com backend (cao-dashboard, cao-collector) is operating normally at the application layer: zero error-status spans, stable p95 request latency (~1.5ms server-side), stable goroutine counts, and no restarts across the last 24h. However, the public HTTP health-check endpoints (/api/readiness, /api/health, /api/v1/health) all return HTTP 404, and the CAO MCP tools/list call also returns HTTP 404 despite correct default-branch authorization — a genuine endpoint-reachability gap that external monitors relying on those routes would miss.
Action: Maintainer should confirm whether the readiness/health/v1-health routes and the MCP tools/list route are intentionally retired or misconfigured in the current deployment, then redeploy or update the health-check path configuration; acceptance check is all three HTTP checks and the MCP tools/list call returning 2xx on the next hosted-health run.
Evidence
| Check |
Outcome |
Detail |
HTTP /api/readiness |
❌ fail |
404, 167ms |
HTTP /api/health |
❌ fail |
404, 145ms |
HTTP /api/v1/health |
❌ fail |
404, 136ms |
CAO MCP tools/list |
❌ fail |
404 (tools_list_http_404), default-branch authorized |
| OpenObserve trace access |
✅ pass |
window=15m |
| OpenObserve metrics access |
✅ pass |
go_memory_allocated, go_goroutine_count readable |
Aggregate telemetry (window 2026-10-07 00:05–04:05 UTC vs. prior 4h)
| Metric |
cao-dashboard |
cao-collector |
| Server spans |
143,051 |
8,643 |
Error-status spans (status_code=2) |
0 |
0 |
span_status=OK (postgres queries) |
4,417 |
— |
| p50 / p95 / p99 server duration (ms) |
0.72 / 1.50 / 2.40 |
0.17 / 5,026.8 / 5,030.7† |
| Postgres query p95 (ms) |
658.7 |
— |
| Goroutines avg (prior window) |
24.48 (24.45) |
26.98 (26.98) |
| Heap allocated avg MB (prior window) |
241.6 (241.6) |
1,709.6 (1,614.9) |
†cao-collector p95/p99 are dominated by long-poll XREADGROUP calls (~5,025ms, ~716/hour, stable), not request failures — expected blocking-read behavior, not a regression.
Prior-window comparison: span volume, error counts, latency, and goroutine counts are stable between the two 4-hour windows; no historical baseline existed in cache (first evaluation), so trend comparison beyond these two windows is unavailable.
Bottlenecks and pressure
- No error spans, queue backlog, or retry evidence in either window.
cao-collector heap-allocated bytes (counter, monotonic) grows steadily (~6MB/15min) over 24h with no observed reset — consistent with normal allocation-rate accumulation for a Go counter metric, not necessarily a leak; go_memory_used{type=stack|other} stayed small and flat (stack ≤1.2MB, other ≤14MB), showing no resident-memory pressure signal.
cao-dashboard heap-allocated average is flat (~241.6MB) across both windows — no pressure.
- Goroutine counts flat for both services (23–27 range) — no leak signal.
Missing telemetry and recommended metrics or APIs
- GC pause duration is not exported by the stable Go runtime instrumentation (
go.memory.allocated, go.memory.used, go.goroutine.count); this is a known instrumentation gap, not a bug. No GC-pause pressure claim can be made.
go_memory_used only reports stack and other memory types — no heap type is exported, so true resident/in-use heap size cannot be distinguished from the monotonic go_memory_allocated counter. Recommend exporting the heap dimension of go.memory.used (bytes) if available, to enable resident-memory trend detection distinct from cumulative allocation.
- No HTTP request/route-level span attributes (
http_route, http_response_status_code) were populated in either service's trace data for this window, limiting per-endpoint latency/error attribution; recommend confirming HTTP server instrumentation emits http.route and http.response.status_code on server spans.
- CAO MCP data-availability tools (
cao_catalog, cao_query) were not exercised because the MCP tools/list call itself returned 404; this run's CAO data-availability evidence is unavailable until that route is restored.
Actions
- Restore or reconfigure the three HTTP health-check routes and the CAO MCP
tools/list route at the expected paths.
- Re-run hosted-health evaluation after the fix to confirm HTTP/MCP checks pass; application-layer telemetry (traces/metrics) shows no corresponding degradation requiring remediation.
cao: githubnext/gh-aw-cao, correlation: (a href="https://github.com/githubnext/gh-aw-cao/actions/runs/37568961540")37568961540-1535(/a)
Generated by SelfCare / Hosted Health · copilot · auto · 180.2 AIC · ⌖ 9.54 AIC · ⊞ 13.3K · ◷
cao.githubnext.combackend (cao-dashboard, cao-collector) is operating normally at the application layer: zero error-status spans, stable p95 request latency (~1.5ms server-side), stable goroutine counts, and no restarts across the last 24h. However, the public HTTP health-check endpoints (/api/readiness,/api/health,/api/v1/health) all return HTTP 404, and the CAO MCPtools/listcall also returns HTTP 404 despite correct default-branch authorization — a genuine endpoint-reachability gap that external monitors relying on those routes would miss.Action: Maintainer should confirm whether the readiness/health/v1-health routes and the MCP tools/list route are intentionally retired or misconfigured in the current deployment, then redeploy or update the health-check path configuration; acceptance check is all three HTTP checks and the MCP
tools/listcall returning 2xx on the next hosted-health run.Evidence
/api/readiness/api/health/api/v1/healthtools/listtools_list_http_404), default-branch authorizedgo_memory_allocated,go_goroutine_countreadableAggregate telemetry (window 2026-10-07 00:05–04:05 UTC vs. prior 4h)
status_code=2)span_status=OK(postgres queries)†cao-collector p95/p99 are dominated by long-poll
XREADGROUPcalls (~5,025ms, ~716/hour, stable), not request failures — expected blocking-read behavior, not a regression.Prior-window comparison: span volume, error counts, latency, and goroutine counts are stable between the two 4-hour windows; no historical baseline existed in cache (first evaluation), so trend comparison beyond these two windows is unavailable.
Bottlenecks and pressure
cao-collectorheap-allocated bytes (counter, monotonic) grows steadily (~6MB/15min) over 24h with no observed reset — consistent with normal allocation-rate accumulation for a Go counter metric, not necessarily a leak;go_memory_used{type=stack|other}stayed small and flat (stack ≤1.2MB, other ≤14MB), showing no resident-memory pressure signal.cao-dashboardheap-allocated average is flat (~241.6MB) across both windows — no pressure.Missing telemetry and recommended metrics or APIs
go.memory.allocated,go.memory.used,go.goroutine.count); this is a known instrumentation gap, not a bug. No GC-pause pressure claim can be made.go_memory_usedonly reportsstackandothermemory types — noheaptype is exported, so true resident/in-use heap size cannot be distinguished from the monotonicgo_memory_allocatedcounter. Recommend exporting theheapdimension ofgo.memory.used(bytes) if available, to enable resident-memory trend detection distinct from cumulative allocation.http_route,http_response_status_code) were populated in either service's trace data for this window, limiting per-endpoint latency/error attribution; recommend confirming HTTP server instrumentation emitshttp.routeandhttp.response.status_codeon server spans.cao_catalog,cao_query) were not exercised because the MCPtools/listcall itself returned 404; this run's CAO data-availability evidence is unavailable until that route is restored.Actions
tools/listroute at the expected paths.