Skip to content

feat(api): shared upstream client with request deadline and distinct outage reporting - #601

Merged
salazarsebas merged 4 commits into
mainfrom
feat/shared-upstream-client
Oct 10, 2026
Merged

salazarsebas merged 4 commits into
mainfrom
feat/shared-upstream-client

Conversation

@salazarsebas

Copy link
Copy Markdown
Member

What

#451 adds one shared upstream read client (lib/stellar/upstream-client.ts) and moves horizonGet and horizonPaginate onto it.

  • A request-scoped deadline (40 s, created once per inbound request in the request context next to the request id) bounds every read, including a paginated read as a whole.
  • Per-attempt timeouts never exceed the remaining budget, and no backoff sleeps past the deadline.
  • 429, 5xx, timeouts and network faults are retried with full-jitter exponential backoff. Retry-After is honored in seconds and HTTP-date form, capped at 5 s, and a hostile value falls back to the jittered backoff.
  • Failures are typed UpstreamError (rate_limited, timeout, unavailable, bad_response) with plain messages that carry no status code, URL or host. 400, 410 and other 4xx are bad_response and are not retried.
  • The shared RPC server gets bounded timeouts (reads 10 s, simulation 20 s on a separate instance). A thin proxy stops waiting on an allowlisted set of promise-returning reads and simulation when the request deadline passes.
  • sendTransaction, getTransaction and pollTransaction are not on that list: they pass through untouched, are called once by the client, and are never abandoned by the deadline.
  • Errors map through domain-errors.ts (503 service_unavailable, or 502 provider_response_unusable for bad_response) and through the account controller. No new error codes. OpenAPI annotations and openapi.json are updated.

#452 moves path finding and the sponsorship reader onto the client and reports outages distinctly.

  • fetchConversionPath returns route | none | unavailable. none is only an answered empty market or an unroutable input. Because the result is a discriminated object, the compiler rejects any call site that used it as a path.
  • Plan: an unavailable price source adds a "price source unavailable, try again" blocker and drops that asset's decision card, so it can never default to "no route, return to issuer".
  • Build time: unavailable is rethrown as the typed 503, never AssetRouteLostError (409 quote_drifted). The paths endpoint answers 503 instead of { path: null }.
  • Sponsorship: every request goes through the shared client and is counted in the 429 and error counters. A transient 5xx is retried before the enumeration degrades to incomplete, and per-owner concurrency drops from 10 to 2 once a 429 has been seen.
  • fetchConversionPath no longer leaks its timer on throw.
  • assessConversions (no production caller) reports priceSourceUnavailable.

Why

A retried Horizon read could take about 43 s, a paginated read multiplied that per page, and the sponsorship scan, direct-read fallback and RPC calls each had their own budget. The worst case exceeded the proxy's 55 s limit, so users saw a generic 504. A transient 5xx or 429 on path finding returned null, indistinguishable from an illiquid asset, and one transient failure in sponsorship blocked closes that would otherwise succeed.

Tests

New: upstream-client, rpc-deadline, path-finding, price-source-unavailable, plus additions to horizon-http, domain-errors and sponsorship-io. Existing tests were adapted where they asserted raw status codes or URLs in messages, or the old null path result.

  • The sponsorship bug-repro tests fail on the unmodified sources: a transient 5xx marks the enumeration incomplete, and the fan-out does not narrow after a 429.
  • The other new tests exercise new modules and fail on main at import.
  • Whole matrix (type-check, lint, format check, tests) is green locally.

Invariants checked

  • The API re-reads exact on-chain state before building. Retries only re-issue the same read.
  • A failed read never looks empty or "none": unavailable is a distinct type at every path-finding call site, and a test per call-site family covers plan, build and the paths endpoint.
  • The client never retries sendTransaction, and the deadline never abandons it. The existing TRY_AGAIN_LATER loop in submit.ts is unchanged and out of scope.
  • User-facing errors are plain language with no status code, URL or host.
  • Logs from the client carry only a target label and the error kind, per docs/reference/logging-and-privacy.mdx. A test asserts the exact log payload.

Risks

Closes #451
Closes #452

@vercel

vercel Bot commented Oct 10, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
lumenwipe Ready Ready Preview Oct 10, 2026 7:59pm UTC
playground Ready Ready Preview Oct 10, 2026 7:59pm UTC

Comment thread apps/api/src/lib/stellar/upstream-client.ts Fixed
@vercel

vercel Bot commented Oct 10, 2026

Copy link
Copy Markdown

Deployment failed for project playground with the following error:

Resource is limited - try again in 24 hours (more than 100, code: "api-deployments-free-per-day").

Learn More: https://vercel.com/salazarsebas?upgradeToPro=build-rate-limit

@vercel

vercel Bot commented Oct 10, 2026

Copy link
Copy Markdown

Deployment failed for project lumenwipe with the following error:

Resource is limited - try again in 24 hours (more than 100, code: "api-deployments-free-per-day").

Learn More: https://vercel.com/salazarsebas?upgradeToPro=build-rate-limit

@salazarsebas salazarsebas self-assigned this Oct 10, 2026
@salazarsebas
salazarsebas merged commit 14d048f into main Oct 10, 2026
20 of 22 checks passed
@salazarsebas
salazarsebas deleted the feat/shared-upstream-client branch October 10, 2026 20:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants