Skip to content

feat(ai): try OpenAI before Gemini when Anthropic is capped - #29

Merged
ralyodio merged 2 commits into
mainfrom
feat/openai-fallback
Aug 13, 2026
Merged

ralyodio merged 2 commits into
mainfrom
feat/openai-fallback

Conversation

@ralyodio

Copy link
Copy Markdown
Contributor

/setup has been dying on this:

400 {"type":"invalid_request_error","message":"You have reached your specified
API usage limits. You will regain access on 2026-09-01 at 00:00 UTC."}

#25 built the chain for exactly this and made it anthropic -> gemini. This adds
OpenAI in between, so the order becomes anthropic → openai → gemini.

OpenAI goes second rather than last because it is the closer substitute: the
prompts and the deterministic quality gates were built around Claude, and a
same-shaped model is likelier to produce a draft that still passes them.

What's here

  • packages/ai/src/openai.ts — OpenAIModel behind the existing TextModel
    interface, plain HTTP for the reason GeminiModel gives.
  • Wired into the boot chain in apps/server/src/index.ts, keyed on
    OPENAI_API_KEY exactly like the other two. Absent keys just shorten the
    chain; nothing refuses to boot.
  • GEMINI_API_KEY / GEMINI_MODEL documented in .env.example — feat(ai): fall back to Gemini when a provider runs out of budget #25 never
    added them.
  • The 503 that told operators to set ANTHROPIC_API_KEY now names all three.

Three non-obvious things in the adapter

  • max_completion_tokens, never max_tokens. The reasoning models reject
    the latter outright, and a fallback that 400s on every call is worse than none.
  • Reasoning-token headroom, the lesson GeminiModel already paid for.
    Sizing the cap to the caller's maxTokens lets the model reason up to the
    limit and return an empty string, which reads downstream as a model that had
    nothing to say.
  • A refusal is HTTP 200 with content: null and refusal set. Read without
    checking it, that is a successful blank draft.

Default model is gpt-5.5, undated. There is no moving -latest alias for the
reasoning line the way Gemini has one (*-chat-latest is the chat-tuned line,
not the same model renamed), so the generation is bumped by hand and
OPENAI_MODEL overrides it without touching this file.

Verified against the live API

The id resolves. The key currently in the vault answers 429 insufficient_quota / credit_balance_exhausted — which isBudgetExhausted
already classifies as budget, so the chain steps past it to Gemini instead of
failing the request. That exact body is now a test.

The billing gate rejects before request validation, so the body shape itself
could not be exercised end to end on a credit-less key; it is covered by unit
tests against the documented format.

Deploying this changes nothing on its own

Production currently has only ANTHROPIC_API_KEY set. That is why the
Gemini fallback from #25 has been inert since it merged — there was never a
second link. Merging this adds a third link that is also unset.

To actually restore drafting:

railway variables --service 84358cda-e67c-4133-8275-40be706ff5f1 \
  --environment production --set GEMINI_API_KEY=...
railway variables --service 84358cda-e67c-4133-8275-40be706ff5f1 \
  --environment production --set OPENAI_API_KEY=...

GEMINI_API_KEY is the one that matters right now: Anthropic is capped until
2026-09-01 and the OpenAI key has no credits, so Gemini is the only link that
can currently answer. The boot log prints the resolved chain
(model chain: anthropic -> openai -> gemini), so it is checkable after deploy.

510 pass, 0 fail; format and typecheck clean.

🤖 Generated with Claude Code

ralyodio and others added 2 commits August 13, 2026 15:30
The chain added in #25 was anthropic -> gemini. This puts OpenAI between
them, because it is the closer substitute: the prompts and the
deterministic gates were built around Claude, and a same-shaped model is
likelier to produce a draft that still passes them. Gemini stays as the
last link rather than the only one.

`OpenAIModel` is plain HTTP for the reason `GeminiModel` gives — one POST
with a documented body is less to keep working than another SDK. Three
things in it are not obvious and each is a bug avoided:

- `max_completion_tokens`, never `max_tokens`. The reasoning models
  reject the latter outright, and a fallback that 400s on every call is
  worse than no fallback at all.
- The same reasoning-token headroom Gemini needed. Sizing the cap to the
  caller's `maxTokens` lets the model reason up to the limit and return
  an empty string, which reads downstream as a model with nothing to say.
- A refusal arrives as HTTP 200 with `content: null` and `refusal` set.
  Read without checking it, that is a successful blank draft.

The default is `gpt-5.5` rather than a dated snapshot. There is no moving
`-latest` alias for the reasoning line to use the way Gemini has one, so
the generation is bumped by hand and `OPENAI_MODEL` overrides it.

Verified against the live API: the id resolves, and the key in the vault
answers 429 `insufficient_quota` / `credit_balance_exhausted` — which
`isBudgetExhausted` already classifies as budget, so the chain steps past
it to Gemini rather than failing the request. That body is now a test.

Also documents `GEMINI_API_KEY`, which #25 never added to `.env.example`,
and corrects the 503 that told operators to set `ANTHROPIC_API_KEY` as
though it were the only key that enables drafting.

510 pass, 0 fail; format and typecheck clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
# Conflicts:
#	docs/prd-implementation-map.md
@ralyodio
ralyodio merged commit cd84997 into main Aug 13, 2026
4 checks passed
@ralyodio
ralyodio deleted the feat/openai-fallback branch August 16, 2026 17:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant