/fleet reports missing from LiteLLM for a server that is serving normally.
What happened
qwen3.8-27b was redeployed to raise its context window. RunPod gives a new IP and port on every pod recreate, so the host moved from 216.243.220.245:16643 to 103.196.86.105:17140. The fleet_hosts entry was updated, the box came up healthy, and /fleet shows serving.
But LiteLLM still holds the route for the old address. Nothing removed it, and nothing pushed a replacement — the model is unreachable through the proxy while looking healthy in the UI.
Why
Routes are created on the transition into SERVING and deleted on the transition into DOWN. Neither fires here: from the fleet's point of view this is a new server that was never DOWN, and the old one vanished from config rather than failing a health check.
So the deregistration path has a hole — a server removed from fleet_hosts never gets a down transition, and its route outlives it.
What should catch it
check_config_skew already computes bases - fleet_bases — LiteLLM holding routes the fleet does not know about. That branch is precisely this case.
It was previously vacuous, because no route carried an api_base at all. Now that routes exist it should fire, and the fix is to act on it rather than only report it: a route whose api_base is not in the current fleet should be deleted.
Worth pairing with reconciliation on startup, so a portal restart converges the proxy to the fleet rather than relying on transitions alone. That also covers the case where llmao restarts while a box is mid-flight.
Workaround
Restart the portal; it reconciles from scratch.
Note
This will be routine rather than exceptional. RunPod reassigns addresses on every recreate, and we have already moved this host twice tonight. Any config change needing a relaunch produces the same orphan.
/fleet reports missing from LiteLLM for a server that is serving normally.
What happened
qwen3.8-27b was redeployed to raise its context window. RunPod gives a new IP and port on every pod recreate, so the host moved from 216.243.220.245:16643 to 103.196.86.105:17140. The fleet_hosts entry was updated, the box came up healthy, and /fleet shows serving.
But LiteLLM still holds the route for the old address. Nothing removed it, and nothing pushed a replacement — the model is unreachable through the proxy while looking healthy in the UI.
Why
Routes are created on the transition into SERVING and deleted on the transition into DOWN. Neither fires here: from the fleet's point of view this is a new server that was never DOWN, and the old one vanished from config rather than failing a health check.
So the deregistration path has a hole — a server removed from fleet_hosts never gets a down transition, and its route outlives it.
What should catch it
check_config_skew already computes bases - fleet_bases — LiteLLM holding routes the fleet does not know about. That branch is precisely this case.
It was previously vacuous, because no route carried an api_base at all. Now that routes exist it should fire, and the fix is to act on it rather than only report it: a route whose api_base is not in the current fleet should be deleted.
Worth pairing with reconciliation on startup, so a portal restart converges the proxy to the fleet rather than relying on transitions alone. That also covers the case where llmao restarts while a box is mid-flight.
Workaround
Restart the portal; it reconciles from scratch.
Note
This will be routine rather than exceptional. RunPod reassigns addresses on every recreate, and we have already moved this host twice tonight. Any config change needing a relaunch produces the same orphan.