Skip to content

fix: a request the upstream refuses no longer fails over or trips breakers - #38

Merged
fylorn merged 1 commit into
devfrom
fix/upstream-4xx-no-failover
Sep 24, 2026
Merged

fylorn merged 1 commit into
devfrom
fix/upstream-4xx-no-failover

Conversation

@fylorn

@fylorn fylorn commented Sep 24, 2026

Copy link
Copy Markdown
Contributor

Summary

check_status wrapped every upstream non-2xx except 429/401/403 as ProviderError, and is_retryable treated ProviderError as retryable. A client's bad request (400) walked every route of the model, and every hop recorded a breaker failure, so one caller could open the breakers of all routes of a model. ProviderHttpError and ProviderTimeout were never constructed in production.

  • check_status returns ProviderHttpError { status, .. }; reqwest timeouts become ProviderTimeout (send and whole-body read).
  • routing::is_upstream_failure decides failover and breaker accounting:
    • fail over and count against the route: 5xx, 408, 429, timeouts, network errors, unreadable bodies, and 401/403 (UpstreamAuthError: the gateway's own credential for that route was refused, which another route may not share);
    • go straight back to the caller and count as the upstream working: any other 4xx. This matches the desktop gateway.
  • Streams use the same rule in record_outcome (from the recorded error tag and status). A tool-inspection cut (PolicyBlocked) no longer counts against the route either, matching the buffered path.
  • protocol_relearn no longer scrapes the status out of the message text. The probe's is_inconclusive does the same, and treats 5xx/408 as inconclusive.
  • Doc comments in transport.rs and routing.rs rewritten to match.

A 5xx now reaches the caller with the upstream's status rather than a blanket 502.

Test plan

  • Integration, new in gateway_failover.rs:
    • a_refused_request_goes_back_without_trying_another_route: two routes that both answer 400; the caller gets the 400 and its reason; only one upstream is hit;
    • refused_requests_do_not_open_the_breaker: four 400s (buffered and streamed) with min_samples = 2; the fifth request still reaches the upstream;
    • server_errors_open_the_breaker: two 500s open it, and the third request does not reach the upstream.
  • Full integration suite locally (--ignored --test-threads=1)
  • fmt, clippy --all-targets -D warnings and --lib, workspace unit tests

🤖 Generated with Claude Code

…akers

`check_status` reported every non-2xx other than 429/401/403 as
`ProviderError`, and failover treated `ProviderError` as retryable. A
caller's bad request (400) therefore walked every route of the model,
and each hop was recorded as a breaker failure: one caller's malformed
requests could shut every route of a model for everyone.

An upstream status is now kept in `ProviderHttpError`, and
`is_upstream_failure` draws the line the desktop gateway draws. 5xx,
408, 429, timeouts, broken connections, unreadable bodies and a refused
gateway credential (401/403) move on to the next route and count
against the route. Any other 4xx goes straight back to the caller and
counts as the upstream working. Streams apply the same rule when their
outcome is recorded. Timeouts are reported as `ProviderTimeout`.

With the status structured, the relearn and probe classifiers no longer
read it back out of the message text. The probe also treats a 5xx as
inconclusive: an outage is about the moment, not the model.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@fylorn
fylorn merged commit 3c35ab8 into dev Sep 24, 2026
@fylorn
fylorn deleted the fix/upstream-4xx-no-failover branch September 24, 2026 06:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

1 participant