Failover

Failover means that when one provider fails, OSW automatically tries the next candidate in the same group. It keeps requests running even when a single provider wobbles, rate-limits, or has its auth expire.

When it switches to the next

SituationBehavior
Network error / connection timeout / streaming idle timeoutSwitch to the next candidate.
401 / 403Switch to the next candidate, and count it as one failure for that provider.
408 / 429Switch to the next candidate.
5xxSwitch to the next candidate.
Other 4xxReturn as-is, no switch.

Once streaming has started, no stitching

Once a streaming response has begun to be sent, a mid-stream break aborts it; the output of multiple channels is never stitched together.

Cooldown

After consecutive failures reach the threshold, that provider (or that provider model) enters cooldown and is temporarily removed from the candidates:

  • It is no longer tried during cooldown, and requests flow straight to other candidates.
  • The cooldown duration grows with the failure count up to a ceiling, then it is automatically put back.
  • Cooldown state is visible on the model rows of a logical model.

Tuning

All parameters live in Settings → Reliability → Failover:

SettingDefaultMeaning
Failure threshold3How many consecutive failures trigger cooldown.
Base cooldown30sThe starting cooldown duration.
Max cooldown5 minThe cooldown duration ceiling.
Request timeout30sTimeout for a single attempt; on timeout it switches to the next.

Timeout is also per provider

A provider's own "request timeout" overrides the default here (see Providers & models).

It only fails over within a group

Failover only happens within a group. A group can be:

  • the models attached to one logical model; or
  • one channel that has explicitly enabled protocol conversion.

No cross-group fallback

If the channel a request uses does not support the native protocol and has no conversion enabled, it will not switch to another group for that reason — that is routing's job (see Request routing).

Observing it

  • Logical models page: a model row's "consecutive failures" and "cooling down" markers.
  • Analytics: the frequency of failover and the distribution of failure causes.
  • Request logs: the attempt detail of a request, showing which candidate it failed on and who it finally landed on.

If the same session repeatedly hits different providers, the upstream's prompt cache gets scattered; the next section, Cache affinity, solves this.