Cache affinity

Cache affinity keeps requests from the same session on the most recently successful provider model, instead of re-picking from the head of the candidates on every request.

Why it's needed

A provider's prompt cache can only hit when "consecutive requests share an exactly identical prefix". If consecutive requests in the same session are split across different providers, or across different replicas of one provider, the cache is built in vain — you keep paying full price for the same context.

Failover itself is right, but "pick from the head of the queue on every request" scatters the session.

How it works

  • Once enabled, OSW remembers the provider model a session last succeeded on and prefers reusing it for later requests in the same session.
  • The affinity record has a TTL: after TTL with no new request it expires naturally, and the next request picks afresh.
  • If the affined candidate is no longer available (cooling down / disabled), the request proceeds through normal failover and records the new winner as the affined candidate.

Tuning

In Settings → Reliability → Cache affinity:

SettingDescription
Enable cache affinityMaster switch.
Affinity TTLAfter this idle period with no new request, the affinity expires. Longer is stickier, but takes longer to rebalance.

TTL is a trade-off

If the TTL is too short, the session re-routes before it has built up a cache, which is as good as off; if it is too long, a slow channel is not swapped out for a while. The default is fine for most scenarios.

  • It complements Failover: failover preserves availability, cache affinity preserves cost.
  • The amount of cache hit is visible in Analytics.

To see how requests are doing, start with Request logs.