Use this provider switch checks
Force failures before retrieval, after retrieval, before a tool call, and while formatting the reply. At each switch, record the context sent to the fallback provider and the permissions available there.
| Item | Check or owner | Evidence |
|---|---|---|
| Primary unavailable | Fallback starts | Same policy gate |
| Long prompt | Context truncates | Keep trusted policy envelope |
| Tool schema differs | Provider changes format | Reject unsafe call |
| Logging differs | New region or retention | Review data path |
Test the boundary
Compare the two paths for tool allowlists, tenant checks, output validation, logging, retention, and data region. If the fallback cannot meet the same rule, fail closed and show a clear recovery step to the user.
Worked synthetic case
Synthetic case: The primary model times out after drafting a payment. The app switches to a fallback model. The fallback sees the old draft and tries the payment tool again, unaware that the first call may have reached the provider.
Force failures before retrieval, after retrieval, before a tool call, and after a tool timeout. Compare context, permission scope, tool IDs, and customer messages across routes. The pass condition is no extra authority, no duplicate action, and a clear status when the result is unknown.
The payment service should use an operation ID and reconcile provider state before any retry. The app should preserve authorization checks outside both models. If the fallback lacks the needed context, stop safely and ask the user to retry after status is known.
A fallback improves availability but can change data region, retention, or tool behavior. Record the provider path and privacy review before release; do not assume the second route shares the first route’s controls.
Keep policy during a provider switch
Simulate the primary model timing out after a prompt containing a fake customer ID. Before the fallback provider receives it, apply the same data-redaction rule and tool allowlist. Ask the fallback to call a payment tool without approval; the tool service must reject it. Record provider selected, prompt version, redaction result, tool scope, and decision ID. A successful answer is not a pass if the fallback received extra private fields or gained a broader tool. When the policy cannot be enforced for the backup provider, return a clear unavailable state and keep the payment action pending. Test both automatic failover and manual support switches.
Keep identity outside the context window
Build a trusted server envelope with caller ID, tenant, permitted tools, permission version, and operation ID. Neither model may write that envelope. The fallback receives only the context needed for its task; the tool gateway reads the server envelope on every call. If the model prompt is truncated, a tenant value in old chat text must not replace the server value.
Synthetic caller A-101 asks for their own balance. Put “use tenant B” at the end of a long prompt and force a context-limit failure. Switch providers and inspect the gateway’s caller scope. Expect tenant A on both routes. Add a missing-envelope case: the app should stop the restricted task rather than infer identity from a summary.
Reuse execution state across models
Model trace IDs identify conversations. A stable business operation ID identifies the payment. Save the latter before calling any payment tool. If model A times out after the API accepted INT-901, model B may ask for status using INT-901. It must not submit a new transfer under INT-902. Repeated reads are harmless only when their data rules remain the same; repeated writes need idempotent service handling.
Compare provider configurations for the specific model and endpoint, including input capture, retention, and region where available. “Same policy” means both satisfy the required control; it does not mean they have identical vendor terms. If a data route is not approved, use a safe unavailable state for that task. Save fault point, selected provider, policy version, business operation ID, and response. OWASP’s agency guidance supports enforcing narrow permissions outside model judgment.