Smart routing
Choose a model, set account selection, and verify fallback without losing conversation affinity.
Choose a model and account strategy for chat traffic, then check the selected account in Dashboard → Usage. Add and test an active provider connection before changing routing.
Choose a model ID
Read GET /v1/models or use the dashboard model picker. Prefer the exact listed ID:
| Model form | Selection |
|---|---|
openai/gpt-4.1 | A provider and upstream model. |
cc/claude-sonnet | A provider alias and model. |
my-node/local-model | A configured compatible node and model. |
daily-coder | A custom model alias. |
coding-default | A named combo. |
A provider prefix before the first slash selects the provider; further slashes remain part of the model ID. Bare unknown model names can be inferred to a provider, so use a prefix when you need an unambiguous route.
Choose an account strategy
Open Dashboard → Settings or Providers → your provider. A provider override takes precedence over the global setting.
| Strategy | Use |
|---|---|
| Fill-first | Prefer the first eligible account. |
| Round Robin | Rotate accounts after the sticky limit, which defaults to three calls and accepts 1–10. |
| Cache Affinity | Keep a conversation on the same eligible account using its session ID; the assignment is stable across restarts. Without a session ID, use fill-first. |
| Quota-weighted | Spread calls randomly across accounts with comparable fresh quota, weighted by remaining capacity. Without comparable quota, use fill-first order. |
The quota-weighted floor defaults to 1% remaining. Thin accounts are excluded from the draw unless all comparable accounts are thin. Exhausted, cooled-down, or routing-floor-blocked accounts stay ineligible.
Inactive, locked, expired, or scope-excluded connections are skipped. Fresh quota may also skip an exhausted account. Missing or stale quota keeps the account eligible; see Quota tracking.
Keep a conversation on one account
Send the same x-session-id on each turn:
curl http://localhost:20128/v1/chat/completions \
-H "Authorization: Bearer $DURINDOOR_KEY" \
-H "x-session-id: review-1" \
-H "Content-Type: application/json" \
-d '{"model":"openai/gpt-4.1","messages":[{"role":"user","content":"Reply OK."}]}'Replace the model with a connected model from your catalog. Confirm the selected account in Usage, then send a second turn with the same session ID. Round Robin keeps an in-memory affinity map; Cache Affinity derives the assignment without that map. An account becoming ineligible can change the selection.
Other accepted session headers, in priority order after x-session-id, are session-id, session_id, x-amp-thread-id, and x-client-request-id. Body fields prompt_cache_key, session_id, and conversation_id are also accepted. Kiro ignores x-client-request-id for affinity.
Use x-connection-id only when you intend to pin an account. An eligible preferred account takes precedence over rotation. A pinned direct stream does not hop to another account.
Check compatibility and failures
DurinDoor translates client formats to provider formats. Matching formats skip translation. Some direct pairs preserve fields that a conversion through OpenAI would lose; compatibility still depends on the provider and model. A non-streaming provider can be returned as a chat-completion stream when the client asks for SSE.
Malformed requests return HTTP 400 before routing. Fix the body rather than adding fallback models. Account failures and rate limits can select another eligible account, then another combo member.
For a direct stream that ends before meaningful content, DurinDoor can try one sibling account during the initial peek window, which defaults to 10 seconds. After content reaches the client, there is no retry. Request errors such as an excessive context length do not trigger an account hop; a combo may still try a different model.
If a started stream stops, inspect request logs for stall_timeout, ttft_timeout, stream_terminated, or stream_error. Check upstream connectivity, provider status, and backoff settings before retrying. A retry after partial output can create a second provider charge.