Model limits
Context windows, output caps, discovery metadata, and oversized request errors.
DurinDoor reports known model limits and checks chat input before dispatch. A missing limit means unknown capacity. It does not mean unlimited capacity or a guaranteed default.
Limit precedence
An explicit operator override wins over live provider metadata and the bundled catalog. Otherwise, lookup uses the provider's model declaration, the exact model ID, the provider registry, then the first matching model-family pattern. Provider and model aliases resolve within their selected provider.
| Limit source | Meaning |
|---|---|
custom | Explicit positive operator context or output override. |
live | Limits from discovery or the persisted provider cache. |
provider | Provider-specific model declaration. |
exact | Exact model declaration in the shared catalog. |
registry | Published per-model registry limits. |
pattern | First model-family pattern with a positive context window. |
default | No authoritative limit; unknown fields are absent. |
The source label identifies the selected context row. It does not identify the origin of every output limit or capability.
AI Horde and Ollama Local use operator or observed live limits rather than static ceilings. Known Codex models use their underlying API capacity; a client's smaller compaction budget does not replace it.
Operator overrides
The pencil beside a built-in or custom model opens the capability editor under Dashboard → Providers. An override of a built-in model changes that row without adding a duplicate catalog entry. Clearing a limit field removes that override and restores the discovered or catalog value.
A known context window does not establish an output ceiling. An output-only declaration likewise does not establish a context window. Defaults and compaction thresholds are separate values.
Input checks and output reservation
For a known total window, chat rejects a request when counted input plus the effective output reservation exceeds that window. A separate known input ceiling can also reject the input. The response is HTTP 400 with input is too long. This error ends the request; a combo does not try another member.
Claude-compatible providers use native token counting when available. Other providers use an estimate of roughly four characters per token. Failed native counts fall back to that estimate. Unknown positive context limits skip the total-window check; an upstream can still reject the request.
An explicit max_tokens, max_completion_tokens, max_output_tokens, or Gemini generationConfig.maxOutputTokens reserves the value after output clamping. Operator caps take precedence. Without a client limit, reservation uses an explicit operator output cap, then a sent or documented generation default, otherwise zero. A published output maximum only clamps; it is not an automatic reservation.
For example, the bundled Kimi K3 declaration distinguishes its 131072-token generation default from its 1048576-token output ceiling. Input plus generated output must still fit the total window.
Discovery metadata
Provider sync determines available models. Shared models.dev metadata enriches those models; it does not grant account access or add a provider transport. Explicit provider values, including false, win over shared and curated metadata. Explicit operator fields win at the consuming endpoint.
The existing model auto-sync runner refreshes the shared snapshot with a 24-hour TTL. Catalog requests read the persisted cache without downloading it. Scheduled refresh requires an enabled, due provider. Manual sync and forced runs refresh the shared snapshot. A failed download retains cached data; a partial successful download can retain older model entries.
providerFetchedAt records the last successful provider roster sync. sharedEntries records matched shared sources and timestamps; sharedFetchedAt summarizes the oldest known matched timestamp. A recent snapshot download does not prove that every retained model entry is current. Without discovery, the curated catalog remains available.
Public catalog fields
| Endpoint or client format | Known limit fields |
|---|---|
OpenAI GET /v1/models | context_length, max_model_len, output aliases, and nested input/output limits. Also modality, tools, reasoning, structured-output, caching, and thinking-budget metadata where declared. |
| Codex catalog | context_window, max_output_tokens. Selected by originator: codex_cli_rs or a Codex user-agent. |
| Anthropic catalog | max_input_tokens, max_tokens, nested capabilities, and max_context_window_tokens. Selected by the presence of anthropic-version, before Codex detection. Unknown native ceilings can be null. |
GET /v1/models/info?id=provider/model | Scoped model capabilities and known limits. Optional kind disambiguates duplicate IDs. Missing id returns 400; an unknown ID returns 404. |
| Chat combo rows | context_length and max_completion_tokens use the smallest known limit across resolved members, including nested combos and aliases. Unknown members are skipped; all-unknown limits are omitted. |
Web search and web fetch combos keep their kind and omit chat token overlays. /v1beta/models currently advertises fixed 128000 input and 8192 output placeholders; those values do not establish actual capacity.
Client tools can ignore these extension fields or substitute their own defaults. Compare the client's selected model settings with the gateway catalog before relying on a reported capacity.
Bundled capacity examples
These are bundled declarations, before live metadata and operator overrides. Query your running instance for its effective values.
| Model route | Context tokens | Output ceiling tokens |
|---|---|---|
| OpenAI GPT-6 Astra/Sol/Luna and GPT-6.1 Sol; known Codex equivalents | 1050000 | 128000 |
Kimi Platform kimi/kimi-k3 | 1048576 | 1048576 |
Kimi Coding kimi-coding/kimi-for-coding | 1048576 | Unknown |
Kimi Coding kimi-coding/k3-256k | 262144 | Unknown |
MiniMax minimax/MiniMax-M3.1-Flash-Preview | 1000000 | Unknown |
The listed OpenAI models also declare a 922000-token input ceiling. Codex capacity does not grant API-model access or every reasoning effort. MiniMax M3.1 Flash Preview requires M Plan or MiniMax Code access. Subscription keys and pay-as-you-go keys are separate.
Capabilities and model access
Input modalities, native output modalities, and server-tool capabilities are separate facts. A text model using an image-generation tool remains a text-output model. Structured-output support does not guarantee prompt-cache behavior. A custom media model sharing a chat model's ID does not inherit chat capabilities or token limits.
Canonical aliases use the target model's provider-scoped capabilities and limits. Retired models leave selectable catalogs. Inventory-only IDs with routingAvailable: false are not callable routes. Restricted, media, and unpublished ceilings remain unknown.
Context values are exact integers. 1050000, 1048576, and 1000000 describe different windows. Account entitlement can impose a smaller usable window, such as a plan-specific Kimi Coding limit. Catalog availability alone does not prove inference access.
See API for request fields and Troubleshooting for recovery.