DurinDoor
Reference

Model limits

Context windows, output caps, discovery metadata, and oversized request errors.

DurinDoor reports known model limits and checks chat input before dispatch. A missing limit means unknown capacity. It does not mean unlimited capacity or a guaranteed default.

Limit precedence

An explicit operator override wins over live provider metadata and the bundled catalog. Otherwise, lookup uses the provider's model declaration, the exact model ID, the provider registry, then the first matching model-family pattern. Provider and model aliases resolve within their selected provider.

Limit sourceMeaning
customExplicit positive operator context or output override.
liveLimits from discovery or the persisted provider cache.
providerProvider-specific model declaration.
exactExact model declaration in the shared catalog.
registryPublished per-model registry limits.
patternFirst model-family pattern with a positive context window.
defaultNo authoritative limit; unknown fields are absent.

The source label identifies the selected context row. It does not identify the origin of every output limit or capability.

AI Horde and Ollama Local use operator or observed live limits rather than static ceilings. Known Codex models use their underlying API capacity; a client's smaller compaction budget does not replace it.

Operator overrides

The pencil beside a built-in or custom model opens the capability editor under Dashboard → Providers. An override of a built-in model changes that row without adding a duplicate catalog entry. Clearing a limit field removes that override and restores the discovered or catalog value.

A known context window does not establish an output ceiling. An output-only declaration likewise does not establish a context window. Defaults and compaction thresholds are separate values.

Input checks and output reservation

For a known total window, chat rejects a request when counted input plus the effective output reservation exceeds that window. A separate known input ceiling can also reject the input. The response is HTTP 400 with input is too long. This error ends the request; a combo does not try another member.

Claude-compatible providers use native token counting when available. Other providers use an estimate of roughly four characters per token. Failed native counts fall back to that estimate. Unknown positive context limits skip the total-window check; an upstream can still reject the request.

An explicit max_tokens, max_completion_tokens, max_output_tokens, or Gemini generationConfig.maxOutputTokens reserves the value after output clamping. Operator caps take precedence. Without a client limit, reservation uses an explicit operator output cap, then a sent or documented generation default, otherwise zero. A published output maximum only clamps; it is not an automatic reservation.

For example, the bundled Kimi K3 declaration distinguishes its 131072-token generation default from its 1048576-token output ceiling. Input plus generated output must still fit the total window.

Discovery metadata

Provider sync determines available models. Shared models.dev metadata enriches those models; it does not grant account access or add a provider transport. Explicit provider values, including false, win over shared and curated metadata. Explicit operator fields win at the consuming endpoint.

The existing model auto-sync runner refreshes the shared snapshot with a 24-hour TTL. Catalog requests read the persisted cache without downloading it. Scheduled refresh requires an enabled, due provider. Manual sync and forced runs refresh the shared snapshot. A failed download retains cached data; a partial successful download can retain older model entries.

providerFetchedAt records the last successful provider roster sync. sharedEntries records matched shared sources and timestamps; sharedFetchedAt summarizes the oldest known matched timestamp. A recent snapshot download does not prove that every retained model entry is current. Without discovery, the curated catalog remains available.

Public catalog fields

Endpoint or client formatKnown limit fields
OpenAI GET /v1/modelscontext_length, max_model_len, output aliases, and nested input/output limits. Also modality, tools, reasoning, structured-output, caching, and thinking-budget metadata where declared.
Codex catalogcontext_window, max_output_tokens. Selected by originator: codex_cli_rs or a Codex user-agent.
Anthropic catalogmax_input_tokens, max_tokens, nested capabilities, and max_context_window_tokens. Selected by the presence of anthropic-version, before Codex detection. Unknown native ceilings can be null.
GET /v1/models/info?id=provider/modelScoped model capabilities and known limits. Optional kind disambiguates duplicate IDs. Missing id returns 400; an unknown ID returns 404.
Chat combo rowscontext_length and max_completion_tokens use the smallest known limit across resolved members, including nested combos and aliases. Unknown members are skipped; all-unknown limits are omitted.

Web search and web fetch combos keep their kind and omit chat token overlays. /v1beta/models currently advertises fixed 128000 input and 8192 output placeholders; those values do not establish actual capacity.

Client tools can ignore these extension fields or substitute their own defaults. Compare the client's selected model settings with the gateway catalog before relying on a reported capacity.

Bundled capacity examples

These are bundled declarations, before live metadata and operator overrides. Query your running instance for its effective values.

Model routeContext tokensOutput ceiling tokens
OpenAI GPT-6 Astra/Sol/Luna and GPT-6.1 Sol; known Codex equivalents1050000128000
Kimi Platform kimi/kimi-k310485761048576
Kimi Coding kimi-coding/kimi-for-coding1048576Unknown
Kimi Coding kimi-coding/k3-256k262144Unknown
MiniMax minimax/MiniMax-M3.1-Flash-Preview1000000Unknown

The listed OpenAI models also declare a 922000-token input ceiling. Codex capacity does not grant API-model access or every reasoning effort. MiniMax M3.1 Flash Preview requires M Plan or MiniMax Code access. Subscription keys and pay-as-you-go keys are separate.

Capabilities and model access

Input modalities, native output modalities, and server-tool capabilities are separate facts. A text model using an image-generation tool remains a text-output model. Structured-output support does not guarantee prompt-cache behavior. A custom media model sharing a chat model's ID does not inherit chat capabilities or token limits.

Canonical aliases use the target model's provider-scoped capabilities and limits. Retired models leave selectable catalogs. Inventory-only IDs with routingAvailable: false are not callable routes. Restricted, media, and unpublished ceilings remain unknown.

Context values are exact integers. 1050000, 1048576, and 1000000 describe different windows. Account entitlement can impose a smaller usable window, such as a plan-specific Kimi Coding limit. Catalog availability alone does not prove inference access.

See API for request fields and Troubleshooting for recovery.

On this page

Edit on GitHub