DurinDoor
Reference

API

Method, path, body, streaming, and provider support for every /v1 route.

Point clients at http://localhost:20128/v1. Production listens on port 20128. Next rewrites /v1/:path* to /api/v1/:path*, the public path is /v1. A second rewrite maps /v1/v1/:path* onto the same handlers for clients that already put /v1 in both the base URL and the path. /codex/:path* and /responses land on POST /v1/responses. /v1beta is a separate tree.

The tables cover the public inference routes and the allowlisted native-provider facade. A listed route does not mean every connected provider can serve it. Pick a model id from GET /v1/models or GET /v1/models/{kind} for that capability.

Key policy, minting, and lifetime caps are on API keys. SDK snippets are on SDK examples.

Authentication

When Require API Key is on, inference requires a valid DurinDoor key. Missing keys return HTTP 401 Missing API key; invalid or expired keys return Invalid API key. With the setting off, trusted loopback HTTP requests can omit a key. Remote HTTP callers still need a valid gateway key or CLI credential at the network boundary, regardless of this setting. Catalog routes have a separate policy below.

Candidates, in order: Authorization: Bearer, x-api-key, x-goog-api-key, then the key query param. Anthropic-shaped clients often send x-api-key; that header is enough. Do not send upstream provider keys to these routes. Use a DurinDoor key from the dashboard.

POST /v1/chat/completions and POST /v1/messages reject a non-JSON Content-Type with HTTP 415 before reading the body. Send application/json.

Request ids

Correlated responses include an x-request-id header containing a server-generated UUID. A client x-request-id does not set that value.

JSON errors also put that id on the body: error.request_id for OpenAI-shaped { error: {...} }, top-level request_id for Anthropic { type: "error", error: {...} } and flat bodies. A validated upstream id, when present and different, is error.upstream_request_id or top-level upstream_request_id. It never replaces the server id. Successful JSON, binary, and SSE bodies pass through without a second parse for correlation.

Catalog

GET /v1 re-exports the models list. GET /v1/models returns chat/LLM rows by default (object: "list"). Send anthropic-version (header presence is enough) for the Anthropic list envelope. Codex CLI (originator: codex_cli_rs or a codex user-agent) gets { models: [...] }.

Expose combos only (Profile, Model catalog) makes GET /v1/models return configured combo names with owned_by: "combo". Default is off. Web combos keep kind so search and fetch stay distinct. HEAD /v1/models returns 200 with an empty body so SDK probes do not wait for the full list.

Capability filters live at GET /v1/models/{kind} for image, tts, stt, embedding, image-to-text, web, rerank, video, music, realtime, moderation, audio, realtime-translation, realtime-transcription, live, systemone, and document-parsing. web includes search and fetch. A single unknown segment still returns Unknown model kind. A provider-prefixed id (GET /v1/models/cc/claude-sonnet-5) returns that model object or HTTP 404 error.code: "model_not_found". HEAD on those paths is a route probe: known kinds and provider-prefixed paths return 200 without proving the model exists.

Rows keep callable id and owned_by. Presentation may add name, provider_name, provider_alias, and gateway_provider. Send id in requests. GET /v1/models/info?id={alias}/{modelId} keeps the registry name and only attaches those extra fields. Optional kind disambiguates duplicate ids. Missing id: 400. Unknown id: 404.

The catalog handlers do not consult requireApiKey. On loopback, reaching /v1/models does not prove that a key is valid. Remote catalog requests must pass the network credential check; dashboard-session access is also allowed for the read-only /api/models and /api/v1/models paths.

MethodPathRequestStreamProviders
GET, OPTIONS/v1none (same body as /v1/models)nocatalog
GET, HEAD, OPTIONS/v1/modelsnonenoconnected chat models; combos when expose-combos-only is on
GET, HEAD, OPTIONS/v1/models/{kind}path kind or provider/modelnokind filter or one LLM id
GET, HEAD, OPTIONS/v1/models/infoquery id, optional kindnoregistry plus search/fetch virtual ids
GET, OPTIONS/v1/provider-plugin-manifestnone; Cache-Control: public, max-age=60noprovider metadata; manifest contract

Chat and completions

Format follows the path: /v1/chat/completions is OpenAI chat (a body with input[] still counts as OpenAI, for Cursor CLI). /v1/messages is always Claude. /v1/responses and /v1/responses/compact are OpenAI Responses. The gateway translates that format to the selected upstream provider.

curl http://localhost:20128/v1/chat/completions \
  -H "Authorization: Bearer YOUR_DURINDOOR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"coding-default","messages":[{"role":"user","content":"Say hello."}],"stream":true}'

Claude-format requests normally end on a user turn. A trailing assistant text prefill becomes a continuation user turn. Trailing tool_use blocks get error tool_result blocks. Opt out per request with X-9Router-Assistant-Prefill: preserve.

POST /v1/completions is a router shim: prompt (string, or a one-element array) becomes one user message, then the chat core runs. Multiple prompts return 400. Streaming SSE chat.completion.chunk frames are rewritten to text_completion. Provider-native /v1/completions is not a goal.

POST /v1/api/chat runs the same chat core, then returns an Ollama-shaped response. Default model name if the body omits model: llama3.2.

POST /v1/responses with "stream": true, or Accept: text/event-stream, opens an early SSE response so slow provider setup does not drop Codex CLI. POST /v1/messages with "stream": true sends Anthropic ping frames on the same timer. Compact is detected from the /v1/responses/compact path; do not send _compact in client JSON.

POST /v1/messages/count_tokens calls a native count on Claude-compatible providers. Anything else, including a missing model, returns { input_tokens, approximate: true } from the heuristic.

A model whose provider has no chat transport and no llm service kind (TTS, STT, search, image-only, or System One providers such as Laya) gets HTTP 400 Model 'provider/model' is not a chat model and cannot be used on the chat endpoints. from every chat endpoint above. The /api/translator/send and /api/translator/translate debug routes run the same check and refuse with Provider {provider} does not serve chat requests instead.

Combos only help when every member can run the client workflow (tools, images, JSON). Mix kinds and the fallback chain fails on the first incompatible hop.

MethodPathRequestStreamProviders
POST, OPTIONS/v1/chat/completionsJSON model, messages; 415 unless application/jsonSSE when stream: trueany connected chat model or combo
POST, HEAD, OPTIONS/v1/completionsJSON model, prompt; extra fields pass throughSSE rewritten to text_completionsame chat core
POST, OPTIONS/v1/responsesJSON model, input (Responses shape)SSE when stream: true or Accept SSEsame chat core
POST, OPTIONS/v1/responses/compactsame as Responses; compact from the pathsame as Responsessame chat core
POST, OPTIONS/v1/messagesJSON model, max_tokens, messages; 415 unless JSONSSE plus Anthropic pings when stream: truesame chat core, Claude wire format
POST, OPTIONS/v1/api/chatJSON model (default llama3.2), chat bodyOllama-shaped, including stream transformsame chat core
POST, OPTIONS/v1/messages/count_tokensJSON model, messagesnonative Claude-compatible count, else estimate

Media

Every media endpoint on this page, plus web search, web fetch, and embeddings below, accepts a request without model (or "model": "auto") except image edit and voices. It then runs the endpoint's media route: the dashboard-ordered models from connected providers of that kind, with fallback. The 400 no_provider_for_kind fires when no connected provider serves the kind, when a saved route has no available model, or when a saved route's models are available but none can run on this endpoint; see media routes for the exact message per case.

Image generation accepts JSON { model, prompt }, optional Accept: text/event-stream, and ?response_format=binary. Image edits accept multipart or provider-native JSON on supported models. xAI accepts image: { url } or image: { file_id }, or up to five sources in images; the adapter preserves storage_options and converts multipart n to a number. sdwebui and comfyui need no stored key.

TTS is JSON { model, input } plus optional language. ?response_format=mp3 (default) or json. Voices: GET /v1/audio/voices?provider={p}&lang=xx with p in elevenlabs, deepgram, inworld, minimax, minimax-cn, edge-tts, local-device. Each voice row includes model ready for /v1/audio/speech.

STT and translation are multipart: model, file. Translation forwards to the provider /audio/translations path. Only OpenAI-format STT providers (OpenAI, Groq, Local Whisper) have one; the others return 400. The transcription and translation handlers declare a maximum route duration of 300 seconds. This is a server execution setting, not a 300-second audio-file limit.

POST /v1/audio/music re-exports POST /v1/music/generations. JSON { model, prompt }. POST /v1/video/generations is the same shape for video-capable models.

/v1/videos is the async job surface: xAI Grok Imagine, MiniMax, MiniMax CN, and OrcaRouter (any provider with a videoConfig that is not the sync veoaifree-web generator). POST /v1/videos/generations, /edits, /extensions create a job. JSON or multipart is forwarded byte-for-byte. The model field (a JSON key or a multipart form field) picks the provider. A provider/ prefix is stripped: JSON is rewritten, and multipart is re-encoded with a new boundary. With no model, the first video-route model whose provider runs async jobs is used, and its id is written into the body (JSON, or the re-encoded multipart form). A bare id must exactly match one connected video model; if two providers serve it, the request returns 400 and asks for provider/model. Combos are rejected. POST returns { request_id, status, ... } and x-9router-connection-id. Poll GET /v1/videos/{request_id} and echo that value as x-connection-id. Done responses include video.url.

Together uses this same async job surface with its native /v2/videos upstream. Common video polls are scoped to the provider connection: two gateway keys authorized for the same account can access that account's jobs. This is not creator-key isolation; the tracked native facade below has a separate ownership contract.

MethodPathRequestStreamProviders
POST, OPTIONS/v1/images/generationsJSON model, promptSSE if Accept SSEimage-kind models; sdwebui, comfyui no-auth
POST, OPTIONS/v1/images/editsmultipart model, image, promptnoimage-edit models
POST, OPTIONS/v1/audio/speechJSON model, input; optional languageno (audio bytes or JSON)tts-kind models
GET, OPTIONS/v1/audio/voicesquery provider (required), langnolisted TTS voice APIs only
POST, OPTIONS/v1/audio/transcriptionsmultipart model, filenostt-kind models
POST, OPTIONS/v1/audio/translationsmultipart model, filenoSTT providers that implement translations
POST, OPTIONS/v1/music/generationsJSON model, promptnomusic-kind models
POST, OPTIONS/v1/audio/musicsame (re-export)nosame as music generations
POST, OPTIONS/v1/video/generationsJSON model, promptnovideo-kind models
POST, GET, OPTIONS/v1/videos/{action}POST generations, edits, or extensions; GET {request_id}noasync video job providers (xAI Grok Imagine, MiniMax, MiniMax CN, Together, OrcaRouter)

Cached ChatGPT Web images

GET /v1/chatgpt-web/image/{id} returns generated image bytes. Loopback callers can retrieve the opaque URL without a key; remote HTTP callers must pass the gateway credential check. A normal remote browser image element cannot attach an authorization header, so direct image rendering has this limitation. Treat the opaque URL as temporary access to that image. The in-memory cache expires entries after 30 minutes, keeps at most 25 entries, and defaults to a total 10 MiB byte cap. Eviction or restart can expire a link sooner. Missing or expired IDs return 404 Image not found or expired. Successful responses use Cache-Control: private, max-age=1800; OPTIONS is supported.

Native provider APIs

Use /v1/native/{provider}/{vendor-path} for vendor-specific fields that the common chat/media APIs cannot express. Supported HTTP providers are openai, anthropic, minimax, minimax-cn, xai, and cohere. The server uses fixed vendor origins and an operation allowlist; clients cannot choose an upstream host or supply provider credentials.

Supply a gateway identity in ?model=provider/model, or in the operation's execution-model field. The server resolves aliases, checks the key's model and account scope, and rewrites execution-model strings without rounding other JSON numbers. Nested Live, transcription, advisor-tool, and batch model fields receive the same checks. Prompts and tool schemas retain their own strings.

Built-in native HTTP operations require a registered model or its canonical alias. Unknown models receive HTTP 400 before account selection or upstream dispatch. Custom System One nodes use their separate operator-managed model registration.

ProviderNative operation families
OpenAIResponses create/retrieve/cancel/delete, Chat Completions, Realtime/translation/transcription ephemeral sessions, Live sessions
AnthropicMessages, native Files and Message Batches
MiniMax / CNOpenAI/Anthropic-compatible text, speech, image, music/lyrics/cover preprocessing, voice management, files, Hailuo/H3 jobs
xAIResponses, image generation/edits, video generation/edits/extensions/polling, TTS and STT
CohereEmbed, Rerank, multipart Audio Transcriptions, image-only Parse

For account-bound operations, send x-connection-id with the account returned in the creation response's x-9router-connection-id. Missing pins return 400; unavailable pins fail instead of switching accounts. Tracked OpenAI Responses and MiniMax/xAI resource operations require the creator key, or local operator access. Completion polling uses the creation model and creator key, and repeated terminal usage does not charge twice. Deleted creator keys cannot redirect charges to the polling key.

Native Anthropic file/batch lists and MiniMax voice-category lists expose the selected account's resources. Share those accounts only with trusted clients. MiniMax POST /v1/get_voice takes voice_type, not an individual voice_id; see the vendor schema.

Native background and batch inference require keys without usage/rate caps: the gateway cannot meter work after a client leaves. Direct ephemeral sessions have the same restriction. Gateway-relayed native WebSockets remain available to scoped keys; see Realtime.

Native JSON/SSE/WS accounting commits reported token usage before the terminal result becomes usable. OpenAI and MiniMax streaming chat requests force stream_options.include_usage: true. Spending the last allowance preserves that result and rejects the next request. Media responses without token usage use a text-input estimate; token/USD caps do not bound vendor charges per character, image, second, or generation. Use vendor account budgets for those charges.

Gemini speech uses Files plus Interactions for transcription and Interactions for WAV speech. Gemini Live uses the gateway's same-process WebSocket relay, not the HTTP facade. NVIDIA hosted Parakeet and Magpie use their published function-specific multipart HTTP endpoints. Together video uses /v2/videos; its translation route supports Whisper, not arbitrary transcription models.

Dify exposes the configured published chat app, not an upstream model catalog. Its stateless chat adapter rejects client tools and media. Key checks use guarded /info and /parameters requests; a selected outbound proxy is rejected because it cannot preserve that probe's DNS guard. Inference retains the selected connection's egress policy.

Tools execute in the client or the vendor's declared server-tool service. DurinDoor does not execute consumer function, computer, or browser tools on the gateway host.

Native background Responses

curl http://localhost:20128/v1/native/openai/v1/responses \
  -H "Authorization: Bearer <gateway-key>" \
  -H "Content-Type: application/json" \
  -d '{"model":"openai/gpt-6.1-sol","input":"Summarize this request.","background":true,"stream":false}'

Keep the returned response id and connection header, then poll the same native path with /responses/{id}?model=openai/gpt-6.1-sol and x-connection-id. Use an uncapped key. The common /v1/responses router retains released provider streaming/conversion behavior.

Embeddings, rerank, and moderation

JSON bodies. Embeddings need model and input. Rerank needs model, query, and documents (array). Moderations need model and input.

MethodPathRequestStreamProviders
POST, OPTIONS/v1/embeddingsJSON model, inputnoembedding-kind models
POST, OPTIONS/v1/rerankJSON model, query, documentsnoCohere/Jina/Voyage-style rerank models
POST, OPTIONS/v1/moderationsJSON model, inputnomoderation models

System One (Jev) decision endpoint

Native pass-through for Jev-shaped decision models, with no chat translation layer. Body needs model, state (the situation to evaluate), and questions (an object, not an array). The provider's JSON response (typed answers, usage) is forwarded unchanged.

Direct TypeSafe uses jev/jev-latest, jev/jev-preview, or jev/jev-1.13.0 with a TypeSafe connection. Existing reseller routes remain separate: opencode-zen/jev-1.13, opencode/jev-1.13-free, and openrouter/typesafe/jev-1.13. Self-hosted Laya exposes laya/auto, laya/english, laya/multilingual, and laya/typed-decisions.

Create a Custom System One node under Media Providers → System One for another compatible server. Enter an API base URL or complete /systemone endpoint, save a connection with an optional key, and register the server's native model ID. Use your-prefix/native-model in requests. Servers without /models, including Laya, support manual registration. Model discovery does not prove inference or authentication on a public listing endpoint.

x-connection-id pins either a credentialed or free route to that account. Aliases receive the same native-model registration and API-key model checks. Structured state values and native controls such as Laya's lang, task, max_len, and min_confidence pass through unchanged. Responses have an 8 MiB proxy limit. Successful usage enters the normal windowed API-key ledger; TypeSafe output is free and input costs $0.042 per million tokens.

Custom System One and Laya URLs retain the resolved-address outbound guard. Outbound proxy pools are unsupported for these guarded URLs; setup rejects that combination instead of falling back to direct traffic. Fixed hosted providers retain their selected egress policy. Ollama remains a chat/embedding backend, not a Jev probability adapter.

MethodPathRequestStreamProviders
POST, OPTIONS/v1/systemoneJSON model, state, questions; optional x-connection-id headernoTypeSafe, native Jev reseller routes, Laya, custom System One nodes

List configured decision models with the dedicated kind; the ordinary chat catalog excludes them:

curl "$DURINDOOR_BASE_URL/v1/models/systemone" \
  -H "Authorization: Bearer $DURINDOOR_API_KEY"

Web search and fetch

Provider is the model. Send provider or model. Catalog ids look like tavily/search and tavily/fetch; the handler strips the suffix when the remainder is a real search or fetch provider. Bare ids and OmniRoute aliases (tavily-search, exa-search, serper-search, google-pse-search, linkup-search, searchapi-search, youcom-search, searxng-search, ollama-search, perplexity-search) also resolve.

Search body: query (required), plus model or provider. Fetch body: url (required), optional format and max_characters. Fetch failures use fallback scope webfetch:<provider> (for Ollama, webfetch:ollama), so a fetch cooldown does not lock that connection for chat.

Search searchConfig providers in the registry today: serper, youcom, searxng, linkup, exa, perplexity, glm, brave-search, searchapi, google-pse, tavily, ollama. Fetch fetchConfig providers: exa, tinyfish, jina-reader, tavily, firecrawl, firecrawl_custom, ollama. Combos work on both routes when every member is that kind.

MethodPathRequestStreamProviders
POST, OPTIONS/v1/searchJSON query plus model or providernosearchConfig / searchViaChat providers and combos
POST, OPTIONS/v1/web/fetchJSON url plus model or provider; optional format, max_charactersnofetchConfig providers and combos

Files and batches

Local store under the data directory (files/<id>), not an upstream Files API. Owner is the DurinDoor API key (or local / operator when keys are not required). JSON POST that is not application/json returns 415. File upload that is not multipart/form-data returns 415.

OpenAI batches: { input_file_id, endpoint, completion_window?, metadata? }. Anthropic batches: { requests: [{ custom_id, params }] }. Cancel stops scheduling, finishes the active row, then marks cancelled. Anthropic results are NDJSON { custom_id, result }.

MethodPathRequestStreamProviders
GET, POST, HEAD, OPTIONS/v1/filesPOST multipart file, purpose (default batch)nolocal disk
GET, DELETE, HEAD, OPTIONS/v1/files/{id}nonenolocal disk
GET, OPTIONS/v1/files/{id}/contentnoneraw byteslocal disk
GET, POST, HEAD, OPTIONS/v1/batchesPOST JSON input_file_id, endpointnolocal executor over chat
GET, OPTIONS/v1/batches/{id}nonenolocal
POST, OPTIONS/v1/batches/{id}/cancelnonenolocal
GET, POST, HEAD, OPTIONS/v1/messages/batchesPOST JSON requests[]nolocal executor, Anthropic surface
GET, OPTIONS/v1/messages/batches/{id}nonenolocal
GET, OPTIONS/v1/messages/batches/{id}/resultsnoneNDJSONlocal
POST, OPTIONS/v1/messages/batches/{id}/cancelnonenolocal

Realtime auth

GET /v1/realtime upgrades to a WebSocket. /v1/realtime/auth checks gateway credentials; /v1/models does not validate keys.

GET /v1/realtime/auth: 200 { ok: true } when the key is admitted (or not required). 401 { error: { message, type: "invalid_request_error", code: "invalid_api_key" } } when missing (and required), invalid, or expired. Authenticate the socket with Authorization: Bearer or the openai-insecure-api-key.<key> subprotocol. The chat-emulation WebSocket is text-only; native realtime models use their declared audio protocol. Event list: Realtime.

MethodPathRequestStreamProviders
GET, OPTIONS/v1/realtime/authsame key headers as chatnonone (auth probe)

Native session creation

These JSON endpoints delegate directly to OpenAI's native session API. Supply a registered gateway model through ?model=openai/model-id or the operation's model field, including nested session or transcription fields where the vendor schema requires them. Other body fields follow the vendor schema.

MethodPathPurpose
POST, OPTIONS/v1/realtime/client_secretsCreate a direct native realtime ephemeral credential.
POST, OPTIONS/v1/realtime/translations/client_secretsCreate a direct translation ephemeral credential.
POST, OPTIONS/v1/realtime/transcription_sessionsCreate a direct transcription session.
POST, OPTIONS/v1/live/sessionsCreate a direct Live session.

Direct sessions require an uncapped key and retain model/account policy checks. A usage-capped key returns 403 Usage-capped API keys cannot mint direct native sessions. Gateway-relayed native WebSockets support scoped keys through Realtime.

Gemini v1beta

/v1beta/:path* rewrites to /api/v1beta/:path*. These routes let a Gemini-native client, such as Gemini CLI or the @google/genai SDK, talk to DurinDoor without an OpenAI shim.

GET /v1beta/models returns a Gemini-style { models: [...] } list built from the static catalog. Every catalog row appears as models/{provider}/{id}. Rows from the gemini provider also appear as plain models/{id} with streamGenerateContent in supportedGenerationMethods. The token limits in that list are fixed placeholders (128000 in, 8192 out), not the real model limits. The list handler does not check a key.

POST /v1beta/models/{model}:generateContent and :streamGenerateContent accept a Gemini request body. {model} is either a bare Gemini id or {provider}/{model}. The URL suffix decides streaming: :streamGenerateContent streams Gemini SSE, :generateContent returns one JSON GenerateContentResponse. The body is converted to the internal chat shape and runs through the same chat core as /v1/chat/completions, so key checks, combos, and fallback behave the same way. inlineData parts (images, PDFs, audio) become data-URL image_url parts in that shape. Parts sent next to a functionResponse in the same content, such as the file a Gemini CLI read_file tool returns, follow the tool message as a user message.

A Gemini text-to-speech request (a Gemini TTS model id, or responseModalities asking for audio) skips that conversion. It is forwarded to Google's generativelanguage.googleapis.com with a saved Gemini credential, after the usual client-key and model-policy checks.

curl "http://localhost:20128/v1beta/models/gemini/gemini-2.5-flash:generateContent" \
  -H "x-goog-api-key: YOUR_DURINDOOR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"contents":[{"role":"user","parts":[{"text":"Say hello."}]}]}'
MethodPathRequestStreamProviders
GET, OPTIONS/v1beta/modelsnonenostatic catalog
POST, OPTIONS/v1beta/models/{model}:generateContentGemini contents bodynoany chat model; Gemini TTS goes to Google
POST, OPTIONS/v1beta/models/{model}:streamGenerateContentGemini contents bodyyesany chat model; Gemini TTS goes to Google

Unknown paths

Unmatched /v1/* returns JSON { error: { type: "not_found", code: "unknown_route", path } } with HTTP 404, not the dashboard HTML shell. Registered routes take precedence over that fallback.

MethodPathRequestStreamProviders
GET, POST, PUT, PATCH, DELETE, HEAD, OPTIONSunmatched /v1/*nonenonone

Management API

Dashboard REST under /api/* is on Management API. The same DurinDoor API key that calls /v1 can call most of those routes; exclusions, redaction, and curl examples live there. MCP tools for the same work: MCP control.

See MCP gateway and compression for gateway features that are not /v1 routes.

On this page

Edit on GitHub