API
Method, path, body, streaming, and provider support for every /v1 route.
Point clients at http://localhost:20128/v1. Production listens on port 20128. Next rewrites /v1/:path* to /api/v1/:path*, the public path is /v1. A second rewrite maps /v1/v1/:path* onto the same handlers for clients that already put /v1 in both the base URL and the path. /codex/:path* and /responses land on POST /v1/responses. /v1beta is a separate tree.
The tables cover the public inference routes and the allowlisted native-provider facade. A listed route does not mean every connected provider can serve it. Pick a model id from GET /v1/models or GET /v1/models/{kind} for that capability.
Key policy, minting, and lifetime caps are on API keys. SDK snippets are on SDK examples.
Authentication
When Require API Key is on, inference requires a valid DurinDoor key. Missing keys return HTTP 401 Missing API key; invalid or expired keys return Invalid API key. With the setting off, trusted loopback HTTP requests can omit a key. Remote HTTP callers still need a valid gateway key or CLI credential at the network boundary, regardless of this setting. Catalog routes have a separate policy below.
Candidates, in order: Authorization: Bearer, x-api-key, x-goog-api-key, then the key query param. Anthropic-shaped clients often send x-api-key; that header is enough. Do not send upstream provider keys to these routes. Use a DurinDoor key from the dashboard.
POST /v1/chat/completions and POST /v1/messages reject a non-JSON Content-Type with HTTP 415 before reading the body. Send application/json.
Request ids
Correlated responses include an x-request-id header containing a server-generated UUID. A client x-request-id does not set that value.
JSON errors also put that id on the body: error.request_id for OpenAI-shaped { error: {...} }, top-level request_id for Anthropic { type: "error", error: {...} } and flat bodies. A validated upstream id, when present and different, is error.upstream_request_id or top-level upstream_request_id. It never replaces the server id. Successful JSON, binary, and SSE bodies pass through without a second parse for correlation.
Catalog
GET /v1 re-exports the models list. GET /v1/models returns chat/LLM rows by default (object: "list"). Send anthropic-version (header presence is enough) for the Anthropic list envelope. Codex CLI (originator: codex_cli_rs or a codex user-agent) gets { models: [...] }.
Expose combos only (Profile, Model catalog) makes GET /v1/models return configured combo names with owned_by: "combo". Default is off. Web combos keep kind so search and fetch stay distinct. HEAD /v1/models returns 200 with an empty body so SDK probes do not wait for the full list.
Capability filters live at GET /v1/models/{kind} for image, tts, stt, embedding, image-to-text, web, rerank, video, music, realtime, moderation, audio, realtime-translation, realtime-transcription, live, systemone, and document-parsing. web includes search and fetch. A single unknown segment still returns Unknown model kind. A provider-prefixed id (GET /v1/models/cc/claude-sonnet-5) returns that model object or HTTP 404 error.code: "model_not_found". HEAD on those paths is a route probe: known kinds and provider-prefixed paths return 200 without proving the model exists.
Rows keep callable id and owned_by. Presentation may add name, provider_name, provider_alias, and gateway_provider. Send id in requests. GET /v1/models/info?id={alias}/{modelId} keeps the registry name and only attaches those extra fields. Optional kind disambiguates duplicate ids. Missing id: 400. Unknown id: 404.
The catalog handlers do not consult requireApiKey. On loopback, reaching /v1/models does not prove that a key is valid. Remote catalog requests must pass the network credential check; dashboard-session access is also allowed for the read-only /api/models and /api/v1/models paths.
| Method | Path | Request | Stream | Providers |
|---|---|---|---|---|
| GET, OPTIONS | /v1 | none (same body as /v1/models) | no | catalog |
| GET, HEAD, OPTIONS | /v1/models | none | no | connected chat models; combos when expose-combos-only is on |
| GET, HEAD, OPTIONS | /v1/models/{kind} | path kind or provider/model | no | kind filter or one LLM id |
| GET, HEAD, OPTIONS | /v1/models/info | query id, optional kind | no | registry plus search/fetch virtual ids |
| GET, OPTIONS | /v1/provider-plugin-manifest | none; Cache-Control: public, max-age=60 | no | provider metadata; manifest contract |
Chat and completions
Format follows the path: /v1/chat/completions is OpenAI chat (a body with input[] still counts as OpenAI, for Cursor CLI). /v1/messages is always Claude. /v1/responses and /v1/responses/compact are OpenAI Responses. The gateway translates that format to the selected upstream provider.
curl http://localhost:20128/v1/chat/completions \
-H "Authorization: Bearer YOUR_DURINDOOR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"coding-default","messages":[{"role":"user","content":"Say hello."}],"stream":true}'Claude-format requests normally end on a user turn. A trailing assistant text prefill becomes a continuation user turn. Trailing tool_use blocks get error tool_result blocks. Opt out per request with X-9Router-Assistant-Prefill: preserve.
POST /v1/completions is a router shim: prompt (string, or a one-element array) becomes one user message, then the chat core runs. Multiple prompts return 400. Streaming SSE chat.completion.chunk frames are rewritten to text_completion. Provider-native /v1/completions is not a goal.
POST /v1/api/chat runs the same chat core, then returns an Ollama-shaped response. Default model name if the body omits model: llama3.2.
POST /v1/responses with "stream": true, or Accept: text/event-stream, opens an early SSE response so slow provider setup does not drop Codex CLI. POST /v1/messages with "stream": true sends Anthropic ping frames on the same timer. Compact is detected from the /v1/responses/compact path; do not send _compact in client JSON.
POST /v1/messages/count_tokens calls a native count on Claude-compatible providers. Anything else, including a missing model, returns { input_tokens, approximate: true } from the heuristic.
A model whose provider has no chat transport and no llm service kind (TTS, STT, search, image-only, or System One providers such as Laya) gets HTTP 400 Model 'provider/model' is not a chat model and cannot be used on the chat endpoints. from every chat endpoint above. The /api/translator/send and /api/translator/translate debug routes run the same check and refuse with Provider {provider} does not serve chat requests instead.
Combos only help when every member can run the client workflow (tools, images, JSON). Mix kinds and the fallback chain fails on the first incompatible hop.
| Method | Path | Request | Stream | Providers |
|---|---|---|---|---|
| POST, OPTIONS | /v1/chat/completions | JSON model, messages; 415 unless application/json | SSE when stream: true | any connected chat model or combo |
| POST, HEAD, OPTIONS | /v1/completions | JSON model, prompt; extra fields pass through | SSE rewritten to text_completion | same chat core |
| POST, OPTIONS | /v1/responses | JSON model, input (Responses shape) | SSE when stream: true or Accept SSE | same chat core |
| POST, OPTIONS | /v1/responses/compact | same as Responses; compact from the path | same as Responses | same chat core |
| POST, OPTIONS | /v1/messages | JSON model, max_tokens, messages; 415 unless JSON | SSE plus Anthropic pings when stream: true | same chat core, Claude wire format |
| POST, OPTIONS | /v1/api/chat | JSON model (default llama3.2), chat body | Ollama-shaped, including stream transform | same chat core |
| POST, OPTIONS | /v1/messages/count_tokens | JSON model, messages | no | native Claude-compatible count, else estimate |
Media
Every media endpoint on this page, plus web search, web fetch, and embeddings below, accepts a request without model (or "model": "auto") except image edit and voices. It then runs the endpoint's media route: the dashboard-ordered models from connected providers of that kind, with fallback. The 400 no_provider_for_kind fires when no connected provider serves the kind, when a saved route has no available model, or when a saved route's models are available but none can run on this endpoint; see media routes for the exact message per case.
Image generation accepts JSON { model, prompt }, optional Accept: text/event-stream, and ?response_format=binary. Image edits accept multipart or provider-native JSON on supported models. xAI accepts image: { url } or image: { file_id }, or up to five sources in images; the adapter preserves storage_options and converts multipart n to a number. sdwebui and comfyui need no stored key.
TTS is JSON { model, input } plus optional language. ?response_format=mp3 (default) or json. Voices: GET /v1/audio/voices?provider={p}&lang=xx with p in elevenlabs, deepgram, inworld, minimax, minimax-cn, edge-tts, local-device. Each voice row includes model ready for /v1/audio/speech.
STT and translation are multipart: model, file. Translation forwards to the provider /audio/translations path. Only OpenAI-format STT providers (OpenAI, Groq, Local Whisper) have one; the others return 400. The transcription and translation handlers declare a maximum route duration of 300 seconds. This is a server execution setting, not a 300-second audio-file limit.
POST /v1/audio/music re-exports POST /v1/music/generations. JSON { model, prompt }. POST /v1/video/generations is the same shape for video-capable models.
/v1/videos is the async job surface: xAI Grok Imagine, MiniMax, MiniMax CN, and OrcaRouter (any provider with a videoConfig that is not the sync veoaifree-web generator). POST /v1/videos/generations, /edits, /extensions create a job. JSON or multipart is forwarded byte-for-byte. The model field (a JSON key or a multipart form field) picks the provider. A provider/ prefix is stripped: JSON is rewritten, and multipart is re-encoded with a new boundary. With no model, the first video-route model whose provider runs async jobs is used, and its id is written into the body (JSON, or the re-encoded multipart form). A bare id must exactly match one connected video model; if two providers serve it, the request returns 400 and asks for provider/model. Combos are rejected. POST returns { request_id, status, ... } and x-9router-connection-id. Poll GET /v1/videos/{request_id} and echo that value as x-connection-id. Done responses include video.url.
Together uses this same async job surface with its native /v2/videos upstream. Common video polls are scoped to the provider connection: two gateway keys authorized for the same account can access that account's jobs. This is not creator-key isolation; the tracked native facade below has a separate ownership contract.
| Method | Path | Request | Stream | Providers |
|---|---|---|---|---|
| POST, OPTIONS | /v1/images/generations | JSON model, prompt | SSE if Accept SSE | image-kind models; sdwebui, comfyui no-auth |
| POST, OPTIONS | /v1/images/edits | multipart model, image, prompt | no | image-edit models |
| POST, OPTIONS | /v1/audio/speech | JSON model, input; optional language | no (audio bytes or JSON) | tts-kind models |
| GET, OPTIONS | /v1/audio/voices | query provider (required), lang | no | listed TTS voice APIs only |
| POST, OPTIONS | /v1/audio/transcriptions | multipart model, file | no | stt-kind models |
| POST, OPTIONS | /v1/audio/translations | multipart model, file | no | STT providers that implement translations |
| POST, OPTIONS | /v1/music/generations | JSON model, prompt | no | music-kind models |
| POST, OPTIONS | /v1/audio/music | same (re-export) | no | same as music generations |
| POST, OPTIONS | /v1/video/generations | JSON model, prompt | no | video-kind models |
| POST, GET, OPTIONS | /v1/videos/{action} | POST generations, edits, or extensions; GET {request_id} | no | async video job providers (xAI Grok Imagine, MiniMax, MiniMax CN, Together, OrcaRouter) |
Cached ChatGPT Web images
GET /v1/chatgpt-web/image/{id} returns generated image bytes. Loopback callers can retrieve the opaque URL without a key; remote HTTP callers must pass the gateway credential check. A normal remote browser image element cannot attach an authorization header, so direct image rendering has this limitation. Treat the opaque URL as temporary access to that image. The in-memory cache expires entries after 30 minutes, keeps at most 25 entries, and defaults to a total 10 MiB byte cap. Eviction or restart can expire a link sooner. Missing or expired IDs return 404 Image not found or expired. Successful responses use Cache-Control: private, max-age=1800; OPTIONS is supported.
Native provider APIs
Use /v1/native/{provider}/{vendor-path} for vendor-specific fields that the common chat/media APIs cannot express. Supported HTTP providers are openai, anthropic, minimax, minimax-cn, xai, and cohere. The server uses fixed vendor origins and an operation allowlist; clients cannot choose an upstream host or supply provider credentials.
Supply a gateway identity in ?model=provider/model, or in the operation's execution-model field. The server resolves aliases, checks the key's model and account scope, and rewrites execution-model strings without rounding other JSON numbers. Nested Live, transcription, advisor-tool, and batch model fields receive the same checks. Prompts and tool schemas retain their own strings.
Built-in native HTTP operations require a registered model or its canonical alias. Unknown models receive HTTP 400 before account selection or upstream dispatch. Custom System One nodes use their separate operator-managed model registration.
| Provider | Native operation families |
|---|---|
| OpenAI | Responses create/retrieve/cancel/delete, Chat Completions, Realtime/translation/transcription ephemeral sessions, Live sessions |
| Anthropic | Messages, native Files and Message Batches |
| MiniMax / CN | OpenAI/Anthropic-compatible text, speech, image, music/lyrics/cover preprocessing, voice management, files, Hailuo/H3 jobs |
| xAI | Responses, image generation/edits, video generation/edits/extensions/polling, TTS and STT |
| Cohere | Embed, Rerank, multipart Audio Transcriptions, image-only Parse |
For account-bound operations, send x-connection-id with the account returned in the creation response's x-9router-connection-id. Missing pins return 400; unavailable pins fail instead of switching accounts. Tracked OpenAI Responses and MiniMax/xAI resource operations require the creator key, or local operator access. Completion polling uses the creation model and creator key, and repeated terminal usage does not charge twice. Deleted creator keys cannot redirect charges to the polling key.
Native Anthropic file/batch lists and MiniMax voice-category lists expose the selected account's resources. Share those accounts only with trusted clients. MiniMax POST /v1/get_voice takes voice_type, not an individual voice_id; see the vendor schema.
Native background and batch inference require keys without usage/rate caps: the gateway cannot meter work after a client leaves. Direct ephemeral sessions have the same restriction. Gateway-relayed native WebSockets remain available to scoped keys; see Realtime.
Native JSON/SSE/WS accounting commits reported token usage before the terminal result becomes usable. OpenAI and MiniMax streaming chat requests force stream_options.include_usage: true. Spending the last allowance preserves that result and rejects the next request. Media responses without token usage use a text-input estimate; token/USD caps do not bound vendor charges per character, image, second, or generation. Use vendor account budgets for those charges.
Gemini speech uses Files plus Interactions for transcription and Interactions for WAV speech. Gemini Live uses the gateway's same-process WebSocket relay, not the HTTP facade. NVIDIA hosted Parakeet and Magpie use their published function-specific multipart HTTP endpoints. Together video uses /v2/videos; its translation route supports Whisper, not arbitrary transcription models.
Dify exposes the configured published chat app, not an upstream model catalog. Its stateless chat adapter rejects client tools and media. Key checks use guarded /info and /parameters requests; a selected outbound proxy is rejected because it cannot preserve that probe's DNS guard. Inference retains the selected connection's egress policy.
Tools execute in the client or the vendor's declared server-tool service. DurinDoor does not execute consumer function, computer, or browser tools on the gateway host.
Native background Responses
curl http://localhost:20128/v1/native/openai/v1/responses \
-H "Authorization: Bearer <gateway-key>" \
-H "Content-Type: application/json" \
-d '{"model":"openai/gpt-6.1-sol","input":"Summarize this request.","background":true,"stream":false}'Keep the returned response id and connection header, then poll the same native path with /responses/{id}?model=openai/gpt-6.1-sol and x-connection-id. Use an uncapped key. The common /v1/responses router retains released provider streaming/conversion behavior.
Embeddings, rerank, and moderation
JSON bodies. Embeddings need model and input. Rerank needs model, query, and documents (array). Moderations need model and input.
| Method | Path | Request | Stream | Providers |
|---|---|---|---|---|
| POST, OPTIONS | /v1/embeddings | JSON model, input | no | embedding-kind models |
| POST, OPTIONS | /v1/rerank | JSON model, query, documents | no | Cohere/Jina/Voyage-style rerank models |
| POST, OPTIONS | /v1/moderations | JSON model, input | no | moderation models |
System One (Jev) decision endpoint
Native pass-through for Jev-shaped decision models, with no chat translation layer. Body needs model, state (the situation to evaluate), and questions (an object, not an array). The provider's JSON response (typed answers, usage) is forwarded unchanged.
Direct TypeSafe uses jev/jev-latest, jev/jev-preview, or jev/jev-1.13.0 with a TypeSafe connection. Existing reseller routes remain separate: opencode-zen/jev-1.13, opencode/jev-1.13-free, and openrouter/typesafe/jev-1.13. Self-hosted Laya exposes laya/auto, laya/english, laya/multilingual, and laya/typed-decisions.
Create a Custom System One node under Media Providers → System One for another compatible server. Enter an API base URL or complete /systemone endpoint, save a connection with an optional key, and register the server's native model ID. Use your-prefix/native-model in requests. Servers without /models, including Laya, support manual registration. Model discovery does not prove inference or authentication on a public listing endpoint.
x-connection-id pins either a credentialed or free route to that account. Aliases receive the same native-model registration and API-key model checks. Structured state values and native controls such as Laya's lang, task, max_len, and min_confidence pass through unchanged. Responses have an 8 MiB proxy limit. Successful usage enters the normal windowed API-key ledger; TypeSafe output is free and input costs $0.042 per million tokens.
Custom System One and Laya URLs retain the resolved-address outbound guard. Outbound proxy pools are unsupported for these guarded URLs; setup rejects that combination instead of falling back to direct traffic. Fixed hosted providers retain their selected egress policy. Ollama remains a chat/embedding backend, not a Jev probability adapter.
| Method | Path | Request | Stream | Providers |
|---|---|---|---|---|
| POST, OPTIONS | /v1/systemone | JSON model, state, questions; optional x-connection-id header | no | TypeSafe, native Jev reseller routes, Laya, custom System One nodes |
List configured decision models with the dedicated kind; the ordinary chat catalog excludes them:
curl "$DURINDOOR_BASE_URL/v1/models/systemone" \
-H "Authorization: Bearer $DURINDOOR_API_KEY"Web search and fetch
Provider is the model. Send provider or model. Catalog ids look like tavily/search and tavily/fetch; the handler strips the suffix when the remainder is a real search or fetch provider. Bare ids and OmniRoute aliases (tavily-search, exa-search, serper-search, google-pse-search, linkup-search, searchapi-search, youcom-search, searxng-search, ollama-search, perplexity-search) also resolve.
Search body: query (required), plus model or provider. Fetch body: url (required), optional format and max_characters. Fetch failures use fallback scope webfetch:<provider> (for Ollama, webfetch:ollama), so a fetch cooldown does not lock that connection for chat.
Search searchConfig providers in the registry today: serper, youcom, searxng, linkup, exa, perplexity, glm, brave-search, searchapi, google-pse, tavily, ollama. Fetch fetchConfig providers: exa, tinyfish, jina-reader, tavily, firecrawl, firecrawl_custom, ollama. Combos work on both routes when every member is that kind.
| Method | Path | Request | Stream | Providers |
|---|---|---|---|---|
| POST, OPTIONS | /v1/search | JSON query plus model or provider | no | searchConfig / searchViaChat providers and combos |
| POST, OPTIONS | /v1/web/fetch | JSON url plus model or provider; optional format, max_characters | no | fetchConfig providers and combos |
Files and batches
Local store under the data directory (files/<id>), not an upstream Files API. Owner is the DurinDoor API key (or local / operator when keys are not required). JSON POST that is not application/json returns 415. File upload that is not multipart/form-data returns 415.
OpenAI batches: { input_file_id, endpoint, completion_window?, metadata? }. Anthropic batches: { requests: [{ custom_id, params }] }. Cancel stops scheduling, finishes the active row, then marks cancelled. Anthropic results are NDJSON { custom_id, result }.
| Method | Path | Request | Stream | Providers |
|---|---|---|---|---|
| GET, POST, HEAD, OPTIONS | /v1/files | POST multipart file, purpose (default batch) | no | local disk |
| GET, DELETE, HEAD, OPTIONS | /v1/files/{id} | none | no | local disk |
| GET, OPTIONS | /v1/files/{id}/content | none | raw bytes | local disk |
| GET, POST, HEAD, OPTIONS | /v1/batches | POST JSON input_file_id, endpoint | no | local executor over chat |
| GET, OPTIONS | /v1/batches/{id} | none | no | local |
| POST, OPTIONS | /v1/batches/{id}/cancel | none | no | local |
| GET, POST, HEAD, OPTIONS | /v1/messages/batches | POST JSON requests[] | no | local executor, Anthropic surface |
| GET, OPTIONS | /v1/messages/batches/{id} | none | no | local |
| GET, OPTIONS | /v1/messages/batches/{id}/results | none | NDJSON | local |
| POST, OPTIONS | /v1/messages/batches/{id}/cancel | none | no | local |
Realtime auth
GET /v1/realtime upgrades to a WebSocket. /v1/realtime/auth checks gateway credentials; /v1/models does not validate keys.
GET /v1/realtime/auth: 200 { ok: true } when the key is admitted (or not required). 401 { error: { message, type: "invalid_request_error", code: "invalid_api_key" } } when missing (and required), invalid, or expired. Authenticate the socket with Authorization: Bearer or the openai-insecure-api-key.<key> subprotocol. The chat-emulation WebSocket is text-only; native realtime models use their declared audio protocol. Event list: Realtime.
| Method | Path | Request | Stream | Providers |
|---|---|---|---|---|
| GET, OPTIONS | /v1/realtime/auth | same key headers as chat | no | none (auth probe) |
Native session creation
These JSON endpoints delegate directly to OpenAI's native session API. Supply a registered gateway model through ?model=openai/model-id or the operation's model field, including nested session or transcription fields where the vendor schema requires them. Other body fields follow the vendor schema.
| Method | Path | Purpose |
|---|---|---|
| POST, OPTIONS | /v1/realtime/client_secrets | Create a direct native realtime ephemeral credential. |
| POST, OPTIONS | /v1/realtime/translations/client_secrets | Create a direct translation ephemeral credential. |
| POST, OPTIONS | /v1/realtime/transcription_sessions | Create a direct transcription session. |
| POST, OPTIONS | /v1/live/sessions | Create a direct Live session. |
Direct sessions require an uncapped key and retain model/account policy checks. A usage-capped key returns 403 Usage-capped API keys cannot mint direct native sessions. Gateway-relayed native WebSockets support scoped keys through Realtime.
Gemini v1beta
/v1beta/:path* rewrites to /api/v1beta/:path*. These routes let a Gemini-native client, such as Gemini CLI or the @google/genai SDK, talk to DurinDoor without an OpenAI shim.
GET /v1beta/models returns a Gemini-style { models: [...] } list built from the static catalog. Every catalog row appears as models/{provider}/{id}. Rows from the gemini provider also appear as plain models/{id} with streamGenerateContent in supportedGenerationMethods. The token limits in that list are fixed placeholders (128000 in, 8192 out), not the real model limits. The list handler does not check a key.
POST /v1beta/models/{model}:generateContent and :streamGenerateContent accept a Gemini request body. {model} is either a bare Gemini id or {provider}/{model}. The URL suffix decides streaming: :streamGenerateContent streams Gemini SSE, :generateContent returns one JSON GenerateContentResponse. The body is converted to the internal chat shape and runs through the same chat core as /v1/chat/completions, so key checks, combos, and fallback behave the same way. inlineData parts (images, PDFs, audio) become data-URL image_url parts in that shape. Parts sent next to a functionResponse in the same content, such as the file a Gemini CLI read_file tool returns, follow the tool message as a user message.
A Gemini text-to-speech request (a Gemini TTS model id, or responseModalities asking for audio) skips that conversion. It is forwarded to Google's generativelanguage.googleapis.com with a saved Gemini credential, after the usual client-key and model-policy checks.
curl "http://localhost:20128/v1beta/models/gemini/gemini-2.5-flash:generateContent" \
-H "x-goog-api-key: YOUR_DURINDOOR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"contents":[{"role":"user","parts":[{"text":"Say hello."}]}]}'| Method | Path | Request | Stream | Providers |
|---|---|---|---|---|
| GET, OPTIONS | /v1beta/models | none | no | static catalog |
| POST, OPTIONS | /v1beta/models/{model}:generateContent | Gemini contents body | no | any chat model; Gemini TTS goes to Google |
| POST, OPTIONS | /v1beta/models/{model}:streamGenerateContent | Gemini contents body | yes | any chat model; Gemini TTS goes to Google |
Unknown paths
Unmatched /v1/* returns JSON { error: { type: "not_found", code: "unknown_route", path } } with HTTP 404, not the dashboard HTML shell. Registered routes take precedence over that fallback.
| Method | Path | Request | Stream | Providers |
|---|---|---|---|---|
| GET, POST, PUT, PATCH, DELETE, HEAD, OPTIONS | unmatched /v1/* | none | no | none |
Management API
Dashboard REST under /api/* is on Management API. The same DurinDoor API key that calls /v1 can call most of those routes; exclusions, redaction, and curl examples live there. MCP tools for the same work: MCP control.
See MCP gateway and compression for gateway features that are not /v1 routes.