Compression
Preview token savers, enable a supported engine, and compare results with a per-request bypass.
Use token savers to reduce eligible outbound prompt content. Reductions depend on the payload and engine. Errors keep the inference request running with the original content for the failed step.
Preview before enabling an engine
- Open Dashboard → Test Savers (
/dashboard/compression-studio). - Paste a representative JSON chat payload without secrets.
- Run the preview and compare the output with your input.
- Check that instructions, tool identities, and error results still convey the information the client needs.
Preview uses the dashboard session and does not call your inference provider. Results report compressed, unchanged, unavailable, or error. An unavailable engine is a catalog placeholder and cannot run in the compression stack.
Enable token savers
Open Dashboard → Token Saver → Settings (/dashboard/token-saver/settings). RTK is on by default. Headroom and the compression stack are off by default. Enable one change at a time, then send a representative request and inspect Token Saver → Statistics.
| Option | Behavior |
|---|---|
| RTK | Compress eligible tool-result content. Skip explicit error results. |
| Headroom | Call the optional Python proxy. Follow Headroom setup. |
| Session dedup | Remove eligible repeated blocks across turns in the compression stack. |
| Caveman | Inject the selected output-style instruction into eligible text-only requests; settings preview shows the exact instruction block. |
| Ponytail | Apply its configured prompt treatment and slash commands. |
| PXPIPE | Apply the bundled image transform where eligible. |
Stack engines currently available are session-dedup, caveman, and headroom. Catalog entries rtk, lite, ccr, relevance, aggressive, llmlingua, and ultra are not dispatched as stack engines. The dedicated RTK toggle still works.
Headroom processes the source-format chat body before translation. RTK normally processes the translated body; Cursor needs RTK before its translator rewrites tool results. The compression stack follows those steps, then the separate Caveman, Ponytail, and PXPIPE toggles apply. Enabling overlapping options can apply more than one transform, so compare results before using a stack on important traffic.
Bypass compression for one request
Send:
X-DurinDoor-Token-Saver: offThis disables all token savers and Ponytail slash-command interception for that chat request. The value off is case-insensitive. The legacy X-9Router-Token-Saver header is still accepted, but the DurinDoor header takes precedence whenever present.
Verify the result
Send a chat history with a substantial successful tool result. Compare Statistics before and after the request. Then repeat with the bypass header and compare the output and statistics.
When a compression-stack engine changes the request and reports compression, a successful chat response includes:
X-DurinDoor-Compression: <engineId>[,<engineId>...]|<overallSavingsPercent>%The percentage compares the estimated stack input and final output. It is not a measured provider bill or a sum of engine percentages. Dedicated RTK savings can appear in Statistics without this stack header. Disabled, unchanged, failed, or upstream-error stack runs omit it.
Recover from an unexpected result
Disable the last engine you enabled, or use the bypass header to compare a request without token savers. Use Test Savers to isolate the changed content before restoring the option.
PXPIPE ships with the app. If status is DEPENDENCY_MISSING, reinstall the app instead of looking for an install endpoint. Dashboard → PXPIPE lists transform events. PXPIPE management requires a dashboard session or machine-bound CLI token; an inference API key does not grant those controls.