01Routing
Routing by rule, with failover
Rules send requests to different upstreams by model, tools, images, extended thinking and more. When an upstream fails before the answer begins, the next one takes over, and each session stays on one upstream so its prompt cache keeps hitting. Auxiliary requests such as title generation and warm-ups can be answered locally without using any quota.
02Protection
Protection against relays: keys replaced, malicious tool calls cut off
A relay sees every request in full and can rewrite every answer. Outbound redaction swaps API keys, private keys, JWTs and connection-string passwords for placeholders before a request leaves and restores them in the response, so the relay never holds the real values. When an answer carries a tool call that downloads and runs code, sends out environment variables or credential files, reads private keys or installs a startup item or scheduled job, tool-call inspection cuts the answer off before the client can run it; hidden characters and prompt injection can be refused as well. The five protections start in Observe, recording without changing anything, and each switches to Enforce on its own.
03Tracing
Every request, traceable
A request shows the rule it matched, each upstream it tried, any conversion between API formats and how its cost was calculated. A finished request can be replayed against another upstream and the two answers compared side by side.
04Cost
Costs stated as they are
Tokens, cost, cache savings, time to first token and generation speed, by model and by upstream. Estimated amounts are marked, and requests without a price are counted separately instead of as zero; prices follow LiteLLM's public list, refreshed daily, or a custom price sheet.