Skip to content
ThinkWatch
Lite 2026.9.26 released→

A local gateway for Claude Code, Codex and other AI clients

Each client is connected once; after that, upstreams and models change without touching its configuration. Every request is recorded with its cost and route; the API keys in it can be replaced before it leaves the machine, and dangerous tool calls a relay slips into an answer can be cut off before the client runs them. For macOS, Windows and Linux, under the MIT License.

ThinkWatch Lite · live traffic
Available
Platforms
macOS 12 or later, Apple silicon
Windows 10 21H2 or later, x64 or ARM64
Linux, x86_64 or aarch64
Interface
English and Simplified Chinese
Updates
Installed by the app itself; through Homebrew for a Homebrew installation
License
MIT, free and open source

Set up in one step, or by following the instructions

  • Claude Code
  • Codex
  • Claude Desktop
  • opencode
  • Cursor
  • Zed
  • Aider
  • DeepSeek Harness
  • Continue
  • Antigravity CLI

One request, four stops through the gateway.

01Routing

Routing by rule, with failover

Rules send requests to different upstreams by model, tools, images, extended thinking and more. When an upstream fails before the answer begins, the next one takes over, and each session stays on one upstream so its prompt cache keeps hitting. Auxiliary requests such as title generation and warm-ups can be answered locally without using any quota.

The Routing page: a map from keys through routes and groups to upstreams, and the routes with the rules each applies in order

02Protection

Protection against relays: keys replaced, malicious tool calls cut off

A relay sees every request in full and can rewrite every answer. Outbound redaction swaps API keys, private keys, JWTs and connection-string passwords for placeholders before a request leaves and restores them in the response, so the relay never holds the real values. When an answer carries a tool call that downloads and runs code, sends out environment variables or credential files, reads private keys or installs a startup item or scheduled job, tool-call inspection cuts the answer off before the client can run it; hidden characters and prompt injection can be refused as well. The five protections start in Observe, recording without changing anything, and each switches to Enforce on its own.

The Security page log: credentials replaced before a request left, one of them matched by a custom rule; a download-and-run tool call cut off; and hidden characters, a delete command and an injected instruction recorded, each with the key, client, model and upstream of its request

03Tracing

Every request, traceable

A request shows the rule it matched, each upstream it tried, any conversion between API formats and how its cost was calculated. A finished request can be replayed against another upstream and the two answers compared side by side.

The Traffic page: each request with its key, model, upstream, time to first token, total time, tokens and cost, with marks for converted formats, redacted keys and a blocked request, and one request answered locally by the gateway

04Cost

Costs stated as they are

Tokens, cost, cache savings, time to first token and generation speed, by model and by upstream. Estimated amounts are marked, and requests without a price are counted separately instead of as zero; prices follow LiteLLM's public list, refreshed daily, or a custom price sheet.

The Overview page for the last 7 days: 83.1M tokens, $68.11 in cost including $0.441 estimated and 13 unpriced requests, and 1,339 requests of which 11 failed, each compared with the prior 7 days; a token trend stacked by model with the periods that had failures marked; and the models ranked by tokens

And the rest, all in one app.

Menu bar, tray and notifications

The macOS menu bar shows today's tokens and cost, in orange when a subscription quota is nearly used up and red once it has run out. Its menu gives the gateway's state, each quota with its reset time and the requests in progress, and copies the gateway address or default key without opening the window. Windows and Linux have the same menu in the tray.

7

set up in one step · 3 more by instructions

Connect once, switch freely

Claude Code, Codex, opencode and four other clients are pointed at the gateway in one step, with the change previewed, the original file backed up and a restore always available; on Windows, Claude Code and Codex inside WSL as well. From then on, switching upstreams happens in the gateway, with no client to reconfigure or restart.

Any upstream, any API format

API keys, Amazon Bedrock, ChatGPT and Z.ai accounts, relays such as OpenRouter and local Ollama models all serve as upstreams, with subscription quotas and reset times shown. Requests are converted between the Anthropic, OpenAI and Gemini APIs, so Codex can also use models that only speak Chat Completions.

The MCP page: the MCP servers configured in Claude Code, Claude Desktop, Cursor, Codex, opencode, Antigravity CLI and Zed side by side, with remote third-party servers and a server configured differently in two clients marked; one high, one medium and one low finding in 11 scanned files

MCP servers, skills and hooks, scanned

The MCP servers of eight clients appear side by side, with third-party remote servers and inconsistent configurations marked, and can be copied or removed between clients. Client configuration, skills, hooks and project instructions are scanned for hidden characters, prompt injection, dangerous commands and overly broad permissions, and a new finding raises a notification.

A key for each client

Connecting a client gives it a key of its own, so traffic and cost are counted per client. Each key has its own route, visible models and concurrency limit, and a rotated key is written into its client's configuration.

A routing dry run: a request from the cursor key for claude-sonnet-5 in the OpenAI Chat Completions format does not match the gemini rule, which says why, matches the catch-all rule and goes to the lowest-cost group, which tries relay and then anthropic, converting the request to Anthropic Messages

Dry run before changing rules

A dry run shows which rule a request would match, why the rules before it did not, and which upstreams would be tried in turn. Nothing is sent and nothing is charged.

System notifications

A system notification reports a gateway that stopped forwarding, a lost remote connection, a quota that ran out, a credential that stopped working, an unreachable proxy, configuration that did not take effect, a dangerous tool call that was cut off, and suspicious content in client configuration.

Remote core

ThinkWatch Core on a server

The gateway also runs on a Linux server as a systemd service, serving clients across the network. The app connects to it and shows the server's traffic, cost and configuration in the same pages, and the clients on the computer can be pointed at the server's gateway in one step.

The control connection is encrypted and authenticated by a Noise handshake, with no certificates involved. The app connects to one core at a time, and the local data stays as it was while it is connected elsewhere.

Adding a remote connection: the name homelab, the address 192.168.1.40, the control port and the key, with a successful test that reports the core version and the gateway at 192.168.1.40:8788
The connection menu at the foot of the sidebar while connected to homelab: this Mac with its local core stopped, homelab selected, build-server, and entries to add or manage connections
Connections are switched from the foot of the sidebar.
The Clients page while connected to homelab, with a note that it changes the client configuration on this Mac to point at homelab's gateway at 192.168.1.40:8788 and does not affect clients on the server
The Clients page changes the clients on this computer.

Architecture

The app is the interface; the gateway is ThinkWatch Core

Lite holds no routing, forwarding or accounting logic. All of it is in ThinkWatch Core, which runs beside the app or on a Linux server.

Desktop app

ThinkWatch Lite

  • Main window, menu bar and tray
  • Starts and supervises the local core
  • Or connects to a core on a server

Gateway

ThinkWatch Core

  • Routing, failover and format conversion
  • Cost accounting and request records
  • The five security protections

Upstreams

Model services

  • Anthropic, OpenAI and Gemini APIs
  • Amazon Bedrock and signed-in accounts
  • Relays and local models

The control channel is a Unix socket on macOS and Linux, a loopback port on Windows and a TCP port to a core on a server; every connection is encrypted and authenticated by a Noise handshake.

Install

Download ThinkWatch Lite

The gateway ships inside the app; nothing else needs to be installed.

Requires macOS 12 or later on Apple silicon.

Homebrew

Recommended

Homebrew installs the app and keeps it up to date.

$ brew install --cask thinkwatchproject/tap/thinkwatch-lite

Disk image

Download, open and drag the app into Applications. The app then updates itself.

Download the disk image

Version 2026.9.26 · sha256

The app is not notarized by Apple. When it is installed from the disk image and macOS blocks the first launch, choose Open Anyway in System Settings › Privacy & Security. The Homebrew install needs no such step.

Requires Windows 10 21H2 or later on x64 or ARM64.

Installer

Installs the app for all users and keeps it up to date.

Version 2026.9.26 · sha256 (x64) · sha256 (ARM64)

The installer is not code-signed. When SmartScreen shows “Windows protected your PC”, choose More info, then Run anyway. Installing needs administrator permission, and WebView2 is downloaded if it is missing.

Requires Ubuntu 22.04, Debian 12, Fedora 36 or a later distribution, on x86_64 or aarch64.

Install script

Recommended

Downloads the AppImage for this machine, checks its sha256, places it in ~/Applications and starts it.

$ curl -fsSL https://github.com/ThinkWatchProject/ThinkWatch-Lite/releases/latest/download/install.sh | sh

AppImage

Download, allow it to run (chmod +x) and open it.

Version 2026.9.26 · sha256 (x86_64) · sha256 (aarch64)

Needs fusermount3 from the fuse3 package, which most desktops include. Keep a downloaded AppImage in a folder the user can write to, such as ~/Applications, so that it can update itself. On GNOME the tray icon needs the AppIndicator extension.

MIT License

Use, modification, and redistribution are permitted.

View on GitHub