Changelog
ThinkWatch release notes
Release notes for each version of ThinkWatch Enterprise, ThinkWatch Lite and ThinkWatch Core, listed from newest to oldest. Subscribe via RSS.
Latest releases
September 2026
-
ThinkWatch Enterprise v2.2.0 GitHub ↗
This release fixes authorization. The gateways never checked
ai_gateway:useormcp_gateway:use, so a user with no role, or with only roles that grant neither, could call every model and MCP tool; a role granted at the scope of one team widened model and tool access platform-wide; and an API key whose allow-list came out empty could call any model. The gateways now require the permission and count only the roles that grant it. The seededteam_managerrole now works when granted at team scope, and seven permissions that nothing ever checked are retired. Some users lose gateway access on upgrade: read the first section before deploying. There is one manual database script and no schema, setting, environment variable or Helm value change.Read before upgrading#
-
Users without a role that grants gateway use lose gateway access. A request to the AI gateway is refused with
403unless a role held by the key’s owner grantsai_gateway:use, and a request to the MCP gateway unless one grantsmcp_gateway:use. The body names the missing permission, in the error format of the caller’s API. The built-indeveloper,admin,super_adminandteam_managerroles grant both;viewergrants neither. Three groups of users are refused after the upgrade where 2.1.0 let them through:- users with no role at all;
- users whose roles are only
viewer, or custom roles without these actions; - users whose only gateway role is assigned at team scope (see the third item below).
Find them before upgrading. This query lists every active user with a live key for a gateway that none of their roles will grant (it counts global assignments and roles attached to the user’s teams, which is what the gateways now read; it does not account for
Denystatements):WITH grants AS ( SELECT id AS role_id, jsonb_path_exists(policy_document, '$.Statement[*] ? (@.Effect == "Allow").Action[*] ? (@ == "*" || @ == "ai_gateway:*" || @ == "ai_gateway:use")') AS ai, jsonb_path_exists(policy_document, '$.Statement[*] ? (@.Effect == "Allow").Action[*] ? (@ == "*" || @ == "mcp_gateway:*" || @ == "mcp_gateway:use")') AS mcp FROM rbac_roles ), held AS ( SELECT user_id, role_id FROM rbac_role_assignments WHERE scope_kind = 'global' UNION SELECT tm.user_id, tra.role_id FROM team_members tm JOIN team_role_assignments tra USING (team_id) ), access AS ( SELECT h.user_id, bool_or(g.ai) AS ai, bool_or(g.mcp) AS mcp FROM held h JOIN grants g USING (role_id) GROUP BY h.user_id ) SELECT u.email, s.surface FROM api_keys k JOIN users u ON u.id = k.user_id CROSS JOIN LATERAL unnest(k.surfaces) AS s(surface) LEFT JOIN access a ON a.user_id = u.id WHERE k.is_active AND k.deleted_at IS NULL AND u.is_active AND u.deleted_at IS NULL AND NOT COALESCE(CASE s.surface WHEN 'ai_gateway' THEN a.ai WHEN 'mcp_gateway' THEN a.mcp END, false) GROUP BY u.email, s.surface ORDER BY u.email, s.surface;Give each of them
developer(or a custom role with the actions) either at global scope (Users → edit the user’s roles, scope Global), or by attaching the role to a team they belong to (Teams → the team → Roles), which every member of that team inherits. The change takes effect within a minute; each user’s permissions are cached for 60 seconds.If SSO users are meant to use the gateway from their first sign-in, set Default Role for New Users (
auth.default_role) in Settings todeveloper. It is empty by default, and it applies only to accounts created after it is set; existing users need the grant above. -
A role that does not grant gateway use no longer widens model or tool access. Model and MCP tool scopes are now the union over the roles that grant
ai_gateway:use/mcp_gateway:useonly. A role without those actions (such asviewer) used to count as “unrestricted”, so adding it to a user limited to some models opened every model to them. Users relying on that lose the extra models. -
A role assigned at team scope no longer grants gateway access. An assignment with scope
team:<id>administers that team from the console (team roster, team limits); it no longer contributes models, MCP tools or gateway use, which it used to do for every request the user made, member of the team or not. Roles attached to a team itself (Teams → Roles), which every member inherits, still count. A user whose only gateway role was a team-scopeddeveloperorteam_managerneeds that role at global scope, or attached to their team. The role assignment editor now says this under the scope picker. -
An empty model allow-list allows nothing. A key whose
allowed_modelsis[], or whose list shares no model with what its owner’s roles grant, used to call any model; it now calls none, asallowed_mcp_tools: []already did on the MCP gateway. The console never saves[](clearing the picker sendsnull), so only keys written through the API are affected. To find them:SELECT id, name, user_id FROM api_keys WHERE allowed_models = '{}' AND is_active AND deleted_at IS NULL;Set such a key’s list tonullto leave it bounded by its owner’s roles only. Model entries still match by prefix, and a key narrowed togpt-4o-miniunder a role grantinggpt-4okeepsgpt-4o-mini. -
Run
db/release_migrations/2026-09-30_retire_unchecked_permissions.sqlonce, after deploying. The server does not apply it. It addsteams:readto theteam_managerrole and removes the seven retired permissions (see Changed) from every role, system and custom. It is idempotent and changes no access for any other role, since nothing checked the removed keys. Without it, the server still starts and works, but:team_managerkeeps the oldteam:readit cannot use, so a team manager who is not a member of the team they manage still gets403opening it, its roster or its roles;- every start logs a warning listing the roles that still name retired permissions.
Reset to defaults on a system role in the console has the same effect for that one role, but does not clean custom roles.
-
A team-scoped
rate_limits:writenow takes effect. It used to require global scope for every subject, so the seededteam_managergranted at team scope could not touch any limit. It now covers the rate limits and budget caps of users in that team and of the API keys they own, never the holder’s own user or keys; role subjects still need global scope. Anyone holdingteam_manager(or a custom role withrate_limits:write) at team scope can now change, lift or delete the limits of every member of that team, administrators included. Review team-scoped assignments of these roles before upgrading.
Security#
- Gateway use is checked, and a missing grant means no access. See
the first three items above: a user with no role, or only roles
without gateway use, could call every model and MCP tool; a
viewer-style role widened a restricted user to every model; a role granted at the scope of any team, even one the user was not a member of, widened model and tool access platform-wide. An explicitDenyonai_gateway:useormcp_gateway:usenow closes that gateway. Action wildcards (ai_gateway:*,*) grant it, as they already did for console permissions. - A key narrowed inside a prefix grant is no longer unrestricted.
The key’s allow-list was intersected with its owner’s role grants
entry by entry, as literal strings. A key limited to
gpt-4o-miniunder a role grantinggpt-4o(a prefix) came out with an empty list, which the gateway read as “no restriction”, so the key could call every model. The intersection now keeps an entry of either side that the other side covers, by prefix for models and by<server>__*pattern for MCP tools, and an empty result allows nothing.
Fixed#
- The seeded
team_managerrole works at team scope. It grantedteam:read, but the team handlers checkteams:read, so a team manager got403listing teams and opening the team, its roster or its roles. It now grantsteams:read, and opening a team acceptsteams:readscoped to that team, where it used to require global scope. Existing installations need the release migration above. - Bulk disable and delete work on API keys’ limits. The bulk
disable and delete routes for rate-limit rules and budget caps
rejected every row stored for an API key (
api_key_lineage, the kind a key’s limits are stored under) as an unknown subject kind. - The console’s effective-permissions preview shows real gateway access. It now counts only roles that grant gateway use and skips team-scoped assignments, as the gateways do, where it used to count every assigned role.
Changed#
- Seven permissions that nothing checked are retired:
team:read,team:write,logs:read_own,logs:read_team,audit_logs:read_own,audit_logs:read_teamandaudit_logs:read_all. Every log endpoint, audit logs included, is gated onlogs:read_allat global scope, and no own- or team-filtered log view exists, so these grants never did anything. They are gone from the permission catalog, the role editor and the seeded roles. Roles that still name them load, and the server logs a warning at start instead of refusing to boot. - Deleting one rate-limit rule or budget cap is bound to the subject
in the path.
DELETEon/api/admin/limits/{kind}/{id}/rules/{rule_id}and/api/admin/limits/{kind}/{id}/budgets/{cap_id}now answers404when the rule or cap belongs to a different subject; it used to delete any row id once the path’s subject was authorized. - Creating or editing a key with an explicit allow-list checks it
against gateway-granting roles only. An
allowed_modelsorallowed_mcp_toolsentry is refused with400when no role of the owner that grants the gateway covers it, so a user without a gateway role can no longer save a non-empty list.
-
-
ThinkWatch Enterprise v2.1.0 GitHub ↗
Amazon Bedrock becomes a provider you can run from the console. It authenticates with a Bedrock API key as well as access keys or the instance role, lists its models for import and for the route editor, and has a working Test Connection. Requests converted for Bedrock now keep their prompt-cache breakpoints and Claude’s thinking settings. The route editor takes any upstream model name, and a route the provider refused can be created anyway. ThinkWatch-Core moves to v0.55.0, whose Bedrock layer this edition now shares with the desktop gateway. No database, setting, environment variable or Helm value changes.
Read before upgrading#
- Test Connection with a saved provider’s secrets needs
providers:update. A test that names a saved provider (provider_id, which is what the Edit dialog sends) used to take onlyproviders:create. A custom role that hasproviders:createwithoutproviders:updatecan no longer test existing providers; the built-in admin role has both. A test with values typed into the request still takesproviders:create. See Security below. - Refusing a route the provider does not serve answers
model_not_served.POST /api/admin/models/{model_id}/routesstill answers400when the import probe found no API the provider serves the model on, buterror.typeis nowmodel_not_served, where it wasbad_request. Scripts that matched onbad_requestfor this case need the new value. The request takes a new optional"force": trueto create the route anyway. error_typeingateway_logsfor failed requests is always the error’s tag. An upstream HTTP error used to log its debug text, such asProviderHttpError { status: 502, message: "…" }, and an upstream rate limit a cut-offUpstreamRateLimited { retry_after_secs: Some. They now logProviderHttpErrorandUpstreamRateLimited, as the metric labels and streamed requests already did. Dashboards or log forwarder queries that grouped by those strings see them merge into one value each.- The server log now carries the upstream’s reply to a 401 or 403,
as it already did for other upstream errors. For Bedrock that reply
names the AWS account and the IAM principal. Callers still see only
“Authentication failed with upstream”, and failover, breakers and
gateway_logstreat these errors as before. - Prompt caching now works on Bedrock routes, and is billed at cache
prices. Requests converted for Bedrock used to lose their
cache_controlbreakpoints, so Claude on Bedrock re-read the whole prompt at full price every turn. Breakpoints now go out as ConversecachePointblocks, to Claude and Nova models only (others reject them), and the cache reads and writes Bedrock reports come back into usage, one-hour writes included. Traffic with repeated prompts, such as Claude Code, costs much less on Bedrock routes than it did; writes cost a little more, atcache_write_weight/cache_write_1h_weight. - Claude’s thinking reaches Bedrock. A request converted for Claude
on Bedrock now carries its thinking and effort settings, in the form
Anthropic’s API uses, along with the sampling limits thinking imposes
(
temperature1, notop_k,top_pat least 0.95). Thinking used to be dropped on the way to Bedrock. Other Bedrock models still get no thinking.
Added#
- Bedrock providers can authenticate with a Bedrock API key. Pick
Bedrock API Key as the authentication mode and paste the key AWS
generated. It is kept like any other provider’s API key, as an
encrypted
Authorization: Bearerheader, and replaced in the provider’s Edit dialog. A Bedrock provider that sends anAuthorizationheader is not SigV4-signed; one without it is signed as before, with access keys or the instance role. Use a long-term key: a short-term one expires within 12 hours. - Bedrock models can be imported, and are suggested in the route
editor. A Bedrock provider now lists what it can be routed to: the
region’s foundation models that can be invoked on demand and answer in
text, and the inference profiles AWS defines, such as
us.anthropic.claude-sonnet-4-5-20250929-v1:0. A model that is only served through an inference profile, as most current models are, is listed under its profiles’ ids and not its own. Whether the account may call a model is still checked one model at a time when it is imported. The provider’s credential needsbedrock:ListFoundationModelsandbedrock:ListInferenceProfiles; theAmazonBedrockLimitedAccesspolicy a long-term API key is created with allows both. - Test Connection for Bedrock providers, however they authenticate: with an API key, access keys or the instance role. It lists the models above, so a wrong credential or a missing permission shows up before any traffic does.
- Create anyway, for a route the provider refused. When the route
editor’s save is refused because the provider does not serve the
model, the error now offers Create anyway: the refusal can be
stale, or be about the probe’s request rather than the model. The
route is created, and its
model_route.createdaudit row records the refusal it overrode in a newrefusal_overriddenkey.
Changed#
- The route editor’s upstream model field takes any model name. For a provider that lists its models, the field used to turn into a list to pick from, so a model the listing leaves out could not be routed to from the console. It now suggests the provider’s models as you type, and takes whatever is typed.
- Bedrock instance-role credentials are cached. They took three IMDSv2 round trips per request; they are now kept until five minutes before they expire, and a burst of requests at expiry makes one trip.
- Chat Completions requests can switch reasoning with
thinking. When a request is converted for another format,thinking.type(enabled/disabled, as DeepSeek, GLM and Kimi write it) now turns reasoning on or off;disabledwins overreasoning_effort. - Core crates at ThinkWatch-Core v0.55.0.
tw-dialect,tw-guardandtw-breakermove from v0.43.0, andtw-bedrockjoins them: SigV4 signing, eventstream unframing, Bedrock’s addresses, region checks and model catalog now come from core. Where credentials come from (provider keys, the instance role) stays in this repository.tw-breakeris unchanged.
Fixed#
- Bedrock providers could not be created or edited in the console.
The region was checked as a URL, so saving failed with
400 Invalid URL. A Bedrock provider’sbase_urlis now checked as an AWS region such asus-east-1, and anything else is refused, since the host is built from it. A provider saved with a URL there never reached Bedrock; set its region in the Edit dialog. - Bedrock models the account may not call were imported anyway. Bedrock refuses such a model (model access not granted, or an IAM or organization policy that denies it) with a 403, and the import probe read every 403 as a credential problem that says nothing about the model. The route was created and failed on first use. Now, when the region’s control plane accepts the same credential, the refusal is recorded as the model’s, with AWS’s reason, and the model is skipped on import like any other refused one. Once access is granted, re-check the provider’s models.
- A Bedrock provider whose stored secret key would not decrypt signed
with an empty secret, and every request failed with AWS’s
SignatureDoesNotMatch. It now refuses its requests with the reason until the keys are saved again, rather than falling back to the instance role, which would call AWS as a different identity. An access key ID saved without a secret is refused the same way. - A Bedrock model id that is an ARN could not be routed. Its
/went into the request path unescaped, adding a path segment AWS could not route. It is now escaped. - A base URL ending in its version segment doubled it. A provider
written as
https://api.openai.com/v1sent requests, and the model listing, to/v1/v1/…and got 404s. A trailing version segment (v1,v1beta, …) that the request path starts with is now written once. Base URLs without one are unchanged. - Typing into a provider’s API key field and clearing it again wiped
the saved key on save. The field sent
Bearerwith nothing after it. A cleared field now keeps the saved key, as a blank one always did.
Security#
- Testing a connection with a saved provider’s secrets takes
providers:update. The test fills each header left blank with the saved provider’s secret and sends it to the URL in the request, andproviders:createwas enough to ask for it. So a user allowed only to create providers could send any saved API key to a server of their own. A test that names a saved provider (provider_id) now takesproviders:update, which lets a user point that provider elsewhere anyway; a test without one still takesproviders:create. A user withproviders:updatealone can now use Test Connection in the Edit dialog, which used to answer403.
- Test Connection with a saved provider’s secrets needs
-
ThinkWatch Lite v2026.9.26 GitHub ↗
Upgrade notes:
- The bundled core is still 0.56.0 (control-plane protocol 30), so a remote server on 0.56.0 keeps working.
- Configuration and request history are kept.
MCP page: the servers table now has a column only for the clients found on this computer: those whose own directory or MCP file exists, or whose MCP file was chosen by hand. It used to draw all eight supported clients, which did not fit the window. The clients that were not detected are listed in one line under the table, where Choose location… points the app at a client installed somewhere else.
Toolbar: the page name is always shown in the window toolbar, next to the sidebar toggle, instead of at the top of each page.
-
ThinkWatch Lite v2026.9.25 GitHub ↗
Upgrade notes:
- The bundled core is now 0.56.0, which speaks control-plane protocol 30. A remote server has to run core 0.56.0 too: this version does not connect to 0.55.x, and 2026.9.24 and earlier do not connect to 0.56.0.
- Request history is kept.
- Strategy groups no longer have the sticky-session switch. A configuration in which it was turned off contains
session_affinity: falseunder that group; core 0.56.0 refuses such a configuration and the app starts in safe mode, naming the line. Delete the line.
Failover:
- An upstream can answer, open a stream and then report an error before sending anything, such as Anthropic’s
overloaded_erroror aresponse.failedwhen a ChatGPT account’s usage limit is reached. The client used to receive that error. A streamed answer is now held until its first content arrives, and an error before that point sends the request to the next upstream of the route. What was received while waiting is passed on unchanged. - Upstreams that reject the credential (401, 403), report no balance (402) or do not have the model (404) now send the request to the next upstream as well. A 400 or 422 does so only when the upstream says the balance or quota is used up or the model is unavailable; other client errors still go straight back. When no upstream is left, the last upstream’s own error reaches the client.
- How long a failing upstream is paused depends on the reason it gives: an insufficient balance pauses it for 30 minutes, a used-up quota until the reset time the upstream names (1 hour when it names none), a rate limit for the wait it asks for (at most 1 hour). Failures without a stated reason pause it after 3 in a row, for 1 minute, doubling with each further pause up to 10 minutes.
- Settings › Failover sets all of these, including how long to wait for a streamed answer’s first content (15 seconds).
Prompt cache across a conversation:
- A conversation now stays with the upstream that answered it, including one that took over after a failover: always within a turn, and in later turns while its prompt cache is still warm. Load-balance groups spread only new conversations.
- The routing decision made at the start of a turn holds until the turn ends, so a rule on input size or images no longer switches upstreams halfway through an agent’s work.
- The request details’ Routing tab says when a request kept its turn’s route or stayed with the upstream that answered before.
Import links: a relay or vendor can hand out a link,
thinkwatch://import?…orhttps://thinkwat.ch/import#…, that pre-fills a new upstream with a name, base URL, protocol, API key and model list. The app shows the settings and the host that will receive requests and the key, and writes nothing and contacts nothing until Create is chosen. A link only adds one upstream: it cannot change existing upstreams, headers, proxies, pricing or routes, and a key that refers to an environment variable is rejected. The parameters and a link builder are in Import links.Token counts: Claude Code asks for a token count on every turn, which only an upstream in Anthropic’s format can answer. Routed to an upstream of another format, or to a relay that does not implement it, the count used to fail; the gateway now answers with its own estimate, sends nothing to that upstream, and shows it in the attempt chain as estimated locally.
Reasoning after a switch of account: reasoning that one account produced is sealed for that account, and another account or key refuses a request that carries it back. After a failover between accounts such a request is now sent again without it, and later turns of the conversation leave it out from the start.
-
ThinkWatch Core v0.56.0 GitHub ↗
This release keeps a conversation on the upstream that holds its prompt cache and decides a turn’s route once, moves a request to the next upstream when a stream fails before its first content, sets an upstream aside for as long as the reason it gives calls for, answers token counts that no upstream can serve with a local estimate, and resends a request whose reasoning another account sealed.
Upgrade notes
- The control-plane protocol version (
CONTROL_API_VERSION) is now 30. ThinkWatch Lite connects only to a core with the same protocol version, and has to regenerate its protocol types from this release to use it. ThinkWatch Lite 2026.9.24 includes 0.55.2 (protocol 28) and does not connect to 0.56.0: a server used with it stays on 0.55.2 until the app is updated to a release that includes 0.56.0.sudo twcore upgrade --version 0.55.2 --restartswitches a server back to 0.55.2. - The request store’s schema is unchanged (22): the request history is kept.
groups[].session_affinityis removed. A configuration that still sets it (onlysession_affinity: falsewas ever written) is refused; delete the line.- Changed in the protocol:
- Removed:
GroupView.session_affinity,GroupView.hurts_cache,GroupInput.session_affinity,DryRunResult.hurts_cache. RequestRouted.affinityandRoutingView.affinity(AffinityView,Stay): whether the request kept its turn’s route and why it stayed with the upstream that answered before.Overview.failover(FailoverView).AttemptOutcomehasestimated.
- Removed:
- New in the configuration: the
failoversection (see Setting a failing upstream aside). - Message codes: new
config.failover_rangeandgw.upstream.stream_opening_error; removedgw.count_tokens.bedrock_model.
A conversation stays where its cache is. Routing rules used to be evaluated for every request, and only
load-balancegroups kept a session on one upstream. A rule oninput_tokensor on images could therefore switch upstreams in the middle of an agent’s turn, and a failover moved the conversation away from the upstream holding its prompt cache for good.- A conversation is recognised by the client’s session header (
x-claude-code-session-id,session_id,conversation_idand a few others) together with the content fingerprint, or by the fingerprint alone. A turn lasts from a user message until the next one; tool results belong to the turn. - The routing decision made at the start of a turn (rule, group, rewrites) holds for the rest of the turn. It is made again when the input approaches the smallest context window among the candidate models, where the price data gives one.
- The next request goes first to the upstream that last answered the conversation, including one that took over after a failover: always within a turn, and across turns while its last answer read or wrote at least 1024 cached tokens less than five minutes ago. An upstream that is set aside, not among the candidates or in another group is not held to.
load-balancetakes turns between new conversations only.session_affinityand the warnings that a group lowers the cache hit rate are gone, since no strategy drops a warm cache in the middle of a conversation any more.
Errors at the start of a stream fail over. An upstream can answer 200, open a stream and report an error as its first event: Anthropic’s
overloaded_error, the Codex backend’sresponse.failedwhen the usage limit is reached, Bedrock’sthrottlingException, an error chunk from a Chat or Gemini upstream. The client had received nothing yet, and the error still reached it. A streamed answer is now held until its first content arrives, and an error before that point moves the request to the next candidate, as an error status does. What was read while waiting is passed on unchanged. The wait ends afterfailover.stream_start_wait_secs(15) or 1 MiB; the last candidate is not held.Setting a failing upstream aside. Failures are sorted by the reason the upstream gives.
- 401, 403, 402 and 404 now move the request to the next candidate. A 400 or 422 does so only when its body names an insufficient balance, a used-up quota or an unavailable model; other 4xx still go to the client unchanged. The last candidate’s own 4xx reaches the client as the upstream wrote it.
- An insufficient balance sets the upstream aside for
no_balance_pause_secs(1800). A used-up quota sets it aside until the reset time the upstream names in the body, in its quota headers or in a GLM 429, and forquota_pause_secs(3600) when it names none. A rate limit withRetry-After(or Gemini’sretryDelay) sets it aside for that long, at mostrate_limit_max_pause_secs(3600). A missing model moves on without counting against the upstream. - Failures without a stated reason pause the upstream after
failures_to_pause(3) in a row, forpause_secs(60), doubling with each further pause up tomax_pause_secs(600) and starting over after a success. - A request with a single candidate is never affected, and when every candidate is set aside they are all tried.
Counting tokens.
/v1/messages/count_tokensand Gemini’s:countTokensrouted to an upstream of another format, or to a same-format upstream that answers 404 or 405, are answered by the gateway with an estimate of the system prompt, messages, tool calls and results, and tool definitions; nothing is sent to that upstream. The answer carriesx-thinkwatch-local: 1, and the request is recorded as a local answer with no cost. A count does not fail over to another format or another model. Counting through an AWS Bedrock upstream still answers 501not_supported, after which Claude Code counts precisely itself.Reasoning sealed by another account. Reasoning items carry content encrypted or signed for the account that produced them. After a failover from one account to another, the upstream refuses them (Responses’
invalid_encrypted_content, Anthropic’s invalidsignaturein athinkingblock). The request is now sent once more without them, and the refused items are left out up front on the conversation’s later turns to that upstream. - The control-plane protocol version (
-
ThinkWatch Lite v2026.9.24 GitHub ↗
Upgrade notes:
- The bundled core is now 0.55.2, which speaks control-plane protocol 28. A remote server has to run core 0.55.2 too: this version does not connect to 0.54.0, and 2026.9.23 and earlier do not connect to 0.55.x.
- Request history is kept.
- A request that a routing rule sent as another model is now priced by the model it was sent as, so its cost can differ from what an earlier version showed (see Costs).
Amazon Bedrock upstreams: Amazon Bedrock is in the list of services.
- Choose the region and one credential: a Bedrock API key; access keys, with a session token for temporary credentials; or an AWS profile, read from the AWS credential files on the machine the core runs on. Each can be written as
${VARIABLE}. With a remote core, the variables and the profile are read on the server. A profile that signs in or runs a command to get its keys cannot be used, and says so. - Check connection verifies the credential and lists the models the region can route to: foundation models, inference profiles such as
us.anthropic.claude-…andglobal.…, and the account’s own application inference profiles. - Requests in every client format are converted to Bedrock’s Converse API and signed one by one. Streaming, usage, output redaction and tool-call inspection work as with other upstreams, and prompt-cache breakpoints and Claude’s thinking are carried over.
- Bedrock takes its own model IDs. A client that asks for Anthropic’s names, such as
claude-sonnet-4-5, needs a routing rule that changes the model to a Bedrock ID. - For a VPC endpoint or a proxy, enter its address and the region.
- When AWS refuses a credential, its message names the account and the IAM identity; clients get the gateway’s own sentence instead.
- Replaying a request to a Bedrock upstream is refused: a replay sends the recorded request unchanged, and Bedrock takes only the Converse format, which no client sends.
Claude Code and Claude Desktop that used Bedrock directly:
- Claude Code with
CLAUDE_CODE_USE_BEDROCKturned on ignored the gateway address and kept calling Bedrock directly, although it showed as connected. The same held for its other provider switches: Mantle, Google Cloud’s Agent Platform, Microsoft Foundry and Claude Platform on AWS. Connecting Claude Code now turns such a switch off in~/.claude/settings.json, and restoring turns it back on exactly as it was. A switch turned on by an organization’s managed settings cannot be overridden, so connecting is refused with the reason. The configuration check reports a switch that was turned on again later. - The confirmation lists the model variables that name no Bedrock model: those requests reach the gateway under Anthropic’s model names and need a routing rule.
- The confirmation says whether the gateway has an Amazon Bedrock upstream, and New Bedrock upstream… opens the new upstream filled in from the client’s own settings: the region, a custom endpoint, and the credential as a
${VARIABLE}or a profile name. A key kept in plain text in the client’s configuration is not copied; the notes say where it is. Claude Desktop’s Bedrock configuration in use before connecting is recognized the same way. - Credentials in the
envblock ofsettings.json, such as the Bedrock API key and access keys that/setup-bedrockwrites there, are masked in the change shown before connecting. A value written as an empty string is shown as"".
Web search through other formats: Claude Code runs each web search as a request that only an upstream in Anthropic’s format can carry out. Sent to a Bedrock, OpenAI or Gemini upstream, the search tool was dropped and the model’s own answer came back as the search results. Such a request now goes to the next upstream of the route, and otherwise fails with a message naming the tool and the upstream.
Costs:
- A request that a routing rule sent as another model is priced by the model it was sent as. The traffic view still shows the model the client asked for.
- Long-context prices apply from the threshold the price data gives (200K for Claude Sonnet 4 and 4.5, 272K for GPT-5.4 and later), and cache reads and writes count toward it.
- A price the price data does not give, such as a one-hour cache write, is filled in and the cost is marked estimated. A Bedrock model the price data does not list borrows the price of the same model and is marked estimated.
Errors partway through a stream: an upstream that starts answering and then reports an error inside the stream, such as Anthropic’s
overloaded_erroror Responses’response.failed, now leaves the request recorded as failed instead of finished, with the usage reported before the error. -
ThinkWatch Core v0.55.2 GitHub ↗
This release stops a Bedrock API key that is not valid from passing the connection check.
Upgrade notes
- The control-plane protocol version is unchanged (28) and so is the request store’s schema (22). ThinkWatch Lite 2026.9.23 includes 0.54.0 (protocol 27) and does not connect to 0.55.x: a server used with it stays on 0.54.0 until the app is updated to a release that includes 0.55.2.
- No message codes change.
Checking a Bedrock upstream. AWS answers a Bedrock API key that is not valid with
AccessDeniedException(“Authentication failed: …”, “Invalid API Key format: …”), the same exception it uses for a valid credential that may not list models. The check took everyAccessDeniedExceptionon the model listing to mean the credential works, so a mistyped key passed with “The credential works, and it has no permission to list models”, and the background model fetch said the same. Only a refusal whose message says the identity is not authorized, which is how IAM denies an action, now counts as missing list permission; any other one reports the key as rejected. AWS’s message is only inspected and never passed on, since it names the account. -
ThinkWatch Core v0.55.1 GitHub ↗
This release stops a request that requires a server-side tool, such as Claude Code’s web search, from being sent without it to an upstream of another format.
Upgrade notes
- The control-plane protocol version is unchanged (28) and so is the request store’s schema (22). ThinkWatch Lite 2026.9.23 includes 0.54.0 (protocol 27) and does not connect to 0.55.x: a server used with it stays on 0.54.0 until the app is updated to a release that includes 0.55.1.
- New message codes:
gw.convert.tool_unsendableandgw.convert.tools_unsendable.
Requests that require a server-side tool. Claude Code runs each web search as a request that carries only the server-side
web_searchtool and forces the model to use it. Converted for an upstream of another format (AWS Bedrock, OpenAI, Gemini), the server tool was dropped, since only the provider it belongs to runs it: the model answered without searching, and Claude Code passed that answer on as the search results. When the tool a request forces cannot be sent in the upstream’s format, that attempt is no longer made. The next upstream of the route is tried, and if none can take the request the client gets a 400 that names the tool and the upstream. Server-side tools offered alongside other tools, which the model is free to leave unused, are still dropped and listed in the request’s details as before. -
ThinkWatch Core v0.55.0 GitHub ↗
This release adds AWS Bedrock upstreams. Requests to them are converted to Converse and signed per request, and prompt-cache breakpoints and Claude’s thinking now survive that conversion. Each request is priced by the model it was actually sent as, long-context prices follow the thresholds the price data gives, and an error an upstream reports partway through a stream now counts as a failure.
Upgrade notes
- The control-plane protocol version (
CONTROL_API_VERSION) is now 28. ThinkWatch Lite connects only to a core with the same protocol version, and has to regenerate its protocol types from this release to use it. ThinkWatch Lite 2026.9.23 includes 0.54.0 (protocol 27) and does not connect to 0.55.0: a server used with it stays on 0.54.0 until the app is updated to a release that includes 0.55.0.sudo twcore upgrade --version 0.54.0 --restartswitches a server back to 0.54.0. - The request store’s schema is unchanged (22): the request history is kept.
- Changed in the protocol:
Protocolhasbedrock.ProviderInput.awsandProviderView.aws(AwsKeys): the access keys, or the name of an AWS profile, and the region to sign for.ProviderView.regionandProviderPreview.region: the region of a Bedrock upstream.AttemptView.model: the model an attempt sent when a routing rule rewrote it.
- New in the configuration: the protocol
bedrockandproviders[].aws. The 200K long-context threshold of a price sheet’sinput_above_200kandoutput_above_200know counts cache reads and writes (see Priced by the model that was sent). - New message codes:
config.credential.:bedrock_oauth,key_and_aws,bedrock_no_credential,aws_not_bedrock,aws_empty,aws_profile_and_keys,aws_empty_profile,aws_signed_header,aws_no_region,aws_bad_region,aws_region_mismatchconfig.aws_profile.:no_home,unreadable,not_found,unsupported,no_keysgw.upstream.:aws_token_expired,aws_profile_expired,bedrock_refused,bedrock_refused_unnamed,sign_failed,eventstream_broken,stream_exception,stream_errorgw.probe.aws_token_expired,gw.probe.bedrock_list_denied,gw.count_tokens.bedrock_model,gw.count_tokens.bedrock_upstream,control.replay_bedrock,l3.sign_failed
AWS Bedrock upstreams. An upstream at
https://bedrock-runtime.<region>.amazonaws.comis recognized asbedrock. It authenticates in one of three ways:- A Bedrock API key in
key, sent asAuthorization: Bearer. - AWS access keys in
aws:access_key_id,secret_access_keyand, for temporary credentials,session_token, each of which can be${VAR}. Every request is signed with them (SigV4) once its body is final. aws.profile: the access keys of a profile in the AWS credential files (~/.aws/credentials,~/.aws/config, or the filesAWS_SHARED_CREDENTIALS_FILEandAWS_CONFIG_FILEname) on the machine core runs on. The files are read again when they change, so a tool that refreshes temporary keys there needs no restart. Nothing is run to obtain a credential: a profile that signs in through IAM Identity Center, runs acredential_processor assumes a role is refused, naming the setting.
Requests in every client format are converted to Converse and ConverseStream. The binary eventstream is turned into server-sent events as it arrives, so usage, the first token, output redaction and tool-call inspection work as they do for any other upstream. The
anthropic-betavalues Bedrock accepts for Claude are carried into the request body. The model list comes from the region’s control plane: the foundation models that can be invoked on demand, the inference profiles AWS defines (us.anthropic.claude-…,global.…), and the account’s application inference profiles, listed by the ARN they are invoked by. Checking the connection and the inference speed test go through the same signing. For a VPC endpoint or a proxy, write its address inbase_urland the region inaws.region.- AWS’s text for a refused credential (HTTP 401 or 403) names the account and the IAM identity. The client receives the gateway’s own sentence instead, and AWS’s text is not recorded as the request’s error.
- An expired session token is named as such; for a profile, the message says to refresh the profile.
- An exception AWS sends inside a stream ends the request as a failure.
count_tokensfor a model that only Bedrock upstreams serve is answered with 501not_supported, the answer Claude Code’s guidance for gateways names; Claude Code then counts with a one-token request. Bedrock’s own CountTokens covers only some older Claude models onbedrock-runtime.- A replay to a Bedrock upstream is refused: a replay sends the recorded request as it was, and Bedrock takes only the Converse format, which no client sends.
The code that speaks to Bedrock (signing, the eventstream, the endpoints and the model catalog) is the new layer-one crate
tw-bedrock, which ThinkWatch Enterprise shares.Converse keeps cache breakpoints and thinking. The conversion to Converse used to drop
cache_control, so a client that cached its prompt never hit the cache through Bedrock. Cache breakpoints now becomecachePointblocks (at most four, as Bedrock allows), with the one-hour TTL where the client asked for it, and cache reads and writes, one-hour writes included, are read back. Claude’s thinking and effort are sent inadditionalModelRequestFields, and its reasoning comes back as reasoning content.Priced by the model that was sent. A routing rule can send a request as another model, and the upstream charges for the model it received. Each attempt now records the model it sent when a rule rewrote it (
AttemptView.model), and a request is priced by the model of the attempt that served it. The row still names the model the client asked for. The list of unpriced models names the model a price has to be set for.- A Bedrock model the price data does not list borrows a price, always marked estimated: an inference profile from the model it is named after, and an Anthropic model on Bedrock from Anthropic’s own name for it.
- Long-context tiers are read with the threshold the price data gives (200K for Claude Sonnet 4 and 4.5, 272K for GPT-5.4 and later), including their cache prices. A request’s input counts its cache reads and writes when the tier is chosen. The 272K tier was not recognized before, and a request with a warm cache was priced as a short one.
- A price the data does not give is filled in, and the cost is marked estimated whenever it is used: a one-hour cache write costs twice the input price (it used to fall back to the five-minute price), and the missing cache prices of a long-context tier grow with its input price.
Errors partway through a stream. An upstream can answer 200, stream part of an answer and then report an error in the stream: Anthropic’s
overloaded_error, Responses’response.failed, anerrorin a Chat or Gemini frame. Such a request was recorded as finished; it is now a failure, with the usage the upstream reported before the error.Listed models.
GET /v1/modelsandGET /v1/models/:modelno longer stamp each model with the time of the request.createdandcreated_atare now the Unix epoch, since the gateway does not know when a model was released. - The control-plane protocol version (
-
ThinkWatch Lite v2026.9.23 GitHub ↗
Upgrade notes: the bundled core stays at 0.54.0, so a remote server running 0.54.0 keeps working with this version and nothing on the server has to change. Request history is kept.
Updates keep the window as it was: after an update, the app comes back the way it was before. If it was running in the menu bar only, it stays in the menu bar; if its window was open, the window opens again. Before, the main window opened after every update. This applies to updates the app installs itself, on every platform, and to
brew upgradeon macOS. A Homebrew upgrade now also shows the “Updated to …” notification when Notices is set to System notification.The update from 2026.9.22 to this version is not covered yet: 2026.9.22 does not record the window when it quits, so this one update still opens the window and shows no “Updated to 2026.9.23” notification.
Sidebar: with the sidebar expanded, hovering a page no longer pops up a tooltip that repeats its name. The page’s shortcut (⌘1 to ⌘9, or Ctrl+1 to Ctrl+9 on Windows and Linux) appears at the end of its row while the pointer is over the row or the row has keyboard focus. The collapsed sidebar still shows the name and the shortcut in a tooltip.
-
ThinkWatch Lite v2026.9.22 GitHub ↗
Upgrade notes:
- The bundled core is now 0.54.0, which speaks control-plane protocol 27. A remote server has to run core 0.54.0 too: this version does not connect to an older core, and 2026.9.21 and earlier do not connect to 0.54.0.
- Request history is emptied. The request store changed format, so the requests recorded by an earlier version are removed when the new core first starts. Settings, keys and upstreams are kept.
Time to first token: latency now counts to the first token of the answer instead of to the response headers. The headers of a streamed response usually come back as soon as the upstream receives the request; the first token comes only after the upstream has queued the request and read the prompt. Traffic, the overview and the upstreams page show this time, and the response-header time moves into the request details.
Generation speed: a new figure, in tokens per second, is measured from the first token to the end of the answer.
- Traffic: hovering a request’s latency shows its first token, total time, generation time and speed.
- Request details: the speed is one of the five figures at the top, and the bar is split into the wait for the first token and the generation.
- Overview: a new Generation speed section gives the median speed by model and by upstream.
- Upstreams: the timing column shows the median first token, with the median speed below it.
Some requests have no speed because it cannot be measured honestly:
- responses that are not streamed;
- cancelled and failed requests;
- generations shorter than half a second;
- answers whose model reasoned without showing it, when the upstream does not report how many tokens the reasoning took.
Menu bar:
- The generation speed in the menu was wrong: a single response that was not streamed could push it to tens of thousands of tokens per second. It now uses the same measurement.
- After the computer slept through midnight, Today kept showing the previous day’s tokens and cost until a request came in or the menu was opened. It now catches up as soon as the computer wakes, and when the clock or the time zone is changed. The traffic page, which shows only the time for today’s requests, and the keys page’s usage over the last 24 hours catch up the same way.
Routing: a model name that is not among the suggestions is kept, in a rule’s model condition, in Change model to and in the dry run. Before, closing the suggestions cleared a name the list did not have, such as a relay’s own model name or a model released after the upstream’s list was fetched, and turned a shorter name back into the suggestion picked earlier.
-
ThinkWatch Core v0.54.0 GitHub ↗
This release records when each response’s first token arrives and how fast the response is generated, and corrects the generation rate that
GET /livereports.Upgrade notes
- The control-plane protocol version (
CONTROL_API_VERSION) is now 27. ThinkWatch Lite connects only to a core with the same protocol version, and has to regenerate its protocol types from this release to use it. ThinkWatch Lite 2026.9.21 includes 0.53.0 (protocol 26) and does not connect to 0.54.0: a server used with it stays on 0.53.0 until the app is updated to a release that includes 0.54.0.sudo twcore upgrade --version 0.53.0 --restartswitches a server back to 0.53.0. - The format of the request history changed (request store schema 22, from 21). On its first start, 0.54.0 replaces the request history and the stored request and response bodies of an earlier version with an empty store; the configuration is kept. Switching back to an earlier version empties it again.
- Changed in the protocol: the new event
RequestFirstToken { id, ttft_ms },RequestFinished.tokens_per_sec,HistoryRow.ttft_msandHistoryRow.tokens_per_sec, and the new endpointsGET /token-rateandGET /token-rate/provider, which returnTokenRateView { model, p50, samples }.GET /latencyandGET /latency/providernow report the time to the first token. No message code changed.
First token. The headers of a streamed response usually come back as soon as the upstream receives the request. The upstream then queues the request and reads the prompt before it produces anything, and the time a user waits is the time to the first token, not to the headers. The gateway now reads each successful streamed response in the upstream’s format and reports the moment the first text, reasoning or tool call arrives; opening frames such as
message_startandresponse.createddo not count. A response that is not streamed arrives whole and has no first token.GET /latencyandGET /latency/providernow give percentiles of this time, so responses that are not streamed are left out of them.url-testgroups still pick the upstream with the fastest response headers.Generation speed. Each finished streamed request has a generation speed: the output produced after the first token, divided by the time from the first token to the end.
- When the model reasoned before the first token without streaming its reasoning, those reasoning tokens were generated outside that time. They are subtracted when the upstream reports how many there were, as OpenAI and Gemini do. When it does not, as with Anthropic, the request has no speed rather than one several times too high.
- Reasoning that is streamed as it happens is counted.
- Responses that are not streamed, cancelled and failed requests, and generations shorter than half a second have no speed.
GET /token-rateandGET /token-rate/providergive the median speed by model and by upstream, with the number of samples.The live rate.
GET /liveused to divide each request’s output by the time after its response headers. A response that is not streamed gets its headers only when generation is done, so that time was a few milliseconds, and one such request could push the rate to tens of thousands of tokens per second. The rate now combines the speeds of the requests that have one, weighted by their generation time. - The control-plane protocol version (
-
ThinkWatch Lite v2026.9.21 GitHub ↗
Upgrade notes: the bundled core is now 0.53.0, which speaks control-plane protocol 26. A remote server has to run core 0.53.0 too: this version does not connect to an older core, and 2026.9.20 and earlier do not connect to 0.53.0. Request history is kept. Upstreams now see ThinkWatch’s User-Agent instead of the client’s, so an upstream that admits only certain clients, such as Kimi For Coding, Bailian Coding Plan or a relay restricted to official clients, needs the new Forward client identity switch in its connection settings.
What upstreams receive:
- Each upstream receives the request and only the headers its API needs. Before, almost every header a client sent went along: Claude Code’s User-Agent,
x-app, SDK details and session ID, Gemini CLI’s installation ID, and a browser client’s cookies and origin. Upstreams now see ThinkWatch’s User-Agent and the credential and headers configured for the upstream. From the client’s request they get only the headers of the upstream’s own protocol:anthropic-*for Anthropic,Idempotency-KeyandX-Client-Request-Idfor OpenAI, and none for Gemini. - Identity fields that clients fill in on their own are removed from the request body, such as Claude Code’s
metadata.user_id, which holds a device ID, an account ID and a session ID. - Forward client identity, a new switch in the connection section of the upstream dialog, sends that upstream the client’s own User-Agent and identity, unchanged. It is off by default, including for upstreams created by the Z.ai sign-in, and ChatGPT accounts do not offer it.
ChatGPT accounts:
- Newer models that the account can use, such as GPT-6-Luna, appear in the upstream’s model list and can be requested. Before, the gateway answered “No upstream serves model gpt-6-luna” even though the account had the model.
- Requests from the Codex in the ChatGPT app no longer pass that app’s attestation and installation ID on to OpenAI. On one account, OpenAI revoked the sign-in within seconds of the first such request.
- Request bodies are compressed with zstd, as Codex does, so long conversations upload faster.
- Each upstream receives the request and only the headers its API needs. Before, almost every header a client sent went along: Claude Code’s User-Agent,
-
ThinkWatch Core v0.53.0 GitHub ↗
This release changes what ThinkWatch sends to upstreams. Each upstream now receives the request and only the headers it needs; the client’s own identity stays with ThinkWatch unless an upstream asks for it. It also fixes two problems with ChatGPT account upstreams.
Upgrade notes
- The control-plane protocol version is now 26 (it was 25). The request store’s schema is unchanged (21). ThinkWatch Lite 2026.9.20 includes 0.52.1 (protocol 25) and does not connect to 0.53.0, so a server used with it stays on 0.52.x until the app is updated to a release that includes 0.53.0.
- Upstreams now see ThinkWatch’s
User-Agentinstead of the client’s. An upstream that admits only certain clients, such as Kimi For Coding, Bailian Coding Plan or a relay restricted to official clients, needs the new provider fieldforward_client_identity: true. - New message code:
config.credential.chatgpt_client_identity.
Each upstream gets only what it needs. Client request headers used to reach the upstream almost unchanged. Only credentials, a few transport headers and, when a request was translated, the client format’s own protocol headers were removed. Claude Code’s
User-Agent,x-app,x-stainless-*and session ID, Gemini CLI’s installation ID and a browser client’sCookieandOriginall went through. Outbound headers now come from three sources only:- Set by ThinkWatch: its own
User-Agent, andcontent-typewhen the body was translated. - The upstream’s configuration: the credential and the headers written there.
OpenAI-OrganizationandOpenAI-Projectbelong here, because they belong to the upstream’s key. - The client’s request, only for the headers the upstream’s protocol uses:
- Anthropic:
anthropic-*, exceptanthropic-dangerous-direct-browser-access. - OpenAI:
Idempotency-KeyandX-Client-Request-Id. - Gemini: none.
- Every protocol:
content-typeandacceptwhen the body passes through unchanged.
- Anthropic:
An Anthropic request without
anthropic-versionnow gets the default one. When the body passes through unchanged, identity fields that clients fill in themselves are removed from it: Claude Code’smetadata.user_id, and the Codex installation ID and turn metadata inclient_metadata.forward_client_identity. With this field on, an upstream also receives the client’s ownUser-Agent,x-app,originator,version,session_idandx-codex-*headers, and the body’s identity fields are kept. These are the client’s own values; nothing is made up. The field is off by default. A ChatGPT account upstream refuses it: that upstream always names ThinkWatch as the sender.ChatGPT account upstreams.
- New models were missing from the list. The Codex backend leaves a model out of the list when its minimum client version is above the
client_versionthat the caller reports. ThinkWatch reported0.0.0, so a model such as gpt-6-luna, which needs Codex 0.155.0, was not listed, and requests for it were refused with “No upstream serves model …” even though the account can use it. ThinkWatch now reports99999.0.0, so the list shows every model the account can use. Requests still name ThinkWatch as their origin. - The ChatGPT app’s identity is no longer forwarded. The Codex inside the ChatGPT desktop app adds an attestation generated by the app (
x-oai-attestation), its installation ID and turn metadata. ThinkWatch forwarded them together with the token it had obtained itself, and on one account OpenAI revoked that token within seconds of the first such request.- From the client, the Codex backend now receives only
x-codex-turn-state,x-openai-internal-codex-responses-liteandx-codex-beta-features. - The attestation, installation ID and turn metadata cannot be set in an upstream’s configured headers.
- From the client, the Codex backend now receives only
- Request bodies are compressed. Requests to the Codex backend are compressed with zstd, which is what Codex does when it talks to that backend itself.
-
ThinkWatch Lite v2026.9.20 GitHub ↗
Upgrade notes: the bundled core stays at 0.52.1, so a remote server running 0.52.1 keeps working with this version and nothing on the server has to change.
Connecting and restoring clients:
- A change is written only as it was shown. If a client rewrites its configuration while the confirmation dialog is open, nothing is written and the dialog shows the change worked out from the file as it is now, to confirm again. This applies to connecting and restoring a client, copying or removing an MCP server, and switching WSL to mirrored networking.
- Restoring no longer deletes sections that connecting had created once something else was added to them: Claude Code’s
envinsettings.json, theprovidersection of a new opencode configuration, and DeepSeek Harness’srefs. Restoring DeepSeek Harness no longer fails with “refs is not a scalar”, and its credentials file keeps itsversionwhen DeepSeek Harness has written its own entries into it. - A configuration file that contains an empty object written as
{ }no longer gets a stray comma when a setting is added to it. Claude Code and Claude Desktop rejected such a file. - In TOML files, a comment at the end of a changed line is kept, and an MCP server copied from a JSON client no longer gets empty strings where it had
null. - When the record of an earlier connection next to a client’s configuration is damaged or belongs to another client, connecting is refused, as restoring already was. Before, it recorded the gateway’s address and key as the original values, and a later restore put them back.
- Keys no longer appear in the change shown before writing: restoring showed the user’s own key, and copying an MCP server showed the environment variables and headers of every server in the file. Messages about a configuration file that cannot be parsed no longer quote the line that failed, and a failed write no longer leaves a temporary copy with the gateway key next to the configuration.
- Backups are readable by their owner only, and a configuration folder that is a symbolic link (for example into a dotfiles repository) is pointed out before writing.
- Pointing Claude Desktop at a new address or key, after a key change or when switching to a remote core, updates its model list too.
- On Windows, a running client is recognized by its program name, so a program such as TombRaider.exe is no longer taken for Aider.
- Uninstalling keeps the data folder when a client could not be restored, since it holds that client’s only full backup. It stops the core before deleting the data, and its confirmation also lists the clients inside WSL that will be restored.
- The “Change path…” dialog no longer freezes the window while a network path that is offline takes its time to answer.
- On a new installation, the data folder is created readable by its owner only, as intended. It holds
config.yamlwith the upstream keys, and on Windows its access rules are their only protection.
Figures:
- When a statistic cannot be read, the Overview says it is unavailable and the Upstreams page shows ”—” with a Retry button, instead of zeros or an empty chart.
- Today’s cost in the menu bar, and the costs on the Keys and Upstreams pages, keep measured, estimated and unpriced apart, as the Overview does. A cost is no longer shown as $0 when the requests could not be priced.
- Very small amounts are shown as ”<$0.0001” instead of “$0.0000”. Numbers change unit after rounding, so they read “10k” and “1.0M” rather than “10.0k” and “1000k”.
- In live mode, requests whose price has not arrived yet make the total a lower bound, marked ”≥”, instead of counting as $0.
- In a session, a turn is marked Unpriced only when a price would fix it. “Show unpriced only” on the Traffic page no longer lists failed requests and requests without usage.
- Requests in progress when a remote connection dropped are corrected from the server’s history instead of staying marked as failed.
Time ranges: “Today” in the menu bar starts at midnight on this computer, also when the core runs on a server in another time zone and on the days clocks change. A custom range such as “since 9/8” keeps its start day while the page stays open. Long custom ranges are drawn with daily or weekly points and always reach the present; ranges over about four months used to stop short of it.
Core and remote connections:
- Switching to a remote core also stops a local core that was still starting or restarting. Before, it kept being started and held the gateway port.
- A core in safe mode is shown as being in safe mode, not as running, and Restart works from there.
- A remote connection that died without a sign, for example after the computer slept or the network changed, is noticed within 45 seconds, and a server that accepts the connection but never answers makes the connection test time out.
- Quitting the app stops the local core instead of leaving it to notice a few seconds later.
Notifications:
- A tool call blocked by the security rules, and something suspicious appearing in a client’s configuration, are reported again each time they happen. Before, only the first one was ever reported, even after a restart. To keep a burst from flooding the screen, each shows at most once every five minutes (per upstream for blocked calls), and one notification then sums up the rest.
- A notification that waits a minute before showing is dropped if the gateway went down meanwhile. The “N more” count no longer counts the same notice twice. The title of an upstream that keeps dropping out no longer grows longer with each request.
- “Updated to …” follows the notification setting.
- The notice list and the settings are saved in a way that survives the app being killed mid-save, and a setting this version does not recognize, such as one written by a newer version, resets only that setting.
Menu bar and tray: the color reflects the most severe quota window, not only the fullest one. On Windows and Linux, menu actions such as Check for Updates report their result in a notification, and messages from the core appear in the interface language. A click on a menu that has just been replaced no longer runs a different item. The menu bar no longer stops updating when the core stops answering.
MCP and the security scan:
- The scan also covers MCP servers configured per project in
~/.claude.json, Claude Code settings that run commands without the model (statusLine,apiKeyHelperand similar), and Codex’snotify. Helpers that print credentials are scanned but not listed, so their commands do not appear on screen. - Skills and folders created after the app started are watched as well, and so is
opencode.jsonc. - A remote MCP server is no longer taken for a local one because its address contains “localhost”.
SKILL.mdfiles with Windows line endings or a byte order mark are read correctly. - A file full of hidden characters no longer stalls the scan; findings are capped per file, with a note saying where the scan stopped. A special file, such as a link to
/dev/zero, no longer blocks it. - Findings in a hook’s command, including hidden characters, are shown on that hook’s row.
Editing:
- A
config.yamlwith Windows line endings no longer shows unsaved changes on opening, and is saved with its line endings as they were. “Show in config file” selects the right entry when a key and a route share a name. - Edit dialogs save against the configuration they were opened with. A change made elsewhere meanwhile is reported as a conflict instead of being overwritten. Switching several items on or off in quick succession no longer fails with “version mismatch”.
- Price fields and the dry run accept a decimal comma.
- In the upstream dialog, the connection check result is cleared when the connection settings change.
- In a key’s model scope, model IDs typed by hand are shown and can be removed even when the model list is empty, and they are compared without regard to case, as the gateway does.
- Saving a route applies its auxiliary request categories as they are after the edit: unticked or deleted conditions no longer keep them assigned to routing.
max_tokensof 0 is refused instead of being dropped on save, and a copy of the fallback rule can be moved ahead of it. - A replayed request goes to the upstream it was priced for.
- Several dialogs no longer act on stale state: the speed test waits for the new quote before it can start, deleting a price sheet from the upstream dialog also clears it from the form, checking for updates from About no longer overwrites settings changed meanwhile, dismissing one key-rotation banner no longer dismisses the others, and Z.ai sign-in accepts an existing upstream whose address ends in a slash.
- A small payload cap in Retention no longer blocks saving that section.
- Ctrl/⌘+B no longer toggles the sidebar while typing in a text field.
- On Windows, choosing a connection at startup and opening the update window no longer risk freezing the app.
Speed: the update and connection windows load a fraction of the code they used to, and the configuration editor loads when it is opened. The change shown before writing a large
~/.claude.jsonappears at once. The Traffic page does less work per event, and the macOS menu bar updates its menu in place while it is open. -
ThinkWatch Lite v2026.9.19 GitHub ↗
Upgrade notes: the bundled core is upgraded to 0.52.1. The request store keeps its format, so the request history is kept. A remote server has to run the same core version, 0.52.1 (ThinkWatch Lite 2026.9.18 cannot connect to a server running 0.52.x), so the server and the app are updated together. When the versions differ, the app gives the command to run on the server,
sudo twcore upgrade --version 0.52.1 --restart. A server upgraded from 0.51.0 keeps its request history as well.Upstreams and proxies: the edit dialogs show the settings as they are written in the configuration.
- The base URL keeps its path. It used to be cut down to its origin, under a note that the credentials in the URL are hidden.
- The API key is filled in and hidden, with an eye button to show it, and the authorization header under Headers follows the same button. A
${NAME}reference stays visible. The other header values are filled in as written. - OAuth upstreams have their token endpoint, refresh token, client ID and client secret filled in, the secrets hidden. The access token field appears only once the credential is new or changed.
- A proxy’s username and password are filled in, the password hidden.
- A save sends the whole definition, so the “leave empty to keep”, “Remove” and “Replace” states are gone.
- A base URL that ends with its version segment, as provider docs give it (
https://api.openai.com/v1), no longer has every request and connection check sent to…/v1/v1/…, and the hint under the base URL now only asks to leave out an endpoint path such as/chat/completions. - A proxy password can be a
${NAME}reference. A password containing${used to be refused. - Z.ai sign-in: entering the name of the existing Z.ai upstream was refused as a name already in use. The dialog now says that the sign-in replaces that upstream’s key, and the sign-in can go ahead.
Routing:
- A rule that rewrites the model was refused with “No upstream serves model …” whenever the target upstream lists its models. That is the rule the app suggests for Claude Desktop with a non-Claude upstream, and for Antigravity CLI and DeepSeek Harness with another provider’s models. Admission, skipping the candidates that cannot serve a request, the cheapest group’s price comparison and the routing dry run now look at the model each upstream is actually asked for. When a rewritten model is not served, the error names both models.
- Requests from Codex and other clients that use the OpenAI Responses or Gemini format are recognized as conversations. A load-balance group with sticky sessions, the default, sent all of them to its first upstream. Each conversation now stays on one upstream while the conversations are spread across the members, and the Traffic page groups them into sessions.
Clients: a client’s row menu on the Clients page (right-click or the … button) has “Change path…”, for a client whose files are somewhere else, such as Claude Code moved with
CLAUDE_CONFIG_DIRor Codex moved withCODEX_HOME. It covers the clients under “Set up by hand” too, such as Cursor and Antigravity CLI. The dialog shows where connecting, MCP management and the security scan read and write. They move together: before anything changes, the dialog lists every location that moves, old and new. “Restore defaults” puts them back. The same dialog opens from the MCP page, by right-clicking a column header or a row under “Skills and hooks” or “Findings”. A connected client has to be restored first. The locations apply to this computer only, and clients inside WSL keep their default locations. Claude Desktop and DeepSeek Harness cannot be moved.Remote connections: editing a remote connection fills in the saved key, hidden behind an eye button, so it can be checked or copied.
Menu bar and tray: the menu now reads, from the top: the gateway’s state, Notices, Today, Quota and In Progress, then Open ThinkWatch Lite, the groups whose upstream can be chosen, Copy Gateway Address and Copy Default Key, then Settings…, Connection, Check for Updates… and Quit ThinkWatch Lite…. On Linux, Open ThinkWatch Lite stays at the very top of the tray menu. “Undo Last Configuration Change” is gone; a configuration change can still be restored from the version history in the app. A quota window named by its length, such as the 30-day window a ChatGPT account can report (shown as
30d), is spelled out in the menu, on the Upstreams page and in notifications: “30 days”, or “30-day window” in a sentence. -
ThinkWatch Core v0.52.1 GitHub ↗
This release fixes two routing problems: a routing rule that rewrites the model was judged by the model the client asked for, and requests in the OpenAI Responses and Gemini formats were never recognized as part of a conversation.
Upgrade notes
- The control-plane protocol version is unchanged (25) and so is the request store’s schema (21). ThinkWatch Lite 2026.9.18 includes 0.51.0 (protocol 24) and does not connect to 0.52.x: a server used with it stays on 0.51.0 until the app is updated to a release that includes 0.52.1.
- New message codes:
gw.model.no_upstream_rewrittenandgw.model.not_allowed_rewritten.
Rewriting the model. A rule that sets the model, for example to send the Claude-style names that Claude Desktop requires to a GLM or DeepSeek upstream, was refused with “No upstream serves model …” whenever the target upstream lists its models. Admission, skipping the candidates that cannot serve a request, and the cheapest group’s price comparison all looked at the model the client wrote, before any rule rewrote it. They now look at the model each upstream will actually be asked for, including a rewrite that applies to one upstream only (
provider_would_be), so admission runs after the rules have decided. When a rewritten model is not served, the error names both models. The routing dry run judges the candidates the same way.Conversations in the Responses and Gemini formats. Requests from Codex and other clients that use the OpenAI Responses or Gemini format were never assigned to a session, because the fingerprint read only the Anthropic and Chat Completions fields. A load-balancing group with session affinity, which is the default, sent all of them to its first upstream, and the traffic view did not group them into sessions. The fingerprint now also reads
instructionsandinput, andsystemInstructionandcontents, and uses the request’sprompt_cache_keywhen it has one. Codex sends one per conversation, and Codex conversations started in the same repository open with identical text. -
ThinkWatch Core v0.52.0 GitHub ↗
This release hands ThinkWatch Lite the upstream and proxy settings as they are written in the configuration, so its edit dialogs can show them, and stops repeating the version segment when an upstream’s base URL ends with it.
Upgrade notes
- The control-plane protocol version (
CONTROL_API_VERSION) is now 25. ThinkWatch Lite connects only to a core with the same protocol version, and has to regenerate its protocol types from this release to use it. ThinkWatch Lite 2026.9.18 includes 0.51.0 (protocol 24) and does not connect to 0.52.0: a server used with it stays on 0.51.0 until the app is updated to a release that includes 0.52.0.sudo twcore upgrade --version 0.51.0 --restartswitches a server back to 0.51.0. - The request store’s schema is unchanged (21). Upgrading from 0.50.x or 0.51.0 keeps the request history.
- Changed in the protocol:
ProviderView.base_url,ProviderView.key(now a string),HeaderView.value, the newOAuthView.refreshandOAuthView.client_secret, andProxyView.auth(which replaceshas_auth) carry the settings as written;base_url_masked,SecretViewandHeaderView.maskedare gone.ProviderInput.base_urlis required,ProviderInput.keyis an optional string and everyHeaderInputhas avalue;ProxyInput.authis optional.SecretChangeandProxyAuthInputare gone, and so is the message codecontrol.pass_has_expansion.
Settings as written. The upstream and proxy views used to mask what they showed. The base URL was cut down to its origin, so
https://bedrock-mantle.us-east-1.api.aws/v1was shown without/v1, with a note that credentials had been hidden when there were none. API keys and header values were masked, and OAuth refresh tokens, client secrets and proxy credentials were left out. The app could not fill its edit dialogs in, and every save carried “keep what is saved” states instead of values. The views now carry the settings as they are written inconfig.yaml, and a save carries the whole definition. The access token of an OAuth credential stays out of the view, and an OAuth credential can still be kept as it is on save: while a dialog is open the gateway may renew the token and write back a new refresh token.GET /configalready returned the file unmasked. Diagnostics bundles and logs still mask.Proxy passwords from the environment. A proxy password may be a
${NAME}reference, as the configuration reference shows (pass: ${PROXY_PASSWORD}). A password containing${used to be refused.Base URLs that end with a version segment. A base URL written up to its version segment, as provider docs give it (
https://api.openai.com/v1) and as the configuration reference says, got every request and every connection check sent to.../v1/v1/...: the client’s path (/v1/chat/completions) was appended as it was. When the base URL’s path ends with the version segment the request path starts with (v1,v2,v1beta, …), that segment is now used once. Other segments are joined as before. - The control-plane protocol version (
-
ThinkWatch Lite v2026.9.18 GitHub ↗
Upgrade notes: the bundled core is upgraded to 0.51.0. The request store has a new format: on its first start, the new core empties the request history, the stored request and response bodies and the security log, and statistics such as usage and cost start again from zero. The configuration is kept. A remote server has to run the same core version, 0.51.0 (ThinkWatch Lite 2026.9.17 cannot connect to a server running 0.51.0), so the server and the app are updated together. When the versions differ, the app gives the command to run on the server,
sudo twcore upgrade --version 0.51.0 --restart. When the core on a server is upgraded from 0.49.x, its first start empties the server’s request history as well.Clients:
- Claude Desktop can be pointed at the gateway in one step, through its third-party inference mode. Afterwards Claude Desktop has to be quit completely and reopened; if the sign-in page appears, choose to continue with the gateway. Conversations in this mode are kept apart from the others. A Claude Desktop managed by an organization cannot be pointed at the gateway.
- DeepSeek Harness can be pointed at the gateway in one step, and its web and desktop apps then send their requests through the gateway. Its web search still goes to DeepSeek directly. The MCP page lists its MCP servers, and the scan covers its patches and skills.
- Antigravity CLI takes the place of Gemini CLI on the Clients page, with steps to set it up by hand and a key of its own. Gemini CLI is no longer listed; a Gemini CLI that is already set up keeps working. The MCP page lists Antigravity CLI’s MCP servers, and the scan covers its hooks, MCP configuration, skills and subagents.
- opencode: once opencode was pointed at the gateway, no ThinkWatch model could be selected. The provider it gets now carries the SDK package and the models the client’s key can use, for opencode 1.x and 2.x. When upstreams or routes change, the Clients page says that opencode’s model list needs updating. The MCP page reads the MCP servers of opencode 2.x (
mcp.servers), and the scan coversopencode.jsonc. - Codex: after Codex is restored, the sessions started while it pointed at the gateway can still be opened, and go to OpenAI directly. The notes shown before pointing Codex at the gateway say that sessions started before and after are listed separately.
- WSL on Windows: Claude Code and Codex inside WSL can be pointed at the gateway, under WSL 1 or under WSL 2 with mirrored networking. Both reach the gateway at 127.0.0.1, and the gateway keeps listening on this computer only. When WSL 2 uses NAT networking, the Clients page can switch it to mirrored networking and restart WSL; mirrored networking needs Windows 11 22H2 or later. When the app switches to a remote core, the clients inside WSL can be pointed at the server as well.
Upstreams: GLM Coding Plan upstreams (Z.ai and BigModel) show their 5-hour and weekly quota and the reset times on the Upstreams page, in the menu bar and in the tray, and credit-based plans show the credits left. A used-up quota sends a notification. The monthly count of calls to Z.ai’s own MCP tools in older plans is not shown.
Traffic: the request details show when a request from DeepSeek Harness carried a session log, and its size. Before such a request reaches an upstream other than DeepSeek, the gateway removes the session log and the other DeepSeek Harness extensions.
Gateway: it answers the startup check of Claude Desktop’s third-party mode and keeps a stream that stays silent for a long time alive.
/v1/filesis no longer forwarded to an upstream. A gateway set to listen on a network interface that is not there yet starts on this computer only and adds the interface once it appears. -
ThinkWatch Core v0.51.0 GitHub ↗
This release stops counting the MCP tool calls of the older GLM Coding Plans as a quota, marks only the weekly window when a model request gets business code 1310, and tells the app at once when a GLM key turns out to have no plan.
Upgrade notes
- The control-plane protocol version (
CONTROL_API_VERSION) is now 24. ThinkWatch Lite connects only to a core with the same protocol version, and has to regenerate its protocol types from this release to use it. ThinkWatch Lite 2026.9.17 includes 0.49.0 (protocol 22) and no released version of the app includes 0.50.0, so a server used with the app stays on 0.49.x until the app is updated to a release that includes 0.51.0.sudo twcore upgrade --version 0.49.1 --restartswitches a server back to 0.49.1. - The request store’s schema is unchanged (21). Upgrading from 0.50.0 keeps the request history. Upgrading from 0.49.x or earlier replaces the request history and the stored bodies with an empty store on the first start, as 0.50.0 does; the configuration is kept. Switching back to 0.49.x empties it again.
- Changed in the protocol: the quota window
monthlyis gone, and aQuotaSeenwith emptywindowsmeans the upstream’s quota was taken down. No endpoint, type or message code changed.
MCP calls are not a quota. The quota endpoint of the older GLM Coding Plans also reports a count (
TIME_LIMIT) of calls to Z.ai’s own MCP tools: search-prime, web-reader and zread. Those calls do not go through the gateway, so the count says nothing about the requests it forwards. Core no longer reads it: there is nomonthlywindow, and a used-up count no longer marks the upstream used up or emitsQuotaExhausted.creditsare reported only for the windows of credit-based plans (CREDIT_LIMIT).1310 is the weekly quota. A 429 with business code 1310 on a model request marks the
weeklywindow used up, whatever its message says. 1308 still marks the 5-hour window. For 1316 to 1321 the window still comes from the message when it names the 5-hour or weekly window; otherwise it is the fuller of the two known windows, or the weekly one when neither is known.A key without a plan. When a GLM key is found to have no plan, core clears the quota it had stored for that upstream and emits a
QuotaSeenwith emptywindows, so the app removes the quota at once instead of on its next read of/quota. - The control-plane protocol version (
-
ThinkWatch Core v0.50.0 GitHub ↗
This release lets Claude Desktop’s third-party mode use the gateway, cleans the extensions DeepSeek Harness adds to its requests before they reach an upstream other than DeepSeek, reads the GLM Coding Plan quota of Z.ai and BigModel upstreams, and recognises requests from Antigravity CLI. A gateway bound to a network interface now starts even when that interface is not there yet.
Upgrade notes
- The format of the request history changed (request store schema 21, from 20). On its first start, 0.50.0 replaces the request history and the stored request and response bodies of an earlier version with an empty store; the configuration is kept. Switching back to an earlier version empties it again.
- The control-plane protocol version (
CONTROL_API_VERSION) is now 23. ThinkWatch Lite connects only to a core with the same protocol version, and has to regenerate its protocol types from this release to use it. ThinkWatch Lite 2026.9.17 includes 0.49.0 (protocol 22) and does not connect to 0.50.0: a server used with it stays on 0.49.x until the app is updated to a release that includes 0.50.0.sudo twcore upgrade --version 0.49.1 --restartswitches a server back to 0.49.1. - New in the protocol:
QuotaWindow.credits(a new type,QuotaCredits { total, used, remaining }), the quota windowmonthly, andsession_log_bytesonRequestStartedandHistoryRow. New message code:gw.files.unsupported. /v1/filesand/filesnow answer 404 on the gateway instead of being forwarded to an upstream.
Claude Desktop. Claude Desktop in third-party inference mode, and the Claude Code engine it embeds, expect a few things of a gateway:
HEAD /api/hello, the probe Desktop sends at startup, is answered with 200 by the gateway itself, without a key and without reaching an upstream. It used to be refused with 401.- The client gives up on a stream that is silent for about five minutes. When a client speaking Anthropic Messages has received nothing for 15 seconds, the gateway writes
event: ping, so a converted upstream that is reasoning or queued at length no longer looks like a dead connection. The timer runs from the last bytes the client received, not from the upstream’s, because conversion drops the upstream’s own keep-alive comments. Pings go only between frames, are not recorded in the stored body or counted as output, and are never sent to clients of other formats. After ten minutes without a byte from the upstream the gateway stops pinging, so a half-open upstream connection ends on the client’s timer instead of lasting forever. /v1/modelsanswers Anthropic clients in the Anthropic shape:type,display_name,created_at,has_more,first_idandlast_id, and, for models whose id reads as Claude,anthropic_family_tier, which Desktop uses to resolve aliases such assonnet. The OpenAI fields stay on the objects, because clients that send onlyx-api-keyare also read as Anthropic.
DeepSeek Harness. DeepSeek Harness sends DeepSeek’s extensions to whichever base URL it is given: top-level
dsh_*fields, among them a session log of the whole conversation of up to 8 MiB per request,role: systementries inmessages,tool_additionandtool_removalblocks withdefer_loadingtools,thinkingenabled without a budget, andx-deepseek-harness-*headers. A request is recognised by its User-Agent or adsh_*field. To DeepSeek’s official endpoint it goes through unchanged. To any other upstream, every request from it, includingcount_tokens, has thedsh_*fields removed, its system entries appended tosystemin order, the tool change blocks anddefer_loadingremoved and listed on the hop as a conversion lists what it drops, and athinkingwithout a budget rewritten to what the model accepts. Thex-deepseek-harness-*headers are not forwarded, and the tool changes beta is taken out ofanthropic-beta, with every line of a header sent on several lines kept.session_log_bytesgives the size of the session log a request carried, so the app can show that it carried one./v1/filesand/filesanswer 404: a file uploaded to the upstream picked at that moment cannot be used by a later request routed to another, and the 404 makes DeepSeek Harness send images inline.GLM Coding Plan quota. An upstream whose
base_urlhost isapi.z.aioropen.bigmodel.cnis read as a GLM Coding Plan upstream. Its responses carry no quota headers, soGET /quotaasks the account’s quota endpoint for each enabled one that is due (at most once a minute, waiting up to 5 seconds), and traffic to one asks at most every 5 minutes; failures back off from 30 seconds to 5 minutes. Windows are5h,weeklyandmonthly(the monthly MCP calls of the older plans), told apart by their length, and a reset time beyond a window’s length is dropped. Credit-based plans reportcreditswith the three numbers the endpoint gives. A 429 carrying a used-up quota code marks the window used up and emitsQuotaExhausted. A key without a plan is left alone for an hour, and a key that had a quota needs two no-plan answers in a row before it is believed, because a transient server error returns the same code.Antigravity CLI. Antigravity CLI sends the Go genai SDK’s default user agent. A user agent that starts with
google-genai-sdkand carries Google’s internal Go toolchain (gl-go/followed bycl/) is now reported as the clientantigravity-cli. Programs built on the SDK with a public Go release are not.Listening on an interface that is not there yet. When
listen.gateway.bindnames a network interface that does not exist or is offline, or an address that is not on this machine yet, twcore used to exit at startup. The gateway now listens on 127.0.0.1 only and says why in the listen status (gw.listen.no_such_nic,gw.listen.nic_offline), asks the system again every three seconds, adds the interface once it appears, moves to its new address when it changes, and drops back to loopback when it goes away. This lets the desktop app on Windows bind the gateway to the WSL adapter, which appears only once WSL has started and changes address on every restart. A loopback that cannot be bound is still fatal at startup. -
ThinkWatch Core v0.49.1 GitHub ↗
This release changes only the server guide, docs/server.md and its Chinese version, so that it describes the desktop app’s side of a remote connection as the app behaves. The binaries behave as in 0.49.0.
Upgrade notes
- The control-plane protocol version is unchanged (22) and so is the request store’s schema (20). A server on 0.49.0 does not need to upgrade, and ThinkWatch Lite 2026.9.17, which includes 0.49.0, connects to 0.49.0 and 0.49.1 alike.
Connecting from the desktop app. The guide names the settings section Connection and the Add remote connection dialog, which also asks for a name. Test connection tests without saving, Save and switch tests before it saves and asks before it switches, and Save stores the connection untested; switching to a remote connection always tests it first and stays on the current connection if the test fails.
Which version a server runs. The app connects only to a core that speaks its control-plane protocol version, which in practice means the core version the app includes, and the latest release can be newer than that. The guide therefore installs and switches with
--version <version>, notes thattwcore upgrade --versionalso installs an older release, and says that recent versions of the app (2026.9.17 and later) showsudo twcore upgrade --version <version> --restartwith the version they need when the versions differ. It also says thattwcore upgradeleaves the data alone, while a version whose request store has another schema starts with an empty request history.While connected to a remote core. The guide adds that the app signs ChatGPT accounts in with a device code only, next to the three things a remote connection cannot do.
-
ThinkWatch Lite v2026.9.17 GitHub ↗
Upgrade notes: the bundled core is upgraded to 0.49.0. The request store has a new format: on its first start, the new core empties the request history, the stored request and response bodies and the security log, and statistics such as usage and cost start again from zero. The configuration is kept. A remote server has to run the same core version, 0.49.0 (ThinkWatch Lite 2026.9.16 cannot connect to a server running 0.49.0), so the server and the app are updated together. When the versions differ, the app gives the command to run on the server,
sudo twcore upgrade --version 0.49.0 --restart, which applies whether the installed version is newer or older. When the core on a server is upgraded from 0.47.0, its first start empties the server’s request history as well.Redesigned interface: every page is rebuilt on a common design, with a summary in the page header, skeletons while data loads and a Retry button when a read fails. Turning keys, upstreams, security rules and similar items on or off can be undone from the notification that follows. On macOS, the sidebar of the main window is translucent.
Overview: the model, cache, latency and upstream rows and the failure marks can be clicked, and open Traffic filtered to what was clicked; a model is matched by its exact name. The model ranking distinguishes Unpriced, No usage and a cost that is actually zero, and when a model has unpriced requests, clicking its cost shows those requests. The latency section now covers the selected time range; it used to cover the current day only.
Traffic: the page header shows the requests in progress, the failed requests and the requests per minute over the last 30 minutes; clicking the failed count shows failed requests only. A request in progress shows how long it has been running, updated every second, and belongs to its session from the moment it starts. Requests that a rule denies before an upstream is chosen, and requests that no upstream could serve, are now recorded, marked Denied or No upstream, and counted as failures. The Routing tab of the request details lists the route, the matched rule, the group, the rules that rewrote the request, the rule that denied it and the reason. The token column shows the total input, including cache reads and cache writes, followed by the output.
Clients and Keys: the Clients page counts clients by state in its header, and its Set up by hand and Not detected sections can be collapsed. On the Keys page, each key shows its requests per hour over the last 24 hours and is marked In progress while one of its requests is running.
Upstreams: each upstream shows its vendor’s logo and its requests per hour over the last 24 hours, with the hours that had failures marked in red; the row menu has a new View traffic item. The email and plan of a ChatGPT account are now read from the saved credential, without a network request, and are the same everywhere; once the sign-in has expired, the account is still named and its quota is no longer queried. Plans carry the names that Codex gives them. A completed ChatGPT sign-in names the account that was signed in, and a completed Z.ai sign-in no longer occasionally leaves out the account name.
Routing: a map at the top of the page shows where requests go, from keys through routes and groups to upstreams. Hovering a table row or a node highlights the paths through it, and a request in progress lights its path according to core’s routing decision. Each rule shows its hits over the last 7 days, a rule without hits in that time is marked No hits, and on the map the lines leaving a route grow thicker with the number of requests. When fewer than 7 days of requests are recorded, the counts are labelled with the span the records cover, such as 5-hour hits.
Security and MCP: the Security page header shows how many protections are enforcing, observing or off, and the exact number of matches in the selected period, broken down by outcome. The log is grouped by day. A single click opens a rule, where a double click was needed before. The MCP page header summarizes the findings by severity and the time of the scan; clicking a server row shows its configuration in each client, with the fields that differ marked.
Settings: an index of the sections on the left follows the scrolling and marks a section with unsaved changes; unsaved changes remain after switching pages. The page header shows the current connection, the gateway address and the version. Uninstall is a confirmation dialog that lists what will be done and reports the result of each step, and its title says when a step failed.
Command palette: ⌘K opens the command palette (on Windows and Linux, Ctrl takes the place of ⌘ here and below). It opens pages; runs actions such as creating items, the inference and connection tests, the routing dry run and switching connections; and searches upstreams, keys, routes, clients, Settings sections, the requests already loaded and more. ⌘1 to ⌘9 switch pages, ⌘R now refreshes the data of the current page, and ? outside a text field shows the keyboard shortcuts.
Menu bar: the running time under In Progress in the menu is now computed by core, so a difference between the clocks of the two machines no longer affects it when the app is connected to a remote core. In the English interface, the text in the Quota and Today parts of the macOS menu bar menu no longer overlaps, and the tray on Windows and Linux names the local connection This computer.
-
ThinkWatch Core v0.49.0 GitHub ↗
This release states how far back the route and rule hit counts reach, and reports the signed-in account in the status of a ChatGPT sign-in. Release pages now begin with a summary of the release and list the downloads for each platform with the commands for a server installation.
Upgrade notes
- The control-plane protocol version (
CONTROL_API_VERSION) is now 22. ThinkWatch Lite connects only to a core with the same protocol version, so a server has to run the core version that the app includes. ThinkWatch Lite 2026.9.16 includes 0.47.0 (protocol 20) and does not connect to 0.49.0: a server used with it stays on 0.47.0 until the app is updated to a release that includes 0.49.0.sudo twcore upgrade --version 0.47.0 --restartswitches a server back to 0.47.0. GET /summary/routesreturns an object,RouteStats { covered_since_ms, routes }, instead of an array ofRouteHits;routesholds what the array held.ChatgptLoginStatus.planis removed.ChatgptLoginStatus.accountcarries the email and the plan.- The request store’s schema is unchanged (20), so 0.48.0 and 0.49.0 keep each other’s request history. Upgrading from 0.47.0 or an earlier version empties the request history on the first start, as 0.48.0 does.
How far back the hit counts reach.
GET /summary/routescounts the stored requests, so its counts reach back only as far as the stored history. On a new installation, after an upgrade that emptied the request history, or when requests are kept for fewer days than the window, the oldest stored request is later than the start of the window, and before it a rule without hits is unknown rather than unused.covered_since_msis the later of the window’s start and the time the oldest stored request started; when it is set, it lies within the window. It isnullwhen the stored history covers no part of the window: the store holds no request, or the oldest one started at or after the end of the window.The account after a ChatGPT sign-in. The status of a ChatGPT sign-in (
GET /chatgpt/login/{id}) reports the account it signed in to inaccount, with the email and the plan. It is the same block asProviderView.oauth.accountand is read from the same access token, the one the sign-in saves in the configuration, so the sign-in and the upstream report the same account.Release pages. A release is titled “ThinkWatch Core” followed by its version. Its page begins with a summary of the release when one is written, followed by a table of the files for each platform, the commands that install the version on a Linux server and switch an existing installation to it, how to verify a download against its
.sha256, and the list of pull requests. The pages of 0.47.0 and 0.48.0 carry the same sections.Documentation. The README describes installing the prebuilt binaries, the Linux install script and running
twcoreas a service on a server, and lists all of the crates, grouped by dependency order, including the three that ThinkWatch Enterprise depends on (tw-dialect,tw-guardandtw-breaker). - The control-plane protocol version (
-
ThinkWatch Core v0.48.0 GitHub ↗
This release reports more of what core knows about each request: the session and the route from the moment a request starts, a snapshot of running requests that a client can replay, and how often each route and rule was used. It also names the account signed in on a ChatGPT account upstream, counts the requests in each cost group that have no price or no usage, and gives the security log totals for the whole window.
Upgrade notes
- The control-plane protocol version (
CONTROL_API_VERSION) is now 21. ThinkWatch Lite connects only to a core with the same protocol version, so a server has to run the core version that the app includes. ThinkWatch Lite 2026.9.16 includes 0.47.0 (protocol 20) and does not connect to 0.48.0: a server used with it stays on 0.47.0 until the app is updated to a release that includes 0.48.0.sudo twcore upgrade --version 0.47.0 --restartswitches a server back to 0.47.0. - The format of the request history changed (request store schema 20). On its first start, 0.48.0 replaces the request history and the stored request and response bodies of an earlier version with an empty store; the configuration is kept. Switching back to an earlier version empties it again.
GET /in-flightreturns an object,InFlight { now_ms, requests: [{ id, events }] }, instead of an array of start events.RequestStarted.session_fpis replaced bysession.ChatgptUsage.emailandChatgptUsage.planare removed; the signed-in account is inProviderView.oauth.account.
Session and route from the start. The gateway assigns a request to a session when the request starts (the same conversation, within 30 minutes of its previous request), and
RequestStartedcarries that session id, the one the request is stored under. The first routing phase now finishes before the start event, soRequestStartedalso carries the route, the rule, the strategy group and the rules that rewrote the request.RequestRoutedcarries the final record, including rewrites from the second phase anddenied_byfor a rule that denied the request there. The stored routing record has the same fields and is written from the start event, so a request that ends before routing completes still shows the route and rule it matched.Requests a rule decided. A request that a rule denies in the first routing phase, or that none of the upstreams a rule chose can serve, used to produce no events and no stored row. It now produces start, routed and failed events and a row, and counts among the requests and failures in
/summary; a denial in the second phase now also produces the routed event. Requests that fail before a rule decides (authentication, model admission, no matching rule) are still not recorded.Live views.
GET /in-flightreturns core’s clock and, for each running request, the events seen so far in order, so that a client connecting while requests are running can replay them and compute elapsed time from core’s timestamps. The running requests inGET /liveaddelapsed_ms,session,route,rule,groupandupstream.Rule hit counts.
GET /summary/routesreports, for a time window (today by default), each route’s requests, failures and last request, and for each rule the requests it decided, the requests it applied to, their failures and the last one. The counts come from the stored routing records, not from the current configuration.Cost groups and the security log. Every group returned by
GET /summary/buckets/byreportsunpriced_requestsandno_usage_requests, so a group with no cost can be told apart as unpriced, missing usage or free.GET /security/eventsaddstotalandby_outcome(recorded, replaced, cut, blocked): the counts for everything the query’s guard and time window match, the same on every page.ChatGPT account upstreams.
ProviderView.oauth.accountgives the email and plan of the account signed in on a ChatGPT account upstream. Both are read from the access token core already keeps with the credential: nothing is requested from the network, nothing new is stored, and the account and user ids in the token are not exposed. The plan is reported as the backend names it, including plans that core does not list.Releases. A release is now published by a single job once every platform has been built and each file matches its SHA-256. Previously the first platform to finish created the release, and made it the latest, while the others were still building, and each build job added another copy of the release notes. The server guide, docs/server.md, states that ThinkWatch Lite keeps a remote connection’s key in a private file in its data directory, and applies to the app on macOS, Windows and Linux. The crates carry homepage and documentation links, and
twcore --helpbegins with the binary’s description.Tests. Tests no longer hand core a port that was just released, which could make tests that bind ports, such as those of the remote control port, fail intermittently on Linux.
- The control-plane protocol version (
-
ThinkWatch Lite v2026.9.16 GitHub ↗
Upgrade note: the proxy configuration now rejects misspelled field names, such as
typwritten fortype. Such fields used to be ignored; after the upgrade, a configuration that contains one does not pass validation and the gateway does not start. A proxy configuration that was edited by hand should be checked before upgrading.Remote connections: the app can now connect to a core running on another machine, such as a Linux server. A remote connection is added in Settings → Connection with its address, port and key, and once the connection test succeeds the app can switch to it. Installing and configuring the server side is described in docs/server.md in the ThinkWatch-Core repository. The app connects to one core at a time: while it is connected to a remote core, the local core is stopped and its data is kept. When switching, the clients already connected on this computer can be pointed at the server in the same step.
When a connection fails, the interface remains usable. At launch the app waits up to about 8 seconds, then shows the Not connected page, where the connection can be retried, edited or switched to this computer. If the connection drops while the app is running, pages become read-only and a notice is shown once; it is removed automatically when the connection returns. The menu in the menu bar, or in the tray on Windows and Linux, has a new Connection submenu. When Option is held at launch (Alt on Windows), or after two launches in a row did not finish, the connection choice is shown first.
In remote mode: the Clients and MCP pages check and change the client configuration on this computer; Settings is divided into two groups, the app on this computer and the configuration on the server; ChatGPT accounts sign in with a device code only; the diagnostic bundle is not offered.
The control channel between the app and core now begins with an encrypted handshake, and local and remote connections use the same key, which is written in the configuration file. Connecting clients, MCP management and the client configuration scan are now carried out by the app itself.
-
ThinkWatch Core v0.47.0 GitHub ↗
This release adds a remote control port, through which ThinkWatch Lite on another machine can manage a core running on a server, together with an install script, a systemd unit and
twcore upgradefor runningtwcoreas a service on Linux. ThinkWatch Lite 2026.9.16 includes this version.Upgrade notes
- A misspelled field under
proxies(for exampletyp:fortype:) is now a configuration error; it used to be ignored. A configuration that contains one does not passtwcore check, and core does not start with it. - Every connection to the control plane now begins with a Noise handshake, and control requests no longer use a bearer token. HTTP clients such as curl cannot call the control plane directly;
twcore callsends a request through the handshake. - Client adoption, MCP editing and the client configuration scan moved to ThinkWatch Lite, which runs on the machine where those files are. The
twcore scanandtwcore clientssubcommands and the related control-plane endpoints and events are removed;POST /clients/{id}/keyremains.
Control key.
listen.control.keyinconfig.yamlholds the key, andtwcore serveadds one to a configuration that has none. Every control transport (the unix socket, the loopback port on Windows and the remote port) starts with aNoise_NNpsk0_25519_ChaChaPoly_BLAKE2shandshake keyed by it, and a wrong key is refused before any HTTP.twcore control-keyprints the key;--rotatereplaces it and closes the connections made with the previous one. The key is masked inGET /config, in the configuration history and in diagnostic bundles, and it cannot be changed through the control plane.Remote control port.
listen.control.remoteopens a TCP port for ThinkWatch Lite on another machine, in addition to the local channel.twcore initwrites the section disabled, with a random port from 20000–32000.twcore remote enable [--bind B] [--port N] [--allow CIDR],twcore remote disableandtwcore remote showchange and report it, and a running core applies a change within a second. A source outsideallow_from(the private ranges by default; loopback is not added automatically) is closed before the handshake, and a source with five failed handshakes within 60 seconds is ignored for 60 seconds. Remote connections cannot shut core down, download diagnostics or changelisten.control.Server deployment.
scripts/install.shinstallstwcoreon Linux (x86_64, aarch64) as a systemd service, with athinkwatchsystem user, data in/var/lib/thinkwatchand environment variables in/etc/thinkwatch/env; it checks the SHA-256 and never starts or restarts the service. Releases now includetwcore-<target>.tar.gzfor both Linux targets, holding the binary,twcore.serviceandLICENSE.twcore upgrade [--check] [--restart] [--version X.Y.Z]replaces a separately installedtwcorewith a release from GitHub after checking its SHA-256 and running it once; it does not replace the copy inside ThinkWatch Lite, which the app updates itself. The guide is docs/server.md.Configuration reference. docs/config.md (and
docs/config.zh-CN.md) describes every section and field ofconfig.yaml. The field tables and the lists of built-in rules are generated from the code, and a test fails when the manual and the code disagree.Fix. On Linux, changing
listen.gateway.bindbetweenalland a single address on the same port now takes effect without a restart; it used to be reported as a port in use by another program. - A misspelled field under
-
ThinkWatch Lite v2026.9.15 中文 GitHub ↗
升级须知:请求记录会清空。上游请求头中如写有 {{client}},升级前请先删除,该写法已不再支持,否则配置无法通过校验、网关不会启动。
新增 Linux 版本(x86_64 与 aarch64 的 AppImage),支持应用内更新。可用一行命令安装:curl -fsSL https://github.com/ThinkWatchProject/ThinkWatch-Lite/releases/latest/download/install.sh | sh
安全页新增三项防护:请求中的隐藏字符(标签字符、双向控制符)、内容规则(内置规则可逐条启停和改处置,另可添加「包含」或「正则」自定义规则)、输出长度上限。三项均可选择关闭、仅记录或拦截,命中会记入安全日志、请求详情和概览。
安全加固:转发时不再跟随上游的重定向,避免上游凭据被带到其他地址;WebSocket 请求中的网关密钥不再发往上游;配置、请求记录等含密钥的文件在创建时即只允许本人读写;工具调用审查可识别更多 SSE 写法;窗口权限按需收紧。
提醒改为按当前状态核对:应用启动或重新连接网关时,已存在的凭据失效、配置未通过校验等问题会补发提醒,已解决的会自动收起,同一问题不会重复提醒。系统通知和界面中的错误原因均显示为中文。
其他:重放请求与正常转发使用同一出站设置;菜单栏的进行中请求和生成速率改由网关统计;网关已在运行时重新打开窗口不再播放启动动画;上游表单的请求头列表第一行显示密钥所用的认证头。
-
ThinkWatch Enterprise v2.0.0 GitHub ↗
Callers now get errors in their own API’s format, and an upstream that refuses a request no longer takes a model’s other routes down with it. The gateway also speaks two more client protocols: Gemini, and the Responses API over a WebSocket. Cached input is billed at cache prices, and a request with no usage report is billed on an estimate instead of at zero. The TOTP requirement, which never took effect before, is now enforced. This is a major release because error bodies, the content-filter preset ids and the TOTP behaviour all change in ways a client or a script can notice.
Read before upgrading#
-
Check
security.totp_requiredbefore you upgrade. In 1.x this setting never took effect: it is stored as a boolean and was read as a string, so it always read as off. From 2.0.0 it is enforced by the server. Find out what it is set to:SELECT value FROM system_settings WHERE key = 'security.totp_required';If it is
true, every console user without TOTP, super admins included, is held at a TOTP setup screen on their next request, and sessions that are already open are held too. Until they set up TOTP, every console and admin endpoint answers 403totp_enrollment_required, except/api/auth/me, logout,register-keyand the TOTP status/setup/verify-setup calls. Setting up TOTP releases the session straight away. API keys are not affected: gateway, MCP and consoletw-key traffic keeps working. While the setting is on,POST /api/auth/totp/disableis refused with 400. The setting must now be a JSON boolean; a string such as"true"is refused on save. -
Gateway error bodies follow each client API’s own format. Status codes and
Retry-Afterare unchanged. Code that reads the errortypeneeds updating:- Chat Completions and Responses keep the
{"error": {"message", "type", …}}shape, buttypeis now OpenAI’s value for the status, not a ThinkWatch tag:authentication_error(401),permission_error(403),not_found_error(404),rate_limit_error(429),invalid_request_error(other 4xx) andserver_error(5xx). The old tags are gone:rate_limited,policy_blocked,provider_http_error,provider_error,provider_timeout,transform_error,network_errorandauth_error. A policy block is nowpermission_errorwith status 403. - Anthropic Messages clients get Anthropic’s body,
{"type": "error", "error": {"type", "message"}}, with Anthropic’s type names (rate_limit_error,overloaded_error,api_error, …). - Gemini clients get Google’s body,
{"error": {"code", "message", "status"}}. - Once a stream has started, a Responses client gets a
response.failedevent, where before it got a Chat-style error frame that SDKs skip, so the stream just stopped. An Anthropic client gets anerrorevent whose type follows the status. - The
error_typefield ingateway_logsand the metric labels are unchanged.
- Chat Completions and Responses keep the
-
Only upstream failures fail over or count against a route’s breaker. In 1.x every non-2xx except 401, 403 and 429 was retried on the model’s other routes and counted as a failure on each of them, so one malformed request could open the breakers on all of a model’s routes.
- Tried on another route and counted against this one: 5xx, 408, 429, 401 and 403 (the upstream refused the gateway’s own credential), timeouts, network errors and unreadable responses.
- Returned to the caller straight away, and counted as the upstream working: every other 4xx.
- What the caller sees: such a 4xx comes back with its own status and the upstream’s reason. In 1.x it came back as a 502 after every route had been tried.
- An upstream 5xx comes back with the upstream’s status (500, 503, …) rather than a blanket 502.
- An upstream timeout is now 504.
- Streams follow the same rule. A stream cut by tool-call inspection no longer counts against the route.
-
Cached input is billed at cache prices. In 1.x cache reads and writes were billed, and debited from budgets and weighted rate limits, as full-price input.
modelsgains three weights,cache_read_weight,cache_write_weightandcache_write_1h_weight. When a weight is unset, it isinput_weighttimes Anthropic’s ratio: 0.1× for a read, 1.25× for a write and 2× for a one-hour write. What this changes:- Traffic with many cache reads (Claude Code, for instance) costs much less than it did.
- Traffic that writes to the cache costs a little more.
- Older OpenAI models discount cache reads less (0.5× or 0.25×). Set the weights on those models yourself.
input_tokensin the log is still the whole input. The log detail gainscache_read_tokens,cache_write_tokensandcache_write_1h.
-
Output length limits now apply to streams. In 1.x,
max_lengthoutput guardrails checked only whole responses, so streamed answers were never checked. Now the frame that would cross the limit is not sent, and the stream ends with an error in the caller’s format. A response served from the cache is also checked against the limit in force. If you set a limit, streamed answers that used to go through can now be cut off. -
Content-filter preset groups are renamed. The groups are now
injection,personaandchinese; they used to bebasic,strictandchinese. This matters only if you call the presets endpoint by group id. Rules you have already added are copies and are not affected. Other changes to the filter:- A rule with an empty pattern is now refused on save.
- Each text part of a message is scanned separately, so a pattern no longer matches across two parts.
- The engine is now shared with ThinkWatch-Core’s
tw-guard. The stored format and the admin API are unchanged.
-
Requests with no usage report are billed on an estimate. In 1.x such a request was billed at zero. This happens when an upstream ignores the request for usage, or when the caller leaves before the final chunk arrives. The estimate is:
- input: about four bytes of the request per token, not counting images and files;
- output: the answer that actually arrived.
Estimated rows carry
usage_estimated: truein their detail and count ingateway_usage_estimated_total. A request with no answer at all is still billed at zero. -
Clients that leave early are now logged. In 1.x a client that disconnected before its response existed left no
gateway_logsrow at all. That covers leaving during auth, limits or routing, or while waiting for a whole (not streamed) answer. Such a request now writes one row: status 499,stream_outcome: client_cancelled,cancelled_before: response, no tokens and no cost. Expect more 499 rows in dashboards and log forwarders. A new counter,gateway_cancelled_before_response_total, counts them.
Database changes#
Both apply on their own at startup, as every schema change does, and both are additive:
modelsgains three nullable columns:cache_read_weight,cache_write_weightandcache_write_1h_weight(ALTER TABLE … ADD COLUMN IF NOT EXISTS,CHECK (>= 0)).system_settingsgets anauth.default_rolerow, seeded empty (no role). Existing rows are left alone (ON CONFLICT DO NOTHING).
Neither is irreversible. A 1.1.0 server runs against the upgraded database: it ignores the new columns and the new setting. What a 1.1.0 server cannot do is price cache tokens from the weights.
Added#
-
Gemini clients. New endpoints:
POST /v1beta/models/{model}:generateContentand:streamGenerateContent, also served under/v1/models/…;GET /v1beta/models, which lists models in Gemini’s format.
These requests get the same limits, budgets, filters, routing with failover, format conversion, inspection, billing and audit as every other endpoint. A Gemini upstream gets the request as it was sent. A stream comes back as SSE with
alt=sse, and as Gemini’s JSON array without it.:countTokensand:embedContentare refused with 400. -
The Responses API over a WebSocket. Connect to
GET /v1/responseswithUpgrade: websocket.- Each
response.createframe is handled like a streamedPOST /v1/responses, with its own limits, routing, billing and audit row. - Turns on one connection run in order. A refused turn fails with
response.failed, and the connection stays open. - The connection keeps its latest response. That lets a turn continue
from it with
previous_response_id, even withstore: false, which is how Codex works, and against any upstream format. - A new counter,
gateway_responses_ws_connections_total, counts connections.
- Each
-
More places to put an API key. Gateway keys are also accepted in
x-api-key(Anthropic SDKs),x-goog-api-keyand?key=(Gemini SDKs), as well asAuthorization: Bearer. Headers are checked first. A key given in the query string is never sent upstream. -
Hidden-text audit events show what the text says. Each item in
foundgainsrevealed, the ASCII that the hidden tag characters spell. -
auth.default_rolecan be set. It is the role that newly registered users and SSO users get. In 1.x, setting it through the admin API reported success but changed nothing, because the setting row did not exist. -
Model editor fields for the three cache weights. Each placeholder shows the value used when the field is left empty.
Changed#
- Default output length for upstreams that require
max_tokens. When the caller sets none, the gateway now sends 32000 for Claude models and 8192 for other models. It used to send 4096, which cut Claude answers short. - More upstreams count as the vendor’s own endpoint. DeepSeek,
Moonshot, Zhipu/Z.ai, DashScope, xAI and
*.amazonaws.comare now recognised, and the check reads the parsed host. A relay URL such ashttps://relay/api.openai.comno longer passes as official. Official endpoints are stricter about request parameters, so the gateway drops or renames some parameters before sending to them. - Hidden-text scanning uses ThinkWatch-Core’s
tw-guard. Same scope, same actions, and nothing is stripped. - Requests forwarded in their own format lose ThinkWatch’s reasoning
signatures. A
tw1.signature written by an earlier format conversion is removed, because Anthropic rejects it. The upstream’s own signatures are kept. - Core crates:
tw-dialect,tw-guardandtw-breakerat ThinkWatch-Core v0.43.0. The code only this edition used (at-rest crypto, SigV4 signing, the gateway error type) moved into this repository. It works the same, and stored secrets decrypt as before. - The server’s SQL moved from the request handlers into repository modules (catalog, dashboard, limits, log forwarding, identity, access, MCP). Every statement is unchanged. New integration tests cover these endpoints and pass on both the old and the new code.
- CI runs on pull requests into
dev, including the whole integration suite against Postgres, Redis and ClickHouse.
Fixed#
- Revoking a user’s default MCP connection always failed with a 500, and the account could not be revoked. The newest remaining account is now made the default.
-
-
ThinkWatch Enterprise v1.1.0 GitHub ↗
The gateway stops rebuilding every request as a chat-shaped message. A request whose route speaks the caller’s own format goes out as the caller sent it; one that crosses formats is converted by ThinkWatch-Core, the same layer the desktop edition uses. Tools, tool choice, system prompt blocks,
metadataandcache_controlnow reach the upstream, where they used to be dropped. Two new checks guard what goes in and out: tool calls an upstream returns, and invisible characters in what a caller sends.Read before upgrading#
- Anthropic routes record more prompt tokens for the same work.
Prompt tokens now count the same for every upstream: plain input plus
cache reads and cache writes, which is OpenAI’s definition.
Anthropic’s own
input_tokensleaves the cached part out, so on those routes prompt tokens, cost and budget use all go up. The price model still charges every prompt token at one rate. - Requests in the upstream’s own format are forwarded as sent. Every
field the caller sends reaches the upstream, along with the caller’s
anthropic-betaandanthropic-versionheaders. Only the model name changes, and PII is swapped for placeholders. An OpenAI-compatible upstream that rejects fields it does not know may now refuse requests that used to succeed, because those fields were stripped before. Try your upstreams with the clients you actually run. - Two checks are on by default, and neither blocks anything yet.
- Tool-call inspection starts in
observemode. - The hidden-character check starts in
warnmode. - Both write audit events, so expect new entries in the audit log and in anything subscribed to it.
- Nothing is refused until you switch to
enforceorblock.
- Tool-call inspection starts in
- The response cache starts cold. The cache key now covers the whole request, so entries written by 1.0.2 are never hit again. They expire on their own.
- One PII value gets one placeholder. Within a request, the same
e-mail address is
{{EMAIL_1}}wherever it appears. It used to get a new number each time, so a model saw one person as several. Saving a PII pattern now also requires the placeholder prefix to be letters, digits or underscores.
Added#
- Tool-call inspection (
security.tool_inspection).- Why: an upstream writes the response, so it can hand the caller
a tool call the model never made, such as
bash("curl … | sh")appended to an ordinary answer. An agent set to auto-approve then runs it. - Rules: a built-in set of dangerous-command rules. Each can be switched off or given a different action, and you can add your own.
- Modes:
off,observe(records hits and changes nothing on the wire) andenforce. - Enforce on a stream: the stream is cut at the frame that would complete a matching call, and the refusal arrives in the caller’s format.
- Enforce on a whole response or a cache hit: refused with 403
(
policy_blocked). A refused answer is neither cached nor billed. - Audit and metrics: every hit is an audit event
(
gateway.tool_call_flaggedorgateway.tool_call_blocked) and counts ingateway_tool_call_flagged_total. - Admin API:
GET /api/admin/settings/tool-inspection/rulesandPOST /api/admin/settings/tool-inspection/test. - Console: a card on the security page, plus a sandbox tab.
- Why: an upstream writes the response, so it can hand the caller
a tool call the model never made, such as
- Hidden-character check (
security.hidden_text:off,log,warnorblock).- What it looks for:
- Unicode tag characters, which carry an instruction invisibly into the model’s context;
- bidirectional overrides, which make text read differently on screen than it is.
- Where: the caller’s messages and the tool results inside them. The system prompt and the model’s own turns are not checked.
- Not flagged: zero-width joiners (emoji), the zero-width non-joiner (Persian) and Cyrillic.
- Actions:
warnwritesgateway.hidden_text_flagged;blockrefuses with 403 and writesgateway.hidden_text_blocked. - Console: a card on the security page.
- What it looks for:
- Tool-call arguments get their PII back. A model asked to e-mail
a@example.comused to call the tool with{{EMAIL_1}}as the address.
Changed#
- One pipeline for
/v1/chat/completions,/v1/messagesand/v1/responses. A same-format request is forwarded as sent. A cross-format request is converted, and the gateway logs what the target format cannot carry. - Content filtering and PII detection read tool results too. That is where an injected instruction, or customer data pulled in by a tool, usually sits.
- Usage is read off the upstream’s own bytes. A streamed response is no longer held in memory for an accounting pass at the end.
- Chat streams are billed on the upstream’s actual usage. They are always sent asking for it. A caller who did not ask for the usage chunk still does not get one. 1.0.2 estimated the count for these.
- Streams send their headers at once. A caller who leaves while the upstream is still thinking is recorded as cancelled.
- Connectivity tests use the live encoder. A route’s test request is built by the same encoder as real traffic, so a passing test means forwarding works.
- The web console loads data through TanStack Query.
- Signing out, including from another tab, clears everything cached.
- After a change, screens refresh in place.
- Polling pauses while the tab is hidden.
- Core crates come from one pinned tag (ThinkWatch-Core v0.40.0), declared once at the workspace root.
Fixed#
- Requests lost their tools, tool choice and non-text content on the
way upstream (ThinkWatch-Core#50). Claude Code’s system prompt, sent
as an array, was dropped whole, and so was every
cache_controlbreakpoint. Each cached prefix was billed as full-price input. - The response cache could serve the wrong answer. Its key covered
only model, messages and
max_tokens, so two requests that differed only in tools,response_format,seedand so on shared one entry. - A tripped route stayed out until its Redis key expired, roughly four cooldowns. It now gets probed once the cooldown is over, and a success closes it.
- The dashboard showed every AI provider’s breaker as
Closed. It now shows the real state of the provider’s routes, reporting the worst one. - The PII “try patterns” endpoint misreported labels. A pattern
named
CUSTOM_EMAILwas reported asCUSTOM. :latestcould point at amainbuild rather than the release, which is what happened for v1.0.2’s server image. Only the release workflow sets:latestnow.- The web image hung for six hours. Its frontend was built under QEMU for arm64, where Node crashed and the step never returned. It is now built once, natively. The image contents are unchanged.
Security#
- Refreshed the web console’s lockfile to clear 56 Dependabot alerts (1 critical, 23 high). All were transitive, and none of them reached the shipped bundle.
- Anthropic routes record more prompt tokens for the same work.
Prompt tokens now count the same for every upstream: plain input plus
cache reads and cache writes, which is OpenAI’s definition.
Anthropic’s own
-
ThinkWatch Lite v2026.9.14 中文 GitHub ↗
上游的 API 密钥和请求头中的 ${变量名} 现在可以读取用户配置的环境变量。macOS 上读取登录 shell 中的变量,~/.zshrc 等文件中 export 的变量都会生效,此前从程序坞或访达打开时读不到;Windows 上每次启动网关都重新读取系统设置中的环境变量。代理相关的变量和 PATH 不会被读取。修改变量后重新打开应用即可生效。
上游表单中,API 密钥与请求头的说明改为写明读取系统环境变量,请求头可一键插入网关密钥名称和 Access Token。
熔断改用与企业版共用的状态机,行为不变:连续失败三次后熔断 60 秒,再放行一次探测;所有上游都不可用时照常放行,只有一个上游时不熔断。
-
ThinkWatch Lite v2026.9.13 中文 GitHub ↗
自动检查发现新版本时改为发送系统通知,点按通知后再打开更新窗口,不再自行弹出。更新窗口只显示版本号,不再显示更新说明。
-
ThinkWatch Lite v2026.9.12 中文 GitHub ↗
Windows:网关监听选择「局域网」时,可以选用系统中的网卡(如「以太网」),此前保存会失败。快捷键改用 Ctrl(Ctrl+F 搜索、Ctrl+B 收起源列表),字号与字体按 Windows 调整,「访达」「菜单栏」等说法改为 Windows 对应的名称。安装版只运行安装目录中的网关,「关于」中显示的路径也已修正。
开机启动的说明改为一句话,直接写明开启后的效果。
-
ThinkWatch Lite v2026.9.11 中文 GitHub ↗
修复 2026.9.10 中所有修改操作都会失败的问题:保存设置、新建或编辑上游、接管客户端等操作会提示控制面需要凭据,本版恢复正常。
Windows:启动时不再弹出网关的命令行窗口;客户端页中的配置文件路径改为 Windows 的写法,Zed 的配置位置以及 Continue、Gemini CLI 的手动配置步骤按 Windows 给出。
报错信息改为按界面语言显示,系统或外部服务返回的原文除外。
-
ThinkWatch Lite v2026.9.10 中文 GitHub ↗
首次提供 Windows 版,分 x64 与 arm64 两个安装程序,在发布页下载。Windows 版未做代码签名,首次运行时 SmartScreen 会拦截,选择「更多信息」→「仍要运行」即可。
新增 Z.ai / BigModel 账号登录:在上游的服务类型中选择「Z.ai / BigModel 账号(登录)」,在浏览器中授权后,账号下会生成一把 API 密钥并写入这个上游。
客户端页中的 Codex 一行同时对应 Codex 命令行和 ChatGPT 桌面应用内置的 Codex,二者共用同一份配置;接管时一并说明,ChatGPT 桌面应用需要重启才会生效。
修复:更新窗口底部的按钮被窗口边缘截断;在接管对话框中展开「完整改动」后,内容超出对话框。
-
ThinkWatch Lite v2026.9.9 中文 GitHub ↗
升级须知:2026.9.8 及更早的版本无法在应用内安装这一版,需从发布页下载磁盘映像(DMG)手动安装一次,此后的版本可照常在应用内更新;通过 Homebrew 安装的执行 brew upgrade 即可。
推理测速支持全部类型的上游。ChatGPT 账号上游此前无法测速,只会返回一句英文提示,现在与其余上游一样发送探测请求并测量首 token 时间,失败原因按界面语言显示。ChatGPT 账号不接受输出上限,费用预估中的上限显示为「不限」、金额显示为「按实际用量」,合计随之标为无法计算;会推理的模型在探测时留出推理所需的输出额度,避免上限过小导致没有可见的输出。
发布内容只保留磁盘映像及其校验值,应用内更新改为下载与手动安装同一份磁盘映像。
-
ThinkWatch Lite v2026.9.8 中文 GitHub ↗
升级须知:升级后本机的请求记录会清空。部分旧设置项已删除,配置文件中仍有这些项时网关无法启动,应用会进入安全模式;删除这些项后重新启动即可恢复。用 ChatGPT 账号登录的上游写入的「billing: subscription」属于这一类。
界面支持英文,默认跟随系统语言,也可在「设置」中选择;菜单栏、系统通知和错误信息使用同一种语言。外观可选浅色、深色或跟随系统。
导航重新整理为概览、流量、客户端、密钥、上游、路由、安全、MCP 和设置。
流量页将请求与会话合为一张表,可按会话归组;每条请求标出发出它的密钥、应用和来源机器,请求详情中的 JSON 会格式化显示。概览的实时图显示最近十分钟,切换页面后保留所选的时间范围。
路由页以密钥为中心重新设计,规则按顺序编辑;试算会给出请求使用的路由,不经过任何规则的请求也会如实说明。
密钥单独成页,显示完整密钥并可复制,标出由哪个客户端接管生成,可安全地更换;每把密钥可设置可见模型和并发上限,全局并发限制已删除。客户端页用一张表列出各客户端的状态和所用密钥,未检测到的客户端也给出手动配置方法;「全部还原」此前只还原 opencode,现在会还原所有已接管的客户端。
安全页只保留出站脱敏和工具调用审查两项全局防护,匹配规则可查看和自定义,并新增安全日志;MCP 扫描单独成页。
计费方式只分按量计费和不计费,ChatGPT 等订阅账号也按价目表计算费用。上游行可直接展开模型列表,缺失的列表会自动获取;ChatGPT 账号上游显示邮箱和套餐,也可在另一台已登录的设备上输入代码完成登录。
菜单栏改为原生样式,显示今日 token 和今日费用,菜单中可查看额度、进行中的请求和提醒。提醒只保留一个总开关,可标为已读;不再监测磁盘空间。
网关监听的访问范围、网卡、端口与放行网段移入「设置」,保存后生效;日志保留时长也可在「设置」中调整。启动时显示启动画面,网关就绪后进入主界面,各页面随数据变化实时更新。
-
ThinkWatch Lite v2026.9.6 中文 GitHub ↗
可以用 ChatGPT 账号新建上游:在浏览器中完成登录后,订阅额度即可供各类客户端使用。登录前会说明相关风险;账号的额度与额度重置卡在该上游的菜单中查看和使用。
不同格式之间的请求可以互相转换:使用 Anthropic、OpenAI Chat、OpenAI Responses 或 Gemini 格式的客户端,可以使用其中任何一种格式的上游。请求详情中会显示转换情况与未能保留的字段。
新增提醒:订阅额度用完、登录失效、上游拒绝凭据、代理不通、磁盘空间不足等情况记录在工具栏的提醒列表中,需要处理的会弹出系统通知。每一类提醒可在「设置」中选择弹出通知、仅在应用内显示或关闭。
上游页重新设计:上游、代理与价目表分为三个标签,新建与编辑均在对话框中完成;上游可以设置自定义价目表与额外的请求头。
请求列表会标出被取消的请求、部分失败的请求与缺少用量的请求。
-
ThinkWatch Enterprise v1.0.2 GitHub ↗
The shared gateway layer moves out into its own repository, and OIDC learns to accept an ID token whose
audcarries more than the client ID. Mostly a fix release otherwise — the deploy and auth items below are the ones worth reading before you upgrade.Added#
OIDC_ADDITIONAL_TRUSTED_AUDIENCES— a comma-separated allowlist of extra ID-token audiences to trust. Some IdPs put something besides the client ID inaud; Zitadel includes the parent project ID, andopenidconnect’s stock verifier rejects every non-client-ID audience, so login failed outright. Leave it unset and verification is byte-for-byte what it was. Set, it widens exactly one check: the extra audiences on a multi-audience token are matched against this list as exact strings — no wildcards, no prefixes. The client ID must still appear inaudregardless. Documented in.env.example. Thanks to @DaniW42 for the report and the patch (#12).- Automatic upstream protocol detection — the gateway learns an upstream’s dialect and relearns it when the upstream changes, instead of relying on a static guess. Protocol is now a property of the route.
- Reverse-proxy identification — the server recognizes its own reverse proxy, which is what makes the dashboard WebSocket and the API docs work behind nginx.
Changed#
azpis validated whenever it is present, per OIDC Core 3.1.3.7 step 5, which has no audience-count precondition. Previously it was only checked when a multi-audience token made it mandatory, so a single-audience token naming this client withazppointing at a different client was accepted — the IdP stating plainly that the token was authorized for somebody else. This can reject a token that 1.0.1 accepted. If your IdP setsazpto something other than your client ID, logins will start failing and the message will say so; that is the IdP to fix, not this setting.- The shared gateway layer now comes from ThinkWatch-Core as a git dependency rather than living in this tree. Six crates — types, protocol, provider, resilience, crypto, and the JSON secret envelope — are one implementation shared with the desktop edition instead of two copies drifting apart. No runtime or API change; it affects you only if you build from source, where the build now needs network access to resolve that dependency. The published images are unaffected.
- Stored secrets are redacted on read and decrypted on use. Provider credentials no longer travel in plaintext through code paths that merely display or list them.
Fixed#
- Deploy — nginx broke the API docs three separate ways; the dashboard WebSocket wasn’t proxied; the prod stack’s healthchecks were wrong; and the dev stack could overwrite production’s ClickHouse credentials. The last one is the reason to read this list.
- Auth — the console logged itself out after every token refresh. Rate-limit counters could be created without an expiry and sit in Redis forever.
- Models — the gateway no longer offers models the upstream refuses,
no longer imports models it won’t serve, and reports every dialect
that was refused rather than only the last. Provider
base_urltrailing slashes are normalized. - Analytics — spend is attributed to a person, not a UUID.
- Audit — a syslog forwarder formatting fix that had CI red on
clippy 1.98 (#11, thanks @DaniW42). Two more instances of the same
lint, plus a
result_large_errfalse positive in the MCP lifecycle stages, measured rather than boxed: theOkvariant is 296 bytes against theErr’s 152, so boxing saves nothing and adds an allocation. - UI — the brand mark stays inside the collapsed sidebar rail.
Contributing#
CONTRIBUTING.mdand a PR template now state the branch contract that had only lived indocs/operations/release.md: routine work targetsdev, andmainis the release line. A workflow comments on PRs opened againstmainfrom anything other thandevor ahotfix/*branch, because GitHub pre-fills the base with the default branch and walks contributors into it.Added#
- (nothing yet)
Changed#
- (nothing yet)
Fixed#
- (nothing yet)
Removed#
- (nothing yet)
Security#
- (nothing yet)
May 2026
-
ThinkWatch Enterprise v1.0.1 GitHub ↗
Release-pipeline validation. No product change — the published binary, REST surface, MCP wire shapes, and audit semantics are identical to
v1.0.0. Operators pinning1.0.0have no reason to bump; those tracking:latestmove forward.Changed#
- Release workflow — arm64 image builds now run on a native
arm64 runner (
ubuntu-24.04-arm) instead of QEMU emulation. v1.0.0’s server image build took 1h24m; this should drop to ~10 min. Multi-arch manifest assembled by a new merge job viadocker buildx imagetools create. - Node 24 opt-in — workflow sets
FORCE_JAVASCRIPT_ACTIONS_TO_NODE24=trueso allactions/*+docker/*run on Node 24 ahead of GitHub’s 2026-06-02 forced cutover.
- Release workflow — arm64 image builds now run on a native
arm64 runner (
-
ThinkWatch Enterprise v1.0.0 GitHub ↗
Stability commitment. No code delta since
0.5.0— this tag marks the point at which the API surface becomes a SemVer commitment.Changed#
- Versioning policy — from this tag onwards every breaking
change (REST routes, MCP wire shapes, audit-row JSON keys,
database schema, public Rust APIs in published crates) requires
a major bump. Operators chasing the
:latesttag on the GHCR images can do so without surprise.
Notes#
- Docker images cut at this tag receive
:latestfor the first time — the release workflow suppresses:lateston0.xand pre-release tags. Pin the version in production rather than tracking:latestunless you have a controlled rollback path.
- Versioning policy — from this tag onwards every breaking
change (REST routes, MCP wire shapes, audit-row JSON keys,
database schema, public Rust APIs in published crates) requires
a major bump. Operators chasing the
-
ThinkWatch Enterprise v0.4.0
See every byte
- Request + response bodies captured for every gateway and MCP call
- PII redaction at capture; per-column TTL splits hot bodies from cold metadata
- S3-compatible body offload — AWS S3, MinIO, Ceph, or the bundled RustFS
- Bloom-filter substring search across captured bodies
- audit:read_bodies permission gates the body viewer
- MCP downstream SSE — clients can subscribe to tools/call event timelines
- Per-model routing with Auto / Manual modes + drag-to-redistribute traffic bar
- Model-level kill switch — pause one model without the whole provider
- Pre-call budget check rejects requests when caps are already exhausted
- allowed_models enforced uniformly on Anthropic Messages + OpenAI Responses
Full-body audit capture#
ThinkWatch positioned itself as a bastion for AI traffic — but every audit row was metadata-only. Token counts, latency, cost. No record of what the user actually asked or what the model actually answered. Compliance teams couldn’t replay incidents, attest “PII X was sent”, or even tell two near-identical requests apart.
v0.4.0 closes the gap.
gateway_logsnow persistsrequest_body+response_body(with byte counts and abody_capture_statusdimension);mcp_logspersists the paralleltool_arguments+tool_result. ZSTD(6) compression keeps cold storage tractable; six dynamic config toggles (capture toggles per field + a globalbody_max_bytes) let operators dial scope without a redeploy. Cache-hit responses now emit a full audit row instead of disappearing into a hole.Per-column TTL + S3 offload#
Hot evidence and cold metadata have very different retention shapes. Per-column TTL lets you keep the last 30 days of bodies and several years of token/cost rows in the same table, without paying ClickHouse for the cross.
For payloads that exceed
body_max_bytes— 200k-context prompts, MCP file-read tools returning whole files, base64 images — bodies offload to S3 instead of being truncated. Backend is config-driven via six env vars (S3_ENDPOINT_URL,S3_BUCKET, etc.) and works against AWS S3, MinIO, Ceph, or the bundled RustFS container shipped in the dev and prod Docker Compose files. Object keys carry a date prefix (bodies/<table>/yyyy/mm/dd/...) so bucket lifecycle rules can age them off independently. The body-viewer endpoint transparently dereferencess3://cells, so the console doesn’t know offload exists.Bloom-filter substring search#
Captured-body columns ship with a
tokenbf_v1(512, 3, 0)index. Auditors hunting for “which conversations mentioned API key XYZ” or “which tool calls touched /etc/passwd” no longer scan every granule — the bloom filter rejects the vast majority of rows up front. Indexes are sized at ~50 KB per granule, paying for themselves the first time someone needs to answer a real compliance question.A new
audit:read_bodiespermission gates the body viewer in the console and the underlying/api/admin/{gateway,mcp}/logs/{id}/bodyendpoint, so the raw payload is preserved as legal record of intent while still requiring a separate grant to read.MCP downstream streaming#
tools/callwas buffered-only on the client side: even when the upstream emitted progress notifications, the gateway swallowed the timeline and returned a single JSON. v0.4.0 negotiates on theAcceptheader — clients includingtext/event-streamreceive each upstream event (notifications/progress, partials, final response) as a discrete SSE event downstream. JSON-only clients keep the buffered shape for backward compat. The audit pipeline still gets the full event timeline either way.Per-model routing#
Model-route configuration is now organised around the admin’s actual mental model: Auto (latency-cost, balanced, or latency-only) or Manual (a drag-to-redistribute traffic bar across peers). The priority-tier failover concept is gone; replaced with nginx-style flat peers + circuit-breaker bypass. Each model can override the global strategy; the dashboard surfaces
routing-projectionso the UI can show “your manual config will cost $X; auto would cost $Y” before you commit.Three new endpoints back the UI:
PATCH /api/admin/model-routes/batch-weights— the drag bar’s commit endpointGET /api/admin/models/{id}/routing-projection— current vs. auto split + projected $/1M tokensGET /api/admin/models/{id}/route-history— 60 one-minute p50/p95 buckets for the latency sparkline
The decision log captures provider chosen + reason + any fallback for every request, and every
gateway_logsrow stamps the actualupstream_modelso post-mortems are unambiguous.Pre-call budget enforcement#
Budget caps used to fire post-flight: a runaway client whose monthly cap was already at zero could still fire requests, and only the post-call debit would notice — by which point the upstream had already burned tokens. v0.4.0 adds a
check_budgetlifecycle stage that runs before the upstream call. Steady-state “your cap is exhausted, stop sending” is now blocked at the gate; the existing crossing alerts continue to cover the concurrent-burst race.allowed_models on every API surface#
The per-API-key
allowed_modelsallowlist is now enforced uniformly on OpenAI Chat Completions, Anthropic Messages, and OpenAI Responses. The same enforcement is wired into the shared lifecycle pipeline, so future surfaces inherit it automatically — no API surface bypasses the model allowlist.Under the hood#
A shared lifecycle pipeline unifies the AI gateway and MCP gateway request paths behind a common
Surfacetrait.check_limits,check_access,check_budget,record_usageare now stages applied symmetrically on both surfaces, eliminating the historical drift between buffered, streaming, and short-circuit branches. The MCPproxy.rsand gatewayproxy.rsboth got split into submodules, and the 1667-line dashboard component on the frontend was carved into focused subcomponents. -
ThinkWatch Enterprise v0.3.0
MCP, the per-user way
- Per-user upstream credentials — every developer authenticates as themselves to GitHub, Linear, Slack
- One-paste OAuth onboarding via Dynamic Client Registration
- MCP Store with bilingual templates (Linear OAuth seeded out of the box)
- Step-by-step registration wizard + auth-mode-aware edit form
- Three-tier upstream subject resolution (JWT + userinfo + discovery)
- Per-credential Test Connection on /connections
- Per-user tool catalogs — different users see different tools based on upstream permissions
- MCP response cache scoped per (user, account_label) — no cross-user leakage
- Security hardening from review — SSRF, cache, rate limits, audit
Per-user upstream credentials#
The MCP gateway no longer pretends every user is the same upstream service account. Each developer authenticates as themselves — via OAuth or PAT — to GitHub, Linear, Slack, and any other MCP-enabled service. The upstream audit trail finally works end-to-end: tickets are assigned to real people, GitHub issues are created by the engineer who actually filed them. Per-key account overrides let one user maintain multiple identities (personal + work GitHub) and pick which one a given API key uses.
One-paste OAuth onboarding#
Paste an MCP server URL into the registration wizard. ThinkWatch handles Dynamic Client Registration with the upstream OAuth server, runs the auth probe, captures the OAuth client credentials at install time, and walks you through the consent screen. 401/403 from anonymous probes is correctly classified as
auth_required(with an amber status indicator on the catalog tile) rather than a hard failure, so partially-protected servers register cleanly.MCP Store#
A bilingual template registry is now built into the gateway. Templates ship with sensible defaults, the necessary OAuth scopes, and end-user-facing notes. The Linear OAuth template is seeded out of the box; more popular services follow. Display labels disambiguate multi-install templates so a personal GitHub install and a work GitHub install appear as distinct tiles instead of two indistinguishable entries.
Step-by-step registration wizard#
MCP server registration was reworked into a guided wizard with auth-mode-aware screens. The edit form refuses to save fields that don’t belong to the current auth mode, and tool-call / install errors are now surfaced in plain language with actionable next steps instead of raw JSON-RPC error codes.
Three-tier upstream subject resolution#
When an upstream MCP server uses OAuth, ThinkWatch resolves the upstream user identity through a three-tier strategy: parse the JWT if present, fall back to the OAuth userinfo endpoint, fall back again to issuer discovery. The resolved subject becomes the cache and audit key, so two users who share an MCP server stay strictly separated downstream.
Per-credential Test Connection#
The
/connectionspage surfaces a per-credential Test Connection button. Verify your OAuth/PAT actually works against the upstream server before committing to it; the result is structured (auth_ok / auth_required / unreachable / tool_call_ok) rather than a single green/red dot, so you know exactly which step is broken.Security hardening#
Following an internal security review, the MCP path now has SSRF protection on probe URLs, the response cache is scoped per
(user, account_label)so OAuth/PAT data never leaks across users, rate limits are applied at every gateway hop, audit records are emitted on tool-call boundaries, and admin foot-gun guards prevent the most common misconfigurations (verifying static tokens at paste time, blocking obviously-wrong credential combinations).
April 2026
-
ThinkWatch Enterprise v0.2.0
Teams, spending controls, and a programmable API
- Teams — isolate keys, budgets, and analytics per business unit
- Rate limits & budget caps with fail-closed enforcement and spend alerts
- RBAC v2 — custom roles, permission history, one-click import/export
- Tokens moved to HttpOnly cookies — XSS can no longer steal sessions
- Full management API with OpenAPI docs and API-key auth
- Real-time dashboard powered by WebSocket
Teams#
Users and API keys can now be organized into teams. Each team gets its own isolated view of the analytics dashboard, its own user list, and its own scope when assigning API keys. A structured scope picker makes member assignment straightforward.
Spending controls#
A new limits engine enforces rate limits and budget caps at every layer — per user, per API key, per provider, and per MCP server. Limits are configured through a panel embedded directly in each edit dialog. When a budget is exhausted the gateway fails closed — no silent overruns. Budget alert thresholds let you set a warning before the hard cap is hit. Streaming requests are metered correctly; weighted token costs are supported for models with asymmetric pricing.
RBAC v2#
The role system has been completely reworked. Built-in and custom roles share a unified table. Custom roles are fully editable in a CodeMirror editor, support cloning from an existing role as a starting point, and track a full permission history so you can see who changed what and when. Roles can be exported and imported as JSON for environment parity. Role assignments are now scope-aware — a role granted at the team level applies only within that team.
Security hardening#
Access and refresh tokens have been migrated from
localStorageto HttpOnly cookies — the JavaScript-visible session surface is now zero. The refresh endpoint binds each token to the originating client IP, so a stolen cookie cannot be replayed from a different network. The WebSocket dashboard connection now uses a one-shot ticket rather than a long-lived credential.Management API#
The gateway now exposes a programmable management API authenticated by API key. A full OpenAPI specification is served alongside the gateway so you can integrate key lifecycle and provisioning into your own tooling or CI pipelines without touching the console.
-
ThinkWatch Enterprise v0.1.0
Public preview
- Multi-format AI API gateway (OpenAI, Anthropic, Gemini, Bedrock)
- MCP gateway with namespace isolation and tool-level RBAC
- ClickHouse-powered audit logs with multi-channel forwarding
- First-run setup wizard and built-in configuration guide
ThinkWatch is now in public preview. This first tagged release ships the dual-port architecture (gateway
:3000, console:3001), virtual API key lifecycle management, sliding-window rate limiting, circuit breakers, and the unified log explorer.The Helm chart and distroless container images are available alongside Docker Compose for self-hosted deployments.