Skip to content
ThinkWatch

Changelog

ThinkWatch release notes

Release notes for each version of ThinkWatch Enterprise, ThinkWatch Lite and ThinkWatch Core, listed from newest to oldest. Subscribe via RSS.

Latest releases

September 2026

  1. ThinkWatch Enterprise v2.2.0 GitHub ↗

    This release fixes authorization. The gateways never checked ai_gateway:use or mcp_gateway:use, so a user with no role, or with only roles that grant neither, could call every model and MCP tool; a role granted at the scope of one team widened model and tool access platform-wide; and an API key whose allow-list came out empty could call any model. The gateways now require the permission and count only the roles that grant it. The seeded team_manager role now works when granted at team scope, and seven permissions that nothing ever checked are retired. Some users lose gateway access on upgrade: read the first section before deploying. There is one manual database script and no schema, setting, environment variable or Helm value change.

    Read before upgrading#

    • Users without a role that grants gateway use lose gateway access. A request to the AI gateway is refused with 403 unless a role held by the key’s owner grants ai_gateway:use, and a request to the MCP gateway unless one grants mcp_gateway:use. The body names the missing permission, in the error format of the caller’s API. The built-in developer, admin, super_admin and team_manager roles grant both; viewer grants neither. Three groups of users are refused after the upgrade where 2.1.0 let them through:

      • users with no role at all;
      • users whose roles are only viewer, or custom roles without these actions;
      • users whose only gateway role is assigned at team scope (see the third item below).

      Find them before upgrading. This query lists every active user with a live key for a gateway that none of their roles will grant (it counts global assignments and roles attached to the user’s teams, which is what the gateways now read; it does not account for Deny statements):

      WITH grants AS (
        SELECT id AS role_id,
               jsonb_path_exists(policy_document,
                 '$.Statement[*] ? (@.Effect == "Allow").Action[*] ? (@ == "*" || @ == "ai_gateway:*" || @ == "ai_gateway:use")') AS ai,
               jsonb_path_exists(policy_document,
                 '$.Statement[*] ? (@.Effect == "Allow").Action[*] ? (@ == "*" || @ == "mcp_gateway:*" || @ == "mcp_gateway:use")') AS mcp
          FROM rbac_roles
      ),
      held AS (
        SELECT user_id, role_id FROM rbac_role_assignments WHERE scope_kind = 'global'
        UNION
        SELECT tm.user_id, tra.role_id
          FROM team_members tm JOIN team_role_assignments tra USING (team_id)
      ),
      access AS (
        SELECT h.user_id, bool_or(g.ai) AS ai, bool_or(g.mcp) AS mcp
          FROM held h JOIN grants g USING (role_id) GROUP BY h.user_id
      )
      SELECT u.email, s.surface
        FROM api_keys k
        JOIN users u ON u.id = k.user_id
        CROSS JOIN LATERAL unnest(k.surfaces) AS s(surface)
        LEFT JOIN access a ON a.user_id = u.id
       WHERE k.is_active AND k.deleted_at IS NULL
         AND u.is_active AND u.deleted_at IS NULL
         AND NOT COALESCE(CASE s.surface WHEN 'ai_gateway' THEN a.ai
                                         WHEN 'mcp_gateway' THEN a.mcp END, false)
       GROUP BY u.email, s.surface
       ORDER BY u.email, s.surface;

      Give each of them developer (or a custom role with the actions) either at global scope (Users → edit the user’s roles, scope Global), or by attaching the role to a team they belong to (Teams → the team → Roles), which every member of that team inherits. The change takes effect within a minute; each user’s permissions are cached for 60 seconds.

      If SSO users are meant to use the gateway from their first sign-in, set Default Role for New Users (auth.default_role) in Settings to developer. It is empty by default, and it applies only to accounts created after it is set; existing users need the grant above.

    • A role that does not grant gateway use no longer widens model or tool access. Model and MCP tool scopes are now the union over the roles that grant ai_gateway:use / mcp_gateway:use only. A role without those actions (such as viewer) used to count as “unrestricted”, so adding it to a user limited to some models opened every model to them. Users relying on that lose the extra models.

    • A role assigned at team scope no longer grants gateway access. An assignment with scope team:<id> administers that team from the console (team roster, team limits); it no longer contributes models, MCP tools or gateway use, which it used to do for every request the user made, member of the team or not. Roles attached to a team itself (Teams → Roles), which every member inherits, still count. A user whose only gateway role was a team-scoped developer or team_manager needs that role at global scope, or attached to their team. The role assignment editor now says this under the scope picker.

    • An empty model allow-list allows nothing. A key whose allowed_models is [], or whose list shares no model with what its owner’s roles grant, used to call any model; it now calls none, as allowed_mcp_tools: [] already did on the MCP gateway. The console never saves [] (clearing the picker sends null), so only keys written through the API are affected. To find them: SELECT id, name, user_id FROM api_keys WHERE allowed_models = '{}' AND is_active AND deleted_at IS NULL; Set such a key’s list to null to leave it bounded by its owner’s roles only. Model entries still match by prefix, and a key narrowed to gpt-4o-mini under a role granting gpt-4o keeps gpt-4o-mini.

    • Run db/release_migrations/2026-09-30_retire_unchecked_permissions.sql once, after deploying. The server does not apply it. It adds teams:read to the team_manager role and removes the seven retired permissions (see Changed) from every role, system and custom. It is idempotent and changes no access for any other role, since nothing checked the removed keys. Without it, the server still starts and works, but:

      • team_manager keeps the old team:read it cannot use, so a team manager who is not a member of the team they manage still gets 403 opening it, its roster or its roles;
      • every start logs a warning listing the roles that still name retired permissions.

      Reset to defaults on a system role in the console has the same effect for that one role, but does not clean custom roles.

    • A team-scoped rate_limits:write now takes effect. It used to require global scope for every subject, so the seeded team_manager granted at team scope could not touch any limit. It now covers the rate limits and budget caps of users in that team and of the API keys they own, never the holder’s own user or keys; role subjects still need global scope. Anyone holding team_manager (or a custom role with rate_limits:write) at team scope can now change, lift or delete the limits of every member of that team, administrators included. Review team-scoped assignments of these roles before upgrading.

    Security#

    • Gateway use is checked, and a missing grant means no access. See the first three items above: a user with no role, or only roles without gateway use, could call every model and MCP tool; a viewer-style role widened a restricted user to every model; a role granted at the scope of any team, even one the user was not a member of, widened model and tool access platform-wide. An explicit Deny on ai_gateway:use or mcp_gateway:use now closes that gateway. Action wildcards (ai_gateway:*, *) grant it, as they already did for console permissions.
    • A key narrowed inside a prefix grant is no longer unrestricted. The key’s allow-list was intersected with its owner’s role grants entry by entry, as literal strings. A key limited to gpt-4o-mini under a role granting gpt-4o (a prefix) came out with an empty list, which the gateway read as “no restriction”, so the key could call every model. The intersection now keeps an entry of either side that the other side covers, by prefix for models and by <server>__* pattern for MCP tools, and an empty result allows nothing.

    Fixed#

    • The seeded team_manager role works at team scope. It granted team:read, but the team handlers check teams:read, so a team manager got 403 listing teams and opening the team, its roster or its roles. It now grants teams:read, and opening a team accepts teams:read scoped to that team, where it used to require global scope. Existing installations need the release migration above.
    • Bulk disable and delete work on API keys’ limits. The bulk disable and delete routes for rate-limit rules and budget caps rejected every row stored for an API key (api_key_lineage, the kind a key’s limits are stored under) as an unknown subject kind.
    • The console’s effective-permissions preview shows real gateway access. It now counts only roles that grant gateway use and skips team-scoped assignments, as the gateways do, where it used to count every assigned role.

    Changed#

    • Seven permissions that nothing checked are retired: team:read, team:write, logs:read_own, logs:read_team, audit_logs:read_own, audit_logs:read_team and audit_logs:read_all. Every log endpoint, audit logs included, is gated on logs:read_all at global scope, and no own- or team-filtered log view exists, so these grants never did anything. They are gone from the permission catalog, the role editor and the seeded roles. Roles that still name them load, and the server logs a warning at start instead of refusing to boot.
    • Deleting one rate-limit rule or budget cap is bound to the subject in the path. DELETE on /api/admin/limits/{kind}/{id}/rules/{rule_id} and /api/admin/limits/{kind}/{id}/budgets/{cap_id} now answers 404 when the rule or cap belongs to a different subject; it used to delete any row id once the path’s subject was authorized.
    • Creating or editing a key with an explicit allow-list checks it against gateway-granting roles only. An allowed_models or allowed_mcp_tools entry is refused with 400 when no role of the owner that grants the gateway covers it, so a user without a gateway role can no longer save a non-empty list.
  2. ThinkWatch Enterprise v2.1.0 GitHub ↗

    Amazon Bedrock becomes a provider you can run from the console. It authenticates with a Bedrock API key as well as access keys or the instance role, lists its models for import and for the route editor, and has a working Test Connection. Requests converted for Bedrock now keep their prompt-cache breakpoints and Claude’s thinking settings. The route editor takes any upstream model name, and a route the provider refused can be created anyway. ThinkWatch-Core moves to v0.55.0, whose Bedrock layer this edition now shares with the desktop gateway. No database, setting, environment variable or Helm value changes.

    Read before upgrading#

    • Test Connection with a saved provider’s secrets needs providers:update. A test that names a saved provider (provider_id, which is what the Edit dialog sends) used to take only providers:create. A custom role that has providers:create without providers:update can no longer test existing providers; the built-in admin role has both. A test with values typed into the request still takes providers:create. See Security below.
    • Refusing a route the provider does not serve answers model_not_served. POST /api/admin/models/{model_id}/routes still answers 400 when the import probe found no API the provider serves the model on, but error.type is now model_not_served, where it was bad_request. Scripts that matched on bad_request for this case need the new value. The request takes a new optional "force": true to create the route anyway.
    • error_type in gateway_logs for failed requests is always the error’s tag. An upstream HTTP error used to log its debug text, such as ProviderHttpError { status: 502, message: "…" }, and an upstream rate limit a cut-off UpstreamRateLimited { retry_after_secs: Some. They now log ProviderHttpError and UpstreamRateLimited, as the metric labels and streamed requests already did. Dashboards or log forwarder queries that grouped by those strings see them merge into one value each.
    • The server log now carries the upstream’s reply to a 401 or 403, as it already did for other upstream errors. For Bedrock that reply names the AWS account and the IAM principal. Callers still see only “Authentication failed with upstream”, and failover, breakers and gateway_logs treat these errors as before.
    • Prompt caching now works on Bedrock routes, and is billed at cache prices. Requests converted for Bedrock used to lose their cache_control breakpoints, so Claude on Bedrock re-read the whole prompt at full price every turn. Breakpoints now go out as Converse cachePoint blocks, to Claude and Nova models only (others reject them), and the cache reads and writes Bedrock reports come back into usage, one-hour writes included. Traffic with repeated prompts, such as Claude Code, costs much less on Bedrock routes than it did; writes cost a little more, at cache_write_weight / cache_write_1h_weight.
    • Claude’s thinking reaches Bedrock. A request converted for Claude on Bedrock now carries its thinking and effort settings, in the form Anthropic’s API uses, along with the sampling limits thinking imposes (temperature 1, no top_k, top_p at least 0.95). Thinking used to be dropped on the way to Bedrock. Other Bedrock models still get no thinking.

    Added#

    • Bedrock providers can authenticate with a Bedrock API key. Pick Bedrock API Key as the authentication mode and paste the key AWS generated. It is kept like any other provider’s API key, as an encrypted Authorization: Bearer header, and replaced in the provider’s Edit dialog. A Bedrock provider that sends an Authorization header is not SigV4-signed; one without it is signed as before, with access keys or the instance role. Use a long-term key: a short-term one expires within 12 hours.
    • Bedrock models can be imported, and are suggested in the route editor. A Bedrock provider now lists what it can be routed to: the region’s foundation models that can be invoked on demand and answer in text, and the inference profiles AWS defines, such as us.anthropic.claude-sonnet-4-5-20250929-v1:0. A model that is only served through an inference profile, as most current models are, is listed under its profiles’ ids and not its own. Whether the account may call a model is still checked one model at a time when it is imported. The provider’s credential needs bedrock:ListFoundationModels and bedrock:ListInferenceProfiles; the AmazonBedrockLimitedAccess policy a long-term API key is created with allows both.
    • Test Connection for Bedrock providers, however they authenticate: with an API key, access keys or the instance role. It lists the models above, so a wrong credential or a missing permission shows up before any traffic does.
    • Create anyway, for a route the provider refused. When the route editor’s save is refused because the provider does not serve the model, the error now offers Create anyway: the refusal can be stale, or be about the probe’s request rather than the model. The route is created, and its model_route.created audit row records the refusal it overrode in a new refusal_overridden key.

    Changed#

    • The route editor’s upstream model field takes any model name. For a provider that lists its models, the field used to turn into a list to pick from, so a model the listing leaves out could not be routed to from the console. It now suggests the provider’s models as you type, and takes whatever is typed.
    • Bedrock instance-role credentials are cached. They took three IMDSv2 round trips per request; they are now kept until five minutes before they expire, and a burst of requests at expiry makes one trip.
    • Chat Completions requests can switch reasoning with thinking. When a request is converted for another format, thinking.type (enabled / disabled, as DeepSeek, GLM and Kimi write it) now turns reasoning on or off; disabled wins over reasoning_effort.
    • Core crates at ThinkWatch-Core v0.55.0. tw-dialect, tw-guard and tw-breaker move from v0.43.0, and tw-bedrock joins them: SigV4 signing, eventstream unframing, Bedrock’s addresses, region checks and model catalog now come from core. Where credentials come from (provider keys, the instance role) stays in this repository. tw-breaker is unchanged.

    Fixed#

    • Bedrock providers could not be created or edited in the console. The region was checked as a URL, so saving failed with 400 Invalid URL. A Bedrock provider’s base_url is now checked as an AWS region such as us-east-1, and anything else is refused, since the host is built from it. A provider saved with a URL there never reached Bedrock; set its region in the Edit dialog.
    • Bedrock models the account may not call were imported anyway. Bedrock refuses such a model (model access not granted, or an IAM or organization policy that denies it) with a 403, and the import probe read every 403 as a credential problem that says nothing about the model. The route was created and failed on first use. Now, when the region’s control plane accepts the same credential, the refusal is recorded as the model’s, with AWS’s reason, and the model is skipped on import like any other refused one. Once access is granted, re-check the provider’s models.
    • A Bedrock provider whose stored secret key would not decrypt signed with an empty secret, and every request failed with AWS’s SignatureDoesNotMatch. It now refuses its requests with the reason until the keys are saved again, rather than falling back to the instance role, which would call AWS as a different identity. An access key ID saved without a secret is refused the same way.
    • A Bedrock model id that is an ARN could not be routed. Its / went into the request path unescaped, adding a path segment AWS could not route. It is now escaped.
    • A base URL ending in its version segment doubled it. A provider written as https://api.openai.com/v1 sent requests, and the model listing, to /v1/v1/… and got 404s. A trailing version segment (v1, v1beta, …) that the request path starts with is now written once. Base URLs without one are unchanged.
    • Typing into a provider’s API key field and clearing it again wiped the saved key on save. The field sent Bearer with nothing after it. A cleared field now keeps the saved key, as a blank one always did.

    Security#

    • Testing a connection with a saved provider’s secrets takes providers:update. The test fills each header left blank with the saved provider’s secret and sends it to the URL in the request, and providers:create was enough to ask for it. So a user allowed only to create providers could send any saved API key to a server of their own. A test that names a saved provider (provider_id) now takes providers:update, which lets a user point that provider elsewhere anyway; a test without one still takes providers:create. A user with providers:update alone can now use Test Connection in the Edit dialog, which used to answer 403.
  3. ThinkWatch Lite v2026.9.26 GitHub ↗

    Upgrade notes:

    • The bundled core is still 0.56.0 (control-plane protocol 30), so a remote server on 0.56.0 keeps working.
    • Configuration and request history are kept.

    MCP page: the servers table now has a column only for the clients found on this computer: those whose own directory or MCP file exists, or whose MCP file was chosen by hand. It used to draw all eight supported clients, which did not fit the window. The clients that were not detected are listed in one line under the table, where Choose location… points the app at a client installed somewhere else.

    Toolbar: the page name is always shown in the window toolbar, next to the sidebar toggle, instead of at the top of each page.

  4. ThinkWatch Lite v2026.9.25 GitHub ↗

    Upgrade notes:

    • The bundled core is now 0.56.0, which speaks control-plane protocol 30. A remote server has to run core 0.56.0 too: this version does not connect to 0.55.x, and 2026.9.24 and earlier do not connect to 0.56.0.
    • Request history is kept.
    • Strategy groups no longer have the sticky-session switch. A configuration in which it was turned off contains session_affinity: false under that group; core 0.56.0 refuses such a configuration and the app starts in safe mode, naming the line. Delete the line.

    Failover:

    • An upstream can answer, open a stream and then report an error before sending anything, such as Anthropic’s overloaded_error or a response.failed when a ChatGPT account’s usage limit is reached. The client used to receive that error. A streamed answer is now held until its first content arrives, and an error before that point sends the request to the next upstream of the route. What was received while waiting is passed on unchanged.
    • Upstreams that reject the credential (401, 403), report no balance (402) or do not have the model (404) now send the request to the next upstream as well. A 400 or 422 does so only when the upstream says the balance or quota is used up or the model is unavailable; other client errors still go straight back. When no upstream is left, the last upstream’s own error reaches the client.
    • How long a failing upstream is paused depends on the reason it gives: an insufficient balance pauses it for 30 minutes, a used-up quota until the reset time the upstream names (1 hour when it names none), a rate limit for the wait it asks for (at most 1 hour). Failures without a stated reason pause it after 3 in a row, for 1 minute, doubling with each further pause up to 10 minutes.
    • Settings › Failover sets all of these, including how long to wait for a streamed answer’s first content (15 seconds).

    Prompt cache across a conversation:

    • A conversation now stays with the upstream that answered it, including one that took over after a failover: always within a turn, and in later turns while its prompt cache is still warm. Load-balance groups spread only new conversations.
    • The routing decision made at the start of a turn holds until the turn ends, so a rule on input size or images no longer switches upstreams halfway through an agent’s work.
    • The request details’ Routing tab says when a request kept its turn’s route or stayed with the upstream that answered before.

    Import links: a relay or vendor can hand out a link, thinkwatch://import?… or https://thinkwat.ch/import#…, that pre-fills a new upstream with a name, base URL, protocol, API key and model list. The app shows the settings and the host that will receive requests and the key, and writes nothing and contacts nothing until Create is chosen. A link only adds one upstream: it cannot change existing upstreams, headers, proxies, pricing or routes, and a key that refers to an environment variable is rejected. The parameters and a link builder are in Import links.

    Token counts: Claude Code asks for a token count on every turn, which only an upstream in Anthropic’s format can answer. Routed to an upstream of another format, or to a relay that does not implement it, the count used to fail; the gateway now answers with its own estimate, sends nothing to that upstream, and shows it in the attempt chain as estimated locally.

    Reasoning after a switch of account: reasoning that one account produced is sealed for that account, and another account or key refuses a request that carries it back. After a failover between accounts such a request is now sent again without it, and later turns of the conversation leave it out from the start.

  5. ThinkWatch Core v0.56.0 GitHub ↗

    This release keeps a conversation on the upstream that holds its prompt cache and decides a turn’s route once, moves a request to the next upstream when a stream fails before its first content, sets an upstream aside for as long as the reason it gives calls for, answers token counts that no upstream can serve with a local estimate, and resends a request whose reasoning another account sealed.

    Upgrade notes

    • The control-plane protocol version (CONTROL_API_VERSION) is now 30. ThinkWatch Lite connects only to a core with the same protocol version, and has to regenerate its protocol types from this release to use it. ThinkWatch Lite 2026.9.24 includes 0.55.2 (protocol 28) and does not connect to 0.56.0: a server used with it stays on 0.55.2 until the app is updated to a release that includes 0.56.0. sudo twcore upgrade --version 0.55.2 --restart switches a server back to 0.55.2.
    • The request store’s schema is unchanged (22): the request history is kept.
    • groups[].session_affinity is removed. A configuration that still sets it (only session_affinity: false was ever written) is refused; delete the line.
    • Changed in the protocol:
      • Removed: GroupView.session_affinity, GroupView.hurts_cache, GroupInput.session_affinity, DryRunResult.hurts_cache.
      • RequestRouted.affinity and RoutingView.affinity (AffinityView, Stay): whether the request kept its turn’s route and why it stayed with the upstream that answered before.
      • Overview.failover (FailoverView).
      • AttemptOutcome has estimated.
    • New in the configuration: the failover section (see Setting a failing upstream aside).
    • Message codes: new config.failover_range and gw.upstream.stream_opening_error; removed gw.count_tokens.bedrock_model.

    A conversation stays where its cache is. Routing rules used to be evaluated for every request, and only load-balance groups kept a session on one upstream. A rule on input_tokens or on images could therefore switch upstreams in the middle of an agent’s turn, and a failover moved the conversation away from the upstream holding its prompt cache for good.

    • A conversation is recognised by the client’s session header (x-claude-code-session-id, session_id, conversation_id and a few others) together with the content fingerprint, or by the fingerprint alone. A turn lasts from a user message until the next one; tool results belong to the turn.
    • The routing decision made at the start of a turn (rule, group, rewrites) holds for the rest of the turn. It is made again when the input approaches the smallest context window among the candidate models, where the price data gives one.
    • The next request goes first to the upstream that last answered the conversation, including one that took over after a failover: always within a turn, and across turns while its last answer read or wrote at least 1024 cached tokens less than five minutes ago. An upstream that is set aside, not among the candidates or in another group is not held to.
    • load-balance takes turns between new conversations only. session_affinity and the warnings that a group lowers the cache hit rate are gone, since no strategy drops a warm cache in the middle of a conversation any more.

    Errors at the start of a stream fail over. An upstream can answer 200, open a stream and report an error as its first event: Anthropic’s overloaded_error, the Codex backend’s response.failed when the usage limit is reached, Bedrock’s throttlingException, an error chunk from a Chat or Gemini upstream. The client had received nothing yet, and the error still reached it. A streamed answer is now held until its first content arrives, and an error before that point moves the request to the next candidate, as an error status does. What was read while waiting is passed on unchanged. The wait ends after failover.stream_start_wait_secs (15) or 1 MiB; the last candidate is not held.

    Setting a failing upstream aside. Failures are sorted by the reason the upstream gives.

    • 401, 403, 402 and 404 now move the request to the next candidate. A 400 or 422 does so only when its body names an insufficient balance, a used-up quota or an unavailable model; other 4xx still go to the client unchanged. The last candidate’s own 4xx reaches the client as the upstream wrote it.
    • An insufficient balance sets the upstream aside for no_balance_pause_secs (1800). A used-up quota sets it aside until the reset time the upstream names in the body, in its quota headers or in a GLM 429, and for quota_pause_secs (3600) when it names none. A rate limit with Retry-After (or Gemini’s retryDelay) sets it aside for that long, at most rate_limit_max_pause_secs (3600). A missing model moves on without counting against the upstream.
    • Failures without a stated reason pause the upstream after failures_to_pause (3) in a row, for pause_secs (60), doubling with each further pause up to max_pause_secs (600) and starting over after a success.
    • A request with a single candidate is never affected, and when every candidate is set aside they are all tried.

    Counting tokens. /v1/messages/count_tokens and Gemini’s :countTokens routed to an upstream of another format, or to a same-format upstream that answers 404 or 405, are answered by the gateway with an estimate of the system prompt, messages, tool calls and results, and tool definitions; nothing is sent to that upstream. The answer carries x-thinkwatch-local: 1, and the request is recorded as a local answer with no cost. A count does not fail over to another format or another model. Counting through an AWS Bedrock upstream still answers 501 not_supported, after which Claude Code counts precisely itself.

    Reasoning sealed by another account. Reasoning items carry content encrypted or signed for the account that produced them. After a failover from one account to another, the upstream refuses them (Responses’ invalid_encrypted_content, Anthropic’s invalid signature in a thinking block). The request is now sent once more without them, and the refused items are left out up front on the conversation’s later turns to that upstream.

  6. ThinkWatch Lite v2026.9.24 GitHub ↗

    Upgrade notes:

    • The bundled core is now 0.55.2, which speaks control-plane protocol 28. A remote server has to run core 0.55.2 too: this version does not connect to 0.54.0, and 2026.9.23 and earlier do not connect to 0.55.x.
    • Request history is kept.
    • A request that a routing rule sent as another model is now priced by the model it was sent as, so its cost can differ from what an earlier version showed (see Costs).

    Amazon Bedrock upstreams: Amazon Bedrock is in the list of services.

    • Choose the region and one credential: a Bedrock API key; access keys, with a session token for temporary credentials; or an AWS profile, read from the AWS credential files on the machine the core runs on. Each can be written as ${VARIABLE}. With a remote core, the variables and the profile are read on the server. A profile that signs in or runs a command to get its keys cannot be used, and says so.
    • Check connection verifies the credential and lists the models the region can route to: foundation models, inference profiles such as us.anthropic.claude-… and global.…, and the account’s own application inference profiles.
    • Requests in every client format are converted to Bedrock’s Converse API and signed one by one. Streaming, usage, output redaction and tool-call inspection work as with other upstreams, and prompt-cache breakpoints and Claude’s thinking are carried over.
    • Bedrock takes its own model IDs. A client that asks for Anthropic’s names, such as claude-sonnet-4-5, needs a routing rule that changes the model to a Bedrock ID.
    • For a VPC endpoint or a proxy, enter its address and the region.
    • When AWS refuses a credential, its message names the account and the IAM identity; clients get the gateway’s own sentence instead.
    • Replaying a request to a Bedrock upstream is refused: a replay sends the recorded request unchanged, and Bedrock takes only the Converse format, which no client sends.

    Claude Code and Claude Desktop that used Bedrock directly:

    • Claude Code with CLAUDE_CODE_USE_BEDROCK turned on ignored the gateway address and kept calling Bedrock directly, although it showed as connected. The same held for its other provider switches: Mantle, Google Cloud’s Agent Platform, Microsoft Foundry and Claude Platform on AWS. Connecting Claude Code now turns such a switch off in ~/.claude/settings.json, and restoring turns it back on exactly as it was. A switch turned on by an organization’s managed settings cannot be overridden, so connecting is refused with the reason. The configuration check reports a switch that was turned on again later.
    • The confirmation lists the model variables that name no Bedrock model: those requests reach the gateway under Anthropic’s model names and need a routing rule.
    • The confirmation says whether the gateway has an Amazon Bedrock upstream, and New Bedrock upstream… opens the new upstream filled in from the client’s own settings: the region, a custom endpoint, and the credential as a ${VARIABLE} or a profile name. A key kept in plain text in the client’s configuration is not copied; the notes say where it is. Claude Desktop’s Bedrock configuration in use before connecting is recognized the same way.
    • Credentials in the env block of settings.json, such as the Bedrock API key and access keys that /setup-bedrock writes there, are masked in the change shown before connecting. A value written as an empty string is shown as "".

    Web search through other formats: Claude Code runs each web search as a request that only an upstream in Anthropic’s format can carry out. Sent to a Bedrock, OpenAI or Gemini upstream, the search tool was dropped and the model’s own answer came back as the search results. Such a request now goes to the next upstream of the route, and otherwise fails with a message naming the tool and the upstream.

    Costs:

    • A request that a routing rule sent as another model is priced by the model it was sent as. The traffic view still shows the model the client asked for.
    • Long-context prices apply from the threshold the price data gives (200K for Claude Sonnet 4 and 4.5, 272K for GPT-5.4 and later), and cache reads and writes count toward it.
    • A price the price data does not give, such as a one-hour cache write, is filled in and the cost is marked estimated. A Bedrock model the price data does not list borrows the price of the same model and is marked estimated.

    Errors partway through a stream: an upstream that starts answering and then reports an error inside the stream, such as Anthropic’s overloaded_error or Responses’ response.failed, now leaves the request recorded as failed instead of finished, with the usage reported before the error.

  7. ThinkWatch Core v0.55.2 GitHub ↗

    This release stops a Bedrock API key that is not valid from passing the connection check.

    Upgrade notes

    • The control-plane protocol version is unchanged (28) and so is the request store’s schema (22). ThinkWatch Lite 2026.9.23 includes 0.54.0 (protocol 27) and does not connect to 0.55.x: a server used with it stays on 0.54.0 until the app is updated to a release that includes 0.55.2.
    • No message codes change.

    Checking a Bedrock upstream. AWS answers a Bedrock API key that is not valid with AccessDeniedException (“Authentication failed: …”, “Invalid API Key format: …”), the same exception it uses for a valid credential that may not list models. The check took every AccessDeniedException on the model listing to mean the credential works, so a mistyped key passed with “The credential works, and it has no permission to list models”, and the background model fetch said the same. Only a refusal whose message says the identity is not authorized, which is how IAM denies an action, now counts as missing list permission; any other one reports the key as rejected. AWS’s message is only inspected and never passed on, since it names the account.

  8. ThinkWatch Core v0.55.1 GitHub ↗

    This release stops a request that requires a server-side tool, such as Claude Code’s web search, from being sent without it to an upstream of another format.

    Upgrade notes

    • The control-plane protocol version is unchanged (28) and so is the request store’s schema (22). ThinkWatch Lite 2026.9.23 includes 0.54.0 (protocol 27) and does not connect to 0.55.x: a server used with it stays on 0.54.0 until the app is updated to a release that includes 0.55.1.
    • New message codes: gw.convert.tool_unsendable and gw.convert.tools_unsendable.

    Requests that require a server-side tool. Claude Code runs each web search as a request that carries only the server-side web_search tool and forces the model to use it. Converted for an upstream of another format (AWS Bedrock, OpenAI, Gemini), the server tool was dropped, since only the provider it belongs to runs it: the model answered without searching, and Claude Code passed that answer on as the search results. When the tool a request forces cannot be sent in the upstream’s format, that attempt is no longer made. The next upstream of the route is tried, and if none can take the request the client gets a 400 that names the tool and the upstream. Server-side tools offered alongside other tools, which the model is free to leave unused, are still dropped and listed in the request’s details as before.

  9. ThinkWatch Core v0.55.0 GitHub ↗

    This release adds AWS Bedrock upstreams. Requests to them are converted to Converse and signed per request, and prompt-cache breakpoints and Claude’s thinking now survive that conversion. Each request is priced by the model it was actually sent as, long-context prices follow the thresholds the price data gives, and an error an upstream reports partway through a stream now counts as a failure.

    Upgrade notes

    • The control-plane protocol version (CONTROL_API_VERSION) is now 28. ThinkWatch Lite connects only to a core with the same protocol version, and has to regenerate its protocol types from this release to use it. ThinkWatch Lite 2026.9.23 includes 0.54.0 (protocol 27) and does not connect to 0.55.0: a server used with it stays on 0.54.0 until the app is updated to a release that includes 0.55.0. sudo twcore upgrade --version 0.54.0 --restart switches a server back to 0.54.0.
    • The request store’s schema is unchanged (22): the request history is kept.
    • Changed in the protocol:
      • Protocol has bedrock.
      • ProviderInput.aws and ProviderView.aws (AwsKeys): the access keys, or the name of an AWS profile, and the region to sign for.
      • ProviderView.region and ProviderPreview.region: the region of a Bedrock upstream.
      • AttemptView.model: the model an attempt sent when a routing rule rewrote it.
    • New in the configuration: the protocol bedrock and providers[].aws. The 200K long-context threshold of a price sheet’s input_above_200k and output_above_200k now counts cache reads and writes (see Priced by the model that was sent).
    • New message codes:
      • config.credential.: bedrock_oauth, key_and_aws, bedrock_no_credential, aws_not_bedrock, aws_empty, aws_profile_and_keys, aws_empty_profile, aws_signed_header, aws_no_region, aws_bad_region, aws_region_mismatch
      • config.aws_profile.: no_home, unreadable, not_found, unsupported, no_keys
      • gw.upstream.: aws_token_expired, aws_profile_expired, bedrock_refused, bedrock_refused_unnamed, sign_failed, eventstream_broken, stream_exception, stream_error
      • gw.probe.aws_token_expired, gw.probe.bedrock_list_denied, gw.count_tokens.bedrock_model, gw.count_tokens.bedrock_upstream, control.replay_bedrock, l3.sign_failed

    AWS Bedrock upstreams. An upstream at https://bedrock-runtime.<region>.amazonaws.com is recognized as bedrock. It authenticates in one of three ways:

    • A Bedrock API key in key, sent as Authorization: Bearer.
    • AWS access keys in aws: access_key_id, secret_access_key and, for temporary credentials, session_token, each of which can be ${VAR}. Every request is signed with them (SigV4) once its body is final.
    • aws.profile: the access keys of a profile in the AWS credential files (~/.aws/credentials, ~/.aws/config, or the files AWS_SHARED_CREDENTIALS_FILE and AWS_CONFIG_FILE name) on the machine core runs on. The files are read again when they change, so a tool that refreshes temporary keys there needs no restart. Nothing is run to obtain a credential: a profile that signs in through IAM Identity Center, runs a credential_process or assumes a role is refused, naming the setting.

    Requests in every client format are converted to Converse and ConverseStream. The binary eventstream is turned into server-sent events as it arrives, so usage, the first token, output redaction and tool-call inspection work as they do for any other upstream. The anthropic-beta values Bedrock accepts for Claude are carried into the request body. The model list comes from the region’s control plane: the foundation models that can be invoked on demand, the inference profiles AWS defines (us.anthropic.claude-…, global.…), and the account’s application inference profiles, listed by the ARN they are invoked by. Checking the connection and the inference speed test go through the same signing. For a VPC endpoint or a proxy, write its address in base_url and the region in aws.region.

    • AWS’s text for a refused credential (HTTP 401 or 403) names the account and the IAM identity. The client receives the gateway’s own sentence instead, and AWS’s text is not recorded as the request’s error.
    • An expired session token is named as such; for a profile, the message says to refresh the profile.
    • An exception AWS sends inside a stream ends the request as a failure.
    • count_tokens for a model that only Bedrock upstreams serve is answered with 501 not_supported, the answer Claude Code’s guidance for gateways names; Claude Code then counts with a one-token request. Bedrock’s own CountTokens covers only some older Claude models on bedrock-runtime.
    • A replay to a Bedrock upstream is refused: a replay sends the recorded request as it was, and Bedrock takes only the Converse format, which no client sends.

    The code that speaks to Bedrock (signing, the eventstream, the endpoints and the model catalog) is the new layer-one crate tw-bedrock, which ThinkWatch Enterprise shares.

    Converse keeps cache breakpoints and thinking. The conversion to Converse used to drop cache_control, so a client that cached its prompt never hit the cache through Bedrock. Cache breakpoints now become cachePoint blocks (at most four, as Bedrock allows), with the one-hour TTL where the client asked for it, and cache reads and writes, one-hour writes included, are read back. Claude’s thinking and effort are sent in additionalModelRequestFields, and its reasoning comes back as reasoning content.

    Priced by the model that was sent. A routing rule can send a request as another model, and the upstream charges for the model it received. Each attempt now records the model it sent when a rule rewrote it (AttemptView.model), and a request is priced by the model of the attempt that served it. The row still names the model the client asked for. The list of unpriced models names the model a price has to be set for.

    • A Bedrock model the price data does not list borrows a price, always marked estimated: an inference profile from the model it is named after, and an Anthropic model on Bedrock from Anthropic’s own name for it.
    • Long-context tiers are read with the threshold the price data gives (200K for Claude Sonnet 4 and 4.5, 272K for GPT-5.4 and later), including their cache prices. A request’s input counts its cache reads and writes when the tier is chosen. The 272K tier was not recognized before, and a request with a warm cache was priced as a short one.
    • A price the data does not give is filled in, and the cost is marked estimated whenever it is used: a one-hour cache write costs twice the input price (it used to fall back to the five-minute price), and the missing cache prices of a long-context tier grow with its input price.

    Errors partway through a stream. An upstream can answer 200, stream part of an answer and then report an error in the stream: Anthropic’s overloaded_error, Responses’ response.failed, an error in a Chat or Gemini frame. Such a request was recorded as finished; it is now a failure, with the usage the upstream reported before the error.

    Listed models. GET /v1/models and GET /v1/models/:model no longer stamp each model with the time of the request. created and created_at are now the Unix epoch, since the gateway does not know when a model was released.

  10. ThinkWatch Lite v2026.9.23 GitHub ↗

    Upgrade notes: the bundled core stays at 0.54.0, so a remote server running 0.54.0 keeps working with this version and nothing on the server has to change. Request history is kept.

    Updates keep the window as it was: after an update, the app comes back the way it was before. If it was running in the menu bar only, it stays in the menu bar; if its window was open, the window opens again. Before, the main window opened after every update. This applies to updates the app installs itself, on every platform, and to brew upgrade on macOS. A Homebrew upgrade now also shows the “Updated to …” notification when Notices is set to System notification.

    The update from 2026.9.22 to this version is not covered yet: 2026.9.22 does not record the window when it quits, so this one update still opens the window and shows no “Updated to 2026.9.23” notification.

    Sidebar: with the sidebar expanded, hovering a page no longer pops up a tooltip that repeats its name. The page’s shortcut (⌘1 to ⌘9, or Ctrl+1 to Ctrl+9 on Windows and Linux) appears at the end of its row while the pointer is over the row or the row has keyboard focus. The collapsed sidebar still shows the name and the shortcut in a tooltip.

  11. ThinkWatch Lite v2026.9.22 GitHub ↗

    Upgrade notes:

    • The bundled core is now 0.54.0, which speaks control-plane protocol 27. A remote server has to run core 0.54.0 too: this version does not connect to an older core, and 2026.9.21 and earlier do not connect to 0.54.0.
    • Request history is emptied. The request store changed format, so the requests recorded by an earlier version are removed when the new core first starts. Settings, keys and upstreams are kept.

    Time to first token: latency now counts to the first token of the answer instead of to the response headers. The headers of a streamed response usually come back as soon as the upstream receives the request; the first token comes only after the upstream has queued the request and read the prompt. Traffic, the overview and the upstreams page show this time, and the response-header time moves into the request details.

    Generation speed: a new figure, in tokens per second, is measured from the first token to the end of the answer.

    • Traffic: hovering a request’s latency shows its first token, total time, generation time and speed.
    • Request details: the speed is one of the five figures at the top, and the bar is split into the wait for the first token and the generation.
    • Overview: a new Generation speed section gives the median speed by model and by upstream.
    • Upstreams: the timing column shows the median first token, with the median speed below it.

    Some requests have no speed because it cannot be measured honestly:

    • responses that are not streamed;
    • cancelled and failed requests;
    • generations shorter than half a second;
    • answers whose model reasoned without showing it, when the upstream does not report how many tokens the reasoning took.

    Menu bar:

    • The generation speed in the menu was wrong: a single response that was not streamed could push it to tens of thousands of tokens per second. It now uses the same measurement.
    • After the computer slept through midnight, Today kept showing the previous day’s tokens and cost until a request came in or the menu was opened. It now catches up as soon as the computer wakes, and when the clock or the time zone is changed. The traffic page, which shows only the time for today’s requests, and the keys page’s usage over the last 24 hours catch up the same way.

    Routing: a model name that is not among the suggestions is kept, in a rule’s model condition, in Change model to and in the dry run. Before, closing the suggestions cleared a name the list did not have, such as a relay’s own model name or a model released after the upstream’s list was fetched, and turned a shorter name back into the suggestion picked earlier.

  12. ThinkWatch Core v0.54.0 GitHub ↗

    This release records when each response’s first token arrives and how fast the response is generated, and corrects the generation rate that GET /live reports.

    Upgrade notes

    • The control-plane protocol version (CONTROL_API_VERSION) is now 27. ThinkWatch Lite connects only to a core with the same protocol version, and has to regenerate its protocol types from this release to use it. ThinkWatch Lite 2026.9.21 includes 0.53.0 (protocol 26) and does not connect to 0.54.0: a server used with it stays on 0.53.0 until the app is updated to a release that includes 0.54.0. sudo twcore upgrade --version 0.53.0 --restart switches a server back to 0.53.0.
    • The format of the request history changed (request store schema 22, from 21). On its first start, 0.54.0 replaces the request history and the stored request and response bodies of an earlier version with an empty store; the configuration is kept. Switching back to an earlier version empties it again.
    • Changed in the protocol: the new event RequestFirstToken { id, ttft_ms }, RequestFinished.tokens_per_sec, HistoryRow.ttft_ms and HistoryRow.tokens_per_sec, and the new endpoints GET /token-rate and GET /token-rate/provider, which return TokenRateView { model, p50, samples }. GET /latency and GET /latency/provider now report the time to the first token. No message code changed.

    First token. The headers of a streamed response usually come back as soon as the upstream receives the request. The upstream then queues the request and reads the prompt before it produces anything, and the time a user waits is the time to the first token, not to the headers. The gateway now reads each successful streamed response in the upstream’s format and reports the moment the first text, reasoning or tool call arrives; opening frames such as message_start and response.created do not count. A response that is not streamed arrives whole and has no first token. GET /latency and GET /latency/provider now give percentiles of this time, so responses that are not streamed are left out of them. url-test groups still pick the upstream with the fastest response headers.

    Generation speed. Each finished streamed request has a generation speed: the output produced after the first token, divided by the time from the first token to the end.

    • When the model reasoned before the first token without streaming its reasoning, those reasoning tokens were generated outside that time. They are subtracted when the upstream reports how many there were, as OpenAI and Gemini do. When it does not, as with Anthropic, the request has no speed rather than one several times too high.
    • Reasoning that is streamed as it happens is counted.
    • Responses that are not streamed, cancelled and failed requests, and generations shorter than half a second have no speed.

    GET /token-rate and GET /token-rate/provider give the median speed by model and by upstream, with the number of samples.

    The live rate. GET /live used to divide each request’s output by the time after its response headers. A response that is not streamed gets its headers only when generation is done, so that time was a few milliseconds, and one such request could push the rate to tens of thousands of tokens per second. The rate now combines the speeds of the requests that have one, weighted by their generation time.

  13. ThinkWatch Lite v2026.9.21 GitHub ↗

    Upgrade notes: the bundled core is now 0.53.0, which speaks control-plane protocol 26. A remote server has to run core 0.53.0 too: this version does not connect to an older core, and 2026.9.20 and earlier do not connect to 0.53.0. Request history is kept. Upstreams now see ThinkWatch’s User-Agent instead of the client’s, so an upstream that admits only certain clients, such as Kimi For Coding, Bailian Coding Plan or a relay restricted to official clients, needs the new Forward client identity switch in its connection settings.

    What upstreams receive:

    • Each upstream receives the request and only the headers its API needs. Before, almost every header a client sent went along: Claude Code’s User-Agent, x-app, SDK details and session ID, Gemini CLI’s installation ID, and a browser client’s cookies and origin. Upstreams now see ThinkWatch’s User-Agent and the credential and headers configured for the upstream. From the client’s request they get only the headers of the upstream’s own protocol: anthropic-* for Anthropic, Idempotency-Key and X-Client-Request-Id for OpenAI, and none for Gemini.
    • Identity fields that clients fill in on their own are removed from the request body, such as Claude Code’s metadata.user_id, which holds a device ID, an account ID and a session ID.
    • Forward client identity, a new switch in the connection section of the upstream dialog, sends that upstream the client’s own User-Agent and identity, unchanged. It is off by default, including for upstreams created by the Z.ai sign-in, and ChatGPT accounts do not offer it.

    ChatGPT accounts:

    • Newer models that the account can use, such as GPT-6-Luna, appear in the upstream’s model list and can be requested. Before, the gateway answered “No upstream serves model gpt-6-luna” even though the account had the model.
    • Requests from the Codex in the ChatGPT app no longer pass that app’s attestation and installation ID on to OpenAI. On one account, OpenAI revoked the sign-in within seconds of the first such request.
    • Request bodies are compressed with zstd, as Codex does, so long conversations upload faster.
  14. ThinkWatch Core v0.53.0 GitHub ↗

    This release changes what ThinkWatch sends to upstreams. Each upstream now receives the request and only the headers it needs; the client’s own identity stays with ThinkWatch unless an upstream asks for it. It also fixes two problems with ChatGPT account upstreams.

    Upgrade notes

    • The control-plane protocol version is now 26 (it was 25). The request store’s schema is unchanged (21). ThinkWatch Lite 2026.9.20 includes 0.52.1 (protocol 25) and does not connect to 0.53.0, so a server used with it stays on 0.52.x until the app is updated to a release that includes 0.53.0.
    • Upstreams now see ThinkWatch’s User-Agent instead of the client’s. An upstream that admits only certain clients, such as Kimi For Coding, Bailian Coding Plan or a relay restricted to official clients, needs the new provider field forward_client_identity: true.
    • New message code: config.credential.chatgpt_client_identity.

    Each upstream gets only what it needs. Client request headers used to reach the upstream almost unchanged. Only credentials, a few transport headers and, when a request was translated, the client format’s own protocol headers were removed. Claude Code’s User-Agent, x-app, x-stainless-* and session ID, Gemini CLI’s installation ID and a browser client’s Cookie and Origin all went through. Outbound headers now come from three sources only:

    • Set by ThinkWatch: its own User-Agent, and content-type when the body was translated.
    • The upstream’s configuration: the credential and the headers written there. OpenAI-Organization and OpenAI-Project belong here, because they belong to the upstream’s key.
    • The client’s request, only for the headers the upstream’s protocol uses:
      • Anthropic: anthropic-*, except anthropic-dangerous-direct-browser-access.
      • OpenAI: Idempotency-Key and X-Client-Request-Id.
      • Gemini: none.
      • Every protocol: content-type and accept when the body passes through unchanged.

    An Anthropic request without anthropic-version now gets the default one. When the body passes through unchanged, identity fields that clients fill in themselves are removed from it: Claude Code’s metadata.user_id, and the Codex installation ID and turn metadata in client_metadata.

    forward_client_identity. With this field on, an upstream also receives the client’s own User-Agent, x-app, originator, version, session_id and x-codex-* headers, and the body’s identity fields are kept. These are the client’s own values; nothing is made up. The field is off by default. A ChatGPT account upstream refuses it: that upstream always names ThinkWatch as the sender.

    ChatGPT account upstreams.

    • New models were missing from the list. The Codex backend leaves a model out of the list when its minimum client version is above the client_version that the caller reports. ThinkWatch reported 0.0.0, so a model such as gpt-6-luna, which needs Codex 0.155.0, was not listed, and requests for it were refused with “No upstream serves model …” even though the account can use it. ThinkWatch now reports 99999.0.0, so the list shows every model the account can use. Requests still name ThinkWatch as their origin.
    • The ChatGPT app’s identity is no longer forwarded. The Codex inside the ChatGPT desktop app adds an attestation generated by the app (x-oai-attestation), its installation ID and turn metadata. ThinkWatch forwarded them together with the token it had obtained itself, and on one account OpenAI revoked that token within seconds of the first such request.
      • From the client, the Codex backend now receives only x-codex-turn-state, x-openai-internal-codex-responses-lite and x-codex-beta-features.
      • The attestation, installation ID and turn metadata cannot be set in an upstream’s configured headers.
    • Request bodies are compressed. Requests to the Codex backend are compressed with zstd, which is what Codex does when it talks to that backend itself.
  15. ThinkWatch Lite v2026.9.20 GitHub ↗

    Upgrade notes: the bundled core stays at 0.52.1, so a remote server running 0.52.1 keeps working with this version and nothing on the server has to change.

    Connecting and restoring clients:

    • A change is written only as it was shown. If a client rewrites its configuration while the confirmation dialog is open, nothing is written and the dialog shows the change worked out from the file as it is now, to confirm again. This applies to connecting and restoring a client, copying or removing an MCP server, and switching WSL to mirrored networking.
    • Restoring no longer deletes sections that connecting had created once something else was added to them: Claude Code’s env in settings.json, the provider section of a new opencode configuration, and DeepSeek Harness’s refs. Restoring DeepSeek Harness no longer fails with “refs is not a scalar”, and its credentials file keeps its version when DeepSeek Harness has written its own entries into it.
    • A configuration file that contains an empty object written as { } no longer gets a stray comma when a setting is added to it. Claude Code and Claude Desktop rejected such a file.
    • In TOML files, a comment at the end of a changed line is kept, and an MCP server copied from a JSON client no longer gets empty strings where it had null.
    • When the record of an earlier connection next to a client’s configuration is damaged or belongs to another client, connecting is refused, as restoring already was. Before, it recorded the gateway’s address and key as the original values, and a later restore put them back.
    • Keys no longer appear in the change shown before writing: restoring showed the user’s own key, and copying an MCP server showed the environment variables and headers of every server in the file. Messages about a configuration file that cannot be parsed no longer quote the line that failed, and a failed write no longer leaves a temporary copy with the gateway key next to the configuration.
    • Backups are readable by their owner only, and a configuration folder that is a symbolic link (for example into a dotfiles repository) is pointed out before writing.
    • Pointing Claude Desktop at a new address or key, after a key change or when switching to a remote core, updates its model list too.
    • On Windows, a running client is recognized by its program name, so a program such as TombRaider.exe is no longer taken for Aider.
    • Uninstalling keeps the data folder when a client could not be restored, since it holds that client’s only full backup. It stops the core before deleting the data, and its confirmation also lists the clients inside WSL that will be restored.
    • The “Change path…” dialog no longer freezes the window while a network path that is offline takes its time to answer.
    • On a new installation, the data folder is created readable by its owner only, as intended. It holds config.yaml with the upstream keys, and on Windows its access rules are their only protection.

    Figures:

    • When a statistic cannot be read, the Overview says it is unavailable and the Upstreams page shows ”—” with a Retry button, instead of zeros or an empty chart.
    • Today’s cost in the menu bar, and the costs on the Keys and Upstreams pages, keep measured, estimated and unpriced apart, as the Overview does. A cost is no longer shown as $0 when the requests could not be priced.
    • Very small amounts are shown as ”<$0.0001” instead of “$0.0000”. Numbers change unit after rounding, so they read “10k” and “1.0M” rather than “10.0k” and “1000k”.
    • In live mode, requests whose price has not arrived yet make the total a lower bound, marked ”≥”, instead of counting as $0.
    • In a session, a turn is marked Unpriced only when a price would fix it. “Show unpriced only” on the Traffic page no longer lists failed requests and requests without usage.
    • Requests in progress when a remote connection dropped are corrected from the server’s history instead of staying marked as failed.

    Time ranges: “Today” in the menu bar starts at midnight on this computer, also when the core runs on a server in another time zone and on the days clocks change. A custom range such as “since 9/8” keeps its start day while the page stays open. Long custom ranges are drawn with daily or weekly points and always reach the present; ranges over about four months used to stop short of it.

    Core and remote connections:

    • Switching to a remote core also stops a local core that was still starting or restarting. Before, it kept being started and held the gateway port.
    • A core in safe mode is shown as being in safe mode, not as running, and Restart works from there.
    • A remote connection that died without a sign, for example after the computer slept or the network changed, is noticed within 45 seconds, and a server that accepts the connection but never answers makes the connection test time out.
    • Quitting the app stops the local core instead of leaving it to notice a few seconds later.

    Notifications:

    • A tool call blocked by the security rules, and something suspicious appearing in a client’s configuration, are reported again each time they happen. Before, only the first one was ever reported, even after a restart. To keep a burst from flooding the screen, each shows at most once every five minutes (per upstream for blocked calls), and one notification then sums up the rest.
    • A notification that waits a minute before showing is dropped if the gateway went down meanwhile. The “N more” count no longer counts the same notice twice. The title of an upstream that keeps dropping out no longer grows longer with each request.
    • “Updated to …” follows the notification setting.
    • The notice list and the settings are saved in a way that survives the app being killed mid-save, and a setting this version does not recognize, such as one written by a newer version, resets only that setting.

    Menu bar and tray: the color reflects the most severe quota window, not only the fullest one. On Windows and Linux, menu actions such as Check for Updates report their result in a notification, and messages from the core appear in the interface language. A click on a menu that has just been replaced no longer runs a different item. The menu bar no longer stops updating when the core stops answering.

    MCP and the security scan:

    • The scan also covers MCP servers configured per project in ~/.claude.json, Claude Code settings that run commands without the model (statusLine, apiKeyHelper and similar), and Codex’s notify. Helpers that print credentials are scanned but not listed, so their commands do not appear on screen.
    • Skills and folders created after the app started are watched as well, and so is opencode.jsonc.
    • A remote MCP server is no longer taken for a local one because its address contains “localhost”. SKILL.md files with Windows line endings or a byte order mark are read correctly.
    • A file full of hidden characters no longer stalls the scan; findings are capped per file, with a note saying where the scan stopped. A special file, such as a link to /dev/zero, no longer blocks it.
    • Findings in a hook’s command, including hidden characters, are shown on that hook’s row.

    Editing:

    • A config.yaml with Windows line endings no longer shows unsaved changes on opening, and is saved with its line endings as they were. “Show in config file” selects the right entry when a key and a route share a name.
    • Edit dialogs save against the configuration they were opened with. A change made elsewhere meanwhile is reported as a conflict instead of being overwritten. Switching several items on or off in quick succession no longer fails with “version mismatch”.
    • Price fields and the dry run accept a decimal comma.
    • In the upstream dialog, the connection check result is cleared when the connection settings change.
    • In a key’s model scope, model IDs typed by hand are shown and can be removed even when the model list is empty, and they are compared without regard to case, as the gateway does.
    • Saving a route applies its auxiliary request categories as they are after the edit: unticked or deleted conditions no longer keep them assigned to routing. max_tokens of 0 is refused instead of being dropped on save, and a copy of the fallback rule can be moved ahead of it.
    • A replayed request goes to the upstream it was priced for.
    • Several dialogs no longer act on stale state: the speed test waits for the new quote before it can start, deleting a price sheet from the upstream dialog also clears it from the form, checking for updates from About no longer overwrites settings changed meanwhile, dismissing one key-rotation banner no longer dismisses the others, and Z.ai sign-in accepts an existing upstream whose address ends in a slash.
    • A small payload cap in Retention no longer blocks saving that section.
    • Ctrl/⌘+B no longer toggles the sidebar while typing in a text field.
    • On Windows, choosing a connection at startup and opening the update window no longer risk freezing the app.

    Speed: the update and connection windows load a fraction of the code they used to, and the configuration editor loads when it is opened. The change shown before writing a large ~/.claude.json appears at once. The Traffic page does less work per event, and the macOS menu bar updates its menu in place while it is open.

  16. ThinkWatch Lite v2026.9.19 GitHub ↗

    Upgrade notes: the bundled core is upgraded to 0.52.1. The request store keeps its format, so the request history is kept. A remote server has to run the same core version, 0.52.1 (ThinkWatch Lite 2026.9.18 cannot connect to a server running 0.52.x), so the server and the app are updated together. When the versions differ, the app gives the command to run on the server, sudo twcore upgrade --version 0.52.1 --restart. A server upgraded from 0.51.0 keeps its request history as well.

    Upstreams and proxies: the edit dialogs show the settings as they are written in the configuration.

    • The base URL keeps its path. It used to be cut down to its origin, under a note that the credentials in the URL are hidden.
    • The API key is filled in and hidden, with an eye button to show it, and the authorization header under Headers follows the same button. A ${NAME} reference stays visible. The other header values are filled in as written.
    • OAuth upstreams have their token endpoint, refresh token, client ID and client secret filled in, the secrets hidden. The access token field appears only once the credential is new or changed.
    • A proxy’s username and password are filled in, the password hidden.
    • A save sends the whole definition, so the “leave empty to keep”, “Remove” and “Replace” states are gone.
    • A base URL that ends with its version segment, as provider docs give it (https://api.openai.com/v1), no longer has every request and connection check sent to …/v1/v1/…, and the hint under the base URL now only asks to leave out an endpoint path such as /chat/completions.
    • A proxy password can be a ${NAME} reference. A password containing ${ used to be refused.
    • Z.ai sign-in: entering the name of the existing Z.ai upstream was refused as a name already in use. The dialog now says that the sign-in replaces that upstream’s key, and the sign-in can go ahead.

    Routing:

    • A rule that rewrites the model was refused with “No upstream serves model …” whenever the target upstream lists its models. That is the rule the app suggests for Claude Desktop with a non-Claude upstream, and for Antigravity CLI and DeepSeek Harness with another provider’s models. Admission, skipping the candidates that cannot serve a request, the cheapest group’s price comparison and the routing dry run now look at the model each upstream is actually asked for. When a rewritten model is not served, the error names both models.
    • Requests from Codex and other clients that use the OpenAI Responses or Gemini format are recognized as conversations. A load-balance group with sticky sessions, the default, sent all of them to its first upstream. Each conversation now stays on one upstream while the conversations are spread across the members, and the Traffic page groups them into sessions.

    Clients: a client’s row menu on the Clients page (right-click or the … button) has “Change path…”, for a client whose files are somewhere else, such as Claude Code moved with CLAUDE_CONFIG_DIR or Codex moved with CODEX_HOME. It covers the clients under “Set up by hand” too, such as Cursor and Antigravity CLI. The dialog shows where connecting, MCP management and the security scan read and write. They move together: before anything changes, the dialog lists every location that moves, old and new. “Restore defaults” puts them back. The same dialog opens from the MCP page, by right-clicking a column header or a row under “Skills and hooks” or “Findings”. A connected client has to be restored first. The locations apply to this computer only, and clients inside WSL keep their default locations. Claude Desktop and DeepSeek Harness cannot be moved.

    Remote connections: editing a remote connection fills in the saved key, hidden behind an eye button, so it can be checked or copied.

    Menu bar and tray: the menu now reads, from the top: the gateway’s state, Notices, Today, Quota and In Progress, then Open ThinkWatch Lite, the groups whose upstream can be chosen, Copy Gateway Address and Copy Default Key, then Settings…, Connection, Check for Updates… and Quit ThinkWatch Lite…. On Linux, Open ThinkWatch Lite stays at the very top of the tray menu. “Undo Last Configuration Change” is gone; a configuration change can still be restored from the version history in the app. A quota window named by its length, such as the 30-day window a ChatGPT account can report (shown as 30d), is spelled out in the menu, on the Upstreams page and in notifications: “30 days”, or “30-day window” in a sentence.

  17. ThinkWatch Core v0.52.1 GitHub ↗

    This release fixes two routing problems: a routing rule that rewrites the model was judged by the model the client asked for, and requests in the OpenAI Responses and Gemini formats were never recognized as part of a conversation.

    Upgrade notes

    • The control-plane protocol version is unchanged (25) and so is the request store’s schema (21). ThinkWatch Lite 2026.9.18 includes 0.51.0 (protocol 24) and does not connect to 0.52.x: a server used with it stays on 0.51.0 until the app is updated to a release that includes 0.52.1.
    • New message codes: gw.model.no_upstream_rewritten and gw.model.not_allowed_rewritten.

    Rewriting the model. A rule that sets the model, for example to send the Claude-style names that Claude Desktop requires to a GLM or DeepSeek upstream, was refused with “No upstream serves model …” whenever the target upstream lists its models. Admission, skipping the candidates that cannot serve a request, and the cheapest group’s price comparison all looked at the model the client wrote, before any rule rewrote it. They now look at the model each upstream will actually be asked for, including a rewrite that applies to one upstream only (provider_would_be), so admission runs after the rules have decided. When a rewritten model is not served, the error names both models. The routing dry run judges the candidates the same way.

    Conversations in the Responses and Gemini formats. Requests from Codex and other clients that use the OpenAI Responses or Gemini format were never assigned to a session, because the fingerprint read only the Anthropic and Chat Completions fields. A load-balancing group with session affinity, which is the default, sent all of them to its first upstream, and the traffic view did not group them into sessions. The fingerprint now also reads instructions and input, and systemInstruction and contents, and uses the request’s prompt_cache_key when it has one. Codex sends one per conversation, and Codex conversations started in the same repository open with identical text.

  18. ThinkWatch Core v0.52.0 GitHub ↗

    This release hands ThinkWatch Lite the upstream and proxy settings as they are written in the configuration, so its edit dialogs can show them, and stops repeating the version segment when an upstream’s base URL ends with it.

    Upgrade notes

    • The control-plane protocol version (CONTROL_API_VERSION) is now 25. ThinkWatch Lite connects only to a core with the same protocol version, and has to regenerate its protocol types from this release to use it. ThinkWatch Lite 2026.9.18 includes 0.51.0 (protocol 24) and does not connect to 0.52.0: a server used with it stays on 0.51.0 until the app is updated to a release that includes 0.52.0. sudo twcore upgrade --version 0.51.0 --restart switches a server back to 0.51.0.
    • The request store’s schema is unchanged (21). Upgrading from 0.50.x or 0.51.0 keeps the request history.
    • Changed in the protocol: ProviderView.base_url, ProviderView.key (now a string), HeaderView.value, the new OAuthView.refresh and OAuthView.client_secret, and ProxyView.auth (which replaces has_auth) carry the settings as written; base_url_masked, SecretView and HeaderView.masked are gone. ProviderInput.base_url is required, ProviderInput.key is an optional string and every HeaderInput has a value; ProxyInput.auth is optional. SecretChange and ProxyAuthInput are gone, and so is the message code control.pass_has_expansion.

    Settings as written. The upstream and proxy views used to mask what they showed. The base URL was cut down to its origin, so https://bedrock-mantle.us-east-1.api.aws/v1 was shown without /v1, with a note that credentials had been hidden when there were none. API keys and header values were masked, and OAuth refresh tokens, client secrets and proxy credentials were left out. The app could not fill its edit dialogs in, and every save carried “keep what is saved” states instead of values. The views now carry the settings as they are written in config.yaml, and a save carries the whole definition. The access token of an OAuth credential stays out of the view, and an OAuth credential can still be kept as it is on save: while a dialog is open the gateway may renew the token and write back a new refresh token. GET /config already returned the file unmasked. Diagnostics bundles and logs still mask.

    Proxy passwords from the environment. A proxy password may be a ${NAME} reference, as the configuration reference shows (pass: ${PROXY_PASSWORD}). A password containing ${ used to be refused.

    Base URLs that end with a version segment. A base URL written up to its version segment, as provider docs give it (https://api.openai.com/v1) and as the configuration reference says, got every request and every connection check sent to .../v1/v1/...: the client’s path (/v1/chat/completions) was appended as it was. When the base URL’s path ends with the version segment the request path starts with (v1, v2, v1beta, …), that segment is now used once. Other segments are joined as before.

  19. ThinkWatch Lite v2026.9.18 GitHub ↗

    Upgrade notes: the bundled core is upgraded to 0.51.0. The request store has a new format: on its first start, the new core empties the request history, the stored request and response bodies and the security log, and statistics such as usage and cost start again from zero. The configuration is kept. A remote server has to run the same core version, 0.51.0 (ThinkWatch Lite 2026.9.17 cannot connect to a server running 0.51.0), so the server and the app are updated together. When the versions differ, the app gives the command to run on the server, sudo twcore upgrade --version 0.51.0 --restart. When the core on a server is upgraded from 0.49.x, its first start empties the server’s request history as well.

    Clients:

    • Claude Desktop can be pointed at the gateway in one step, through its third-party inference mode. Afterwards Claude Desktop has to be quit completely and reopened; if the sign-in page appears, choose to continue with the gateway. Conversations in this mode are kept apart from the others. A Claude Desktop managed by an organization cannot be pointed at the gateway.
    • DeepSeek Harness can be pointed at the gateway in one step, and its web and desktop apps then send their requests through the gateway. Its web search still goes to DeepSeek directly. The MCP page lists its MCP servers, and the scan covers its patches and skills.
    • Antigravity CLI takes the place of Gemini CLI on the Clients page, with steps to set it up by hand and a key of its own. Gemini CLI is no longer listed; a Gemini CLI that is already set up keeps working. The MCP page lists Antigravity CLI’s MCP servers, and the scan covers its hooks, MCP configuration, skills and subagents.
    • opencode: once opencode was pointed at the gateway, no ThinkWatch model could be selected. The provider it gets now carries the SDK package and the models the client’s key can use, for opencode 1.x and 2.x. When upstreams or routes change, the Clients page says that opencode’s model list needs updating. The MCP page reads the MCP servers of opencode 2.x (mcp.servers), and the scan covers opencode.jsonc.
    • Codex: after Codex is restored, the sessions started while it pointed at the gateway can still be opened, and go to OpenAI directly. The notes shown before pointing Codex at the gateway say that sessions started before and after are listed separately.
    • WSL on Windows: Claude Code and Codex inside WSL can be pointed at the gateway, under WSL 1 or under WSL 2 with mirrored networking. Both reach the gateway at 127.0.0.1, and the gateway keeps listening on this computer only. When WSL 2 uses NAT networking, the Clients page can switch it to mirrored networking and restart WSL; mirrored networking needs Windows 11 22H2 or later. When the app switches to a remote core, the clients inside WSL can be pointed at the server as well.

    Upstreams: GLM Coding Plan upstreams (Z.ai and BigModel) show their 5-hour and weekly quota and the reset times on the Upstreams page, in the menu bar and in the tray, and credit-based plans show the credits left. A used-up quota sends a notification. The monthly count of calls to Z.ai’s own MCP tools in older plans is not shown.

    Traffic: the request details show when a request from DeepSeek Harness carried a session log, and its size. Before such a request reaches an upstream other than DeepSeek, the gateway removes the session log and the other DeepSeek Harness extensions.

    Gateway: it answers the startup check of Claude Desktop’s third-party mode and keeps a stream that stays silent for a long time alive. /v1/files is no longer forwarded to an upstream. A gateway set to listen on a network interface that is not there yet starts on this computer only and adds the interface once it appears.

  20. ThinkWatch Core v0.51.0 GitHub ↗

    This release stops counting the MCP tool calls of the older GLM Coding Plans as a quota, marks only the weekly window when a model request gets business code 1310, and tells the app at once when a GLM key turns out to have no plan.

    Upgrade notes

    • The control-plane protocol version (CONTROL_API_VERSION) is now 24. ThinkWatch Lite connects only to a core with the same protocol version, and has to regenerate its protocol types from this release to use it. ThinkWatch Lite 2026.9.17 includes 0.49.0 (protocol 22) and no released version of the app includes 0.50.0, so a server used with the app stays on 0.49.x until the app is updated to a release that includes 0.51.0. sudo twcore upgrade --version 0.49.1 --restart switches a server back to 0.49.1.
    • The request store’s schema is unchanged (21). Upgrading from 0.50.0 keeps the request history. Upgrading from 0.49.x or earlier replaces the request history and the stored bodies with an empty store on the first start, as 0.50.0 does; the configuration is kept. Switching back to 0.49.x empties it again.
    • Changed in the protocol: the quota window monthly is gone, and a QuotaSeen with empty windows means the upstream’s quota was taken down. No endpoint, type or message code changed.

    MCP calls are not a quota. The quota endpoint of the older GLM Coding Plans also reports a count (TIME_LIMIT) of calls to Z.ai’s own MCP tools: search-prime, web-reader and zread. Those calls do not go through the gateway, so the count says nothing about the requests it forwards. Core no longer reads it: there is no monthly window, and a used-up count no longer marks the upstream used up or emits QuotaExhausted. credits are reported only for the windows of credit-based plans (CREDIT_LIMIT).

    1310 is the weekly quota. A 429 with business code 1310 on a model request marks the weekly window used up, whatever its message says. 1308 still marks the 5-hour window. For 1316 to 1321 the window still comes from the message when it names the 5-hour or weekly window; otherwise it is the fuller of the two known windows, or the weekly one when neither is known.

    A key without a plan. When a GLM key is found to have no plan, core clears the quota it had stored for that upstream and emits a QuotaSeen with empty windows, so the app removes the quota at once instead of on its next read of /quota.

  21. ThinkWatch Core v0.50.0 GitHub ↗

    This release lets Claude Desktop’s third-party mode use the gateway, cleans the extensions DeepSeek Harness adds to its requests before they reach an upstream other than DeepSeek, reads the GLM Coding Plan quota of Z.ai and BigModel upstreams, and recognises requests from Antigravity CLI. A gateway bound to a network interface now starts even when that interface is not there yet.

    Upgrade notes

    • The format of the request history changed (request store schema 21, from 20). On its first start, 0.50.0 replaces the request history and the stored request and response bodies of an earlier version with an empty store; the configuration is kept. Switching back to an earlier version empties it again.
    • The control-plane protocol version (CONTROL_API_VERSION) is now 23. ThinkWatch Lite connects only to a core with the same protocol version, and has to regenerate its protocol types from this release to use it. ThinkWatch Lite 2026.9.17 includes 0.49.0 (protocol 22) and does not connect to 0.50.0: a server used with it stays on 0.49.x until the app is updated to a release that includes 0.50.0. sudo twcore upgrade --version 0.49.1 --restart switches a server back to 0.49.1.
    • New in the protocol: QuotaWindow.credits (a new type, QuotaCredits { total, used, remaining }), the quota window monthly, and session_log_bytes on RequestStarted and HistoryRow. New message code: gw.files.unsupported.
    • /v1/files and /files now answer 404 on the gateway instead of being forwarded to an upstream.

    Claude Desktop. Claude Desktop in third-party inference mode, and the Claude Code engine it embeds, expect a few things of a gateway:

    • HEAD /api/hello, the probe Desktop sends at startup, is answered with 200 by the gateway itself, without a key and without reaching an upstream. It used to be refused with 401.
    • The client gives up on a stream that is silent for about five minutes. When a client speaking Anthropic Messages has received nothing for 15 seconds, the gateway writes event: ping, so a converted upstream that is reasoning or queued at length no longer looks like a dead connection. The timer runs from the last bytes the client received, not from the upstream’s, because conversion drops the upstream’s own keep-alive comments. Pings go only between frames, are not recorded in the stored body or counted as output, and are never sent to clients of other formats. After ten minutes without a byte from the upstream the gateway stops pinging, so a half-open upstream connection ends on the client’s timer instead of lasting forever.
    • /v1/models answers Anthropic clients in the Anthropic shape: type, display_name, created_at, has_more, first_id and last_id, and, for models whose id reads as Claude, anthropic_family_tier, which Desktop uses to resolve aliases such as sonnet. The OpenAI fields stay on the objects, because clients that send only x-api-key are also read as Anthropic.

    DeepSeek Harness. DeepSeek Harness sends DeepSeek’s extensions to whichever base URL it is given: top-level dsh_* fields, among them a session log of the whole conversation of up to 8 MiB per request, role: system entries in messages, tool_addition and tool_removal blocks with defer_loading tools, thinking enabled without a budget, and x-deepseek-harness-* headers. A request is recognised by its User-Agent or a dsh_* field. To DeepSeek’s official endpoint it goes through unchanged. To any other upstream, every request from it, including count_tokens, has the dsh_* fields removed, its system entries appended to system in order, the tool change blocks and defer_loading removed and listed on the hop as a conversion lists what it drops, and a thinking without a budget rewritten to what the model accepts. The x-deepseek-harness-* headers are not forwarded, and the tool changes beta is taken out of anthropic-beta, with every line of a header sent on several lines kept. session_log_bytes gives the size of the session log a request carried, so the app can show that it carried one. /v1/files and /files answer 404: a file uploaded to the upstream picked at that moment cannot be used by a later request routed to another, and the 404 makes DeepSeek Harness send images inline.

    GLM Coding Plan quota. An upstream whose base_url host is api.z.ai or open.bigmodel.cn is read as a GLM Coding Plan upstream. Its responses carry no quota headers, so GET /quota asks the account’s quota endpoint for each enabled one that is due (at most once a minute, waiting up to 5 seconds), and traffic to one asks at most every 5 minutes; failures back off from 30 seconds to 5 minutes. Windows are 5h, weekly and monthly (the monthly MCP calls of the older plans), told apart by their length, and a reset time beyond a window’s length is dropped. Credit-based plans report credits with the three numbers the endpoint gives. A 429 carrying a used-up quota code marks the window used up and emits QuotaExhausted. A key without a plan is left alone for an hour, and a key that had a quota needs two no-plan answers in a row before it is believed, because a transient server error returns the same code.

    Antigravity CLI. Antigravity CLI sends the Go genai SDK’s default user agent. A user agent that starts with google-genai-sdk and carries Google’s internal Go toolchain (gl-go/ followed by cl/) is now reported as the client antigravity-cli. Programs built on the SDK with a public Go release are not.

    Listening on an interface that is not there yet. When listen.gateway.bind names a network interface that does not exist or is offline, or an address that is not on this machine yet, twcore used to exit at startup. The gateway now listens on 127.0.0.1 only and says why in the listen status (gw.listen.no_such_nic, gw.listen.nic_offline), asks the system again every three seconds, adds the interface once it appears, moves to its new address when it changes, and drops back to loopback when it goes away. This lets the desktop app on Windows bind the gateway to the WSL adapter, which appears only once WSL has started and changes address on every restart. A loopback that cannot be bound is still fatal at startup.

  22. ThinkWatch Core v0.49.1 GitHub ↗

    This release changes only the server guide, docs/server.md and its Chinese version, so that it describes the desktop app’s side of a remote connection as the app behaves. The binaries behave as in 0.49.0.

    Upgrade notes

    • The control-plane protocol version is unchanged (22) and so is the request store’s schema (20). A server on 0.49.0 does not need to upgrade, and ThinkWatch Lite 2026.9.17, which includes 0.49.0, connects to 0.49.0 and 0.49.1 alike.

    Connecting from the desktop app. The guide names the settings section Connection and the Add remote connection dialog, which also asks for a name. Test connection tests without saving, Save and switch tests before it saves and asks before it switches, and Save stores the connection untested; switching to a remote connection always tests it first and stays on the current connection if the test fails.

    Which version a server runs. The app connects only to a core that speaks its control-plane protocol version, which in practice means the core version the app includes, and the latest release can be newer than that. The guide therefore installs and switches with --version <version>, notes that twcore upgrade --version also installs an older release, and says that recent versions of the app (2026.9.17 and later) show sudo twcore upgrade --version <version> --restart with the version they need when the versions differ. It also says that twcore upgrade leaves the data alone, while a version whose request store has another schema starts with an empty request history.

    While connected to a remote core. The guide adds that the app signs ChatGPT accounts in with a device code only, next to the three things a remote connection cannot do.

  23. ThinkWatch Lite v2026.9.17 GitHub ↗

    Upgrade notes: the bundled core is upgraded to 0.49.0. The request store has a new format: on its first start, the new core empties the request history, the stored request and response bodies and the security log, and statistics such as usage and cost start again from zero. The configuration is kept. A remote server has to run the same core version, 0.49.0 (ThinkWatch Lite 2026.9.16 cannot connect to a server running 0.49.0), so the server and the app are updated together. When the versions differ, the app gives the command to run on the server, sudo twcore upgrade --version 0.49.0 --restart, which applies whether the installed version is newer or older. When the core on a server is upgraded from 0.47.0, its first start empties the server’s request history as well.

    Redesigned interface: every page is rebuilt on a common design, with a summary in the page header, skeletons while data loads and a Retry button when a read fails. Turning keys, upstreams, security rules and similar items on or off can be undone from the notification that follows. On macOS, the sidebar of the main window is translucent.

    Overview: the model, cache, latency and upstream rows and the failure marks can be clicked, and open Traffic filtered to what was clicked; a model is matched by its exact name. The model ranking distinguishes Unpriced, No usage and a cost that is actually zero, and when a model has unpriced requests, clicking its cost shows those requests. The latency section now covers the selected time range; it used to cover the current day only.

    Traffic: the page header shows the requests in progress, the failed requests and the requests per minute over the last 30 minutes; clicking the failed count shows failed requests only. A request in progress shows how long it has been running, updated every second, and belongs to its session from the moment it starts. Requests that a rule denies before an upstream is chosen, and requests that no upstream could serve, are now recorded, marked Denied or No upstream, and counted as failures. The Routing tab of the request details lists the route, the matched rule, the group, the rules that rewrote the request, the rule that denied it and the reason. The token column shows the total input, including cache reads and cache writes, followed by the output.

    Clients and Keys: the Clients page counts clients by state in its header, and its Set up by hand and Not detected sections can be collapsed. On the Keys page, each key shows its requests per hour over the last 24 hours and is marked In progress while one of its requests is running.

    Upstreams: each upstream shows its vendor’s logo and its requests per hour over the last 24 hours, with the hours that had failures marked in red; the row menu has a new View traffic item. The email and plan of a ChatGPT account are now read from the saved credential, without a network request, and are the same everywhere; once the sign-in has expired, the account is still named and its quota is no longer queried. Plans carry the names that Codex gives them. A completed ChatGPT sign-in names the account that was signed in, and a completed Z.ai sign-in no longer occasionally leaves out the account name.

    Routing: a map at the top of the page shows where requests go, from keys through routes and groups to upstreams. Hovering a table row or a node highlights the paths through it, and a request in progress lights its path according to core’s routing decision. Each rule shows its hits over the last 7 days, a rule without hits in that time is marked No hits, and on the map the lines leaving a route grow thicker with the number of requests. When fewer than 7 days of requests are recorded, the counts are labelled with the span the records cover, such as 5-hour hits.

    Security and MCP: the Security page header shows how many protections are enforcing, observing or off, and the exact number of matches in the selected period, broken down by outcome. The log is grouped by day. A single click opens a rule, where a double click was needed before. The MCP page header summarizes the findings by severity and the time of the scan; clicking a server row shows its configuration in each client, with the fields that differ marked.

    Settings: an index of the sections on the left follows the scrolling and marks a section with unsaved changes; unsaved changes remain after switching pages. The page header shows the current connection, the gateway address and the version. Uninstall is a confirmation dialog that lists what will be done and reports the result of each step, and its title says when a step failed.

    Command palette: ⌘K opens the command palette (on Windows and Linux, Ctrl takes the place of ⌘ here and below). It opens pages; runs actions such as creating items, the inference and connection tests, the routing dry run and switching connections; and searches upstreams, keys, routes, clients, Settings sections, the requests already loaded and more. ⌘1 to ⌘9 switch pages, ⌘R now refreshes the data of the current page, and ? outside a text field shows the keyboard shortcuts.

    Menu bar: the running time under In Progress in the menu is now computed by core, so a difference between the clocks of the two machines no longer affects it when the app is connected to a remote core. In the English interface, the text in the Quota and Today parts of the macOS menu bar menu no longer overlaps, and the tray on Windows and Linux names the local connection This computer.

  24. ThinkWatch Core v0.49.0 GitHub ↗

    This release states how far back the route and rule hit counts reach, and reports the signed-in account in the status of a ChatGPT sign-in. Release pages now begin with a summary of the release and list the downloads for each platform with the commands for a server installation.

    Upgrade notes

    • The control-plane protocol version (CONTROL_API_VERSION) is now 22. ThinkWatch Lite connects only to a core with the same protocol version, so a server has to run the core version that the app includes. ThinkWatch Lite 2026.9.16 includes 0.47.0 (protocol 20) and does not connect to 0.49.0: a server used with it stays on 0.47.0 until the app is updated to a release that includes 0.49.0. sudo twcore upgrade --version 0.47.0 --restart switches a server back to 0.47.0.
    • GET /summary/routes returns an object, RouteStats { covered_since_ms, routes }, instead of an array of RouteHits; routes holds what the array held.
    • ChatgptLoginStatus.plan is removed. ChatgptLoginStatus.account carries the email and the plan.
    • The request store’s schema is unchanged (20), so 0.48.0 and 0.49.0 keep each other’s request history. Upgrading from 0.47.0 or an earlier version empties the request history on the first start, as 0.48.0 does.

    How far back the hit counts reach. GET /summary/routes counts the stored requests, so its counts reach back only as far as the stored history. On a new installation, after an upgrade that emptied the request history, or when requests are kept for fewer days than the window, the oldest stored request is later than the start of the window, and before it a rule without hits is unknown rather than unused. covered_since_ms is the later of the window’s start and the time the oldest stored request started; when it is set, it lies within the window. It is null when the stored history covers no part of the window: the store holds no request, or the oldest one started at or after the end of the window.

    The account after a ChatGPT sign-in. The status of a ChatGPT sign-in (GET /chatgpt/login/{id}) reports the account it signed in to in account, with the email and the plan. It is the same block as ProviderView.oauth.account and is read from the same access token, the one the sign-in saves in the configuration, so the sign-in and the upstream report the same account.

    Release pages. A release is titled “ThinkWatch Core” followed by its version. Its page begins with a summary of the release when one is written, followed by a table of the files for each platform, the commands that install the version on a Linux server and switch an existing installation to it, how to verify a download against its .sha256, and the list of pull requests. The pages of 0.47.0 and 0.48.0 carry the same sections.

    Documentation. The README describes installing the prebuilt binaries, the Linux install script and running twcore as a service on a server, and lists all of the crates, grouped by dependency order, including the three that ThinkWatch Enterprise depends on (tw-dialect, tw-guard and tw-breaker).

  25. ThinkWatch Core v0.48.0 GitHub ↗

    This release reports more of what core knows about each request: the session and the route from the moment a request starts, a snapshot of running requests that a client can replay, and how often each route and rule was used. It also names the account signed in on a ChatGPT account upstream, counts the requests in each cost group that have no price or no usage, and gives the security log totals for the whole window.

    Upgrade notes

    • The control-plane protocol version (CONTROL_API_VERSION) is now 21. ThinkWatch Lite connects only to a core with the same protocol version, so a server has to run the core version that the app includes. ThinkWatch Lite 2026.9.16 includes 0.47.0 (protocol 20) and does not connect to 0.48.0: a server used with it stays on 0.47.0 until the app is updated to a release that includes 0.48.0. sudo twcore upgrade --version 0.47.0 --restart switches a server back to 0.47.0.
    • The format of the request history changed (request store schema 20). On its first start, 0.48.0 replaces the request history and the stored request and response bodies of an earlier version with an empty store; the configuration is kept. Switching back to an earlier version empties it again.
    • GET /in-flight returns an object, InFlight { now_ms, requests: [{ id, events }] }, instead of an array of start events. RequestStarted.session_fp is replaced by session. ChatgptUsage.email and ChatgptUsage.plan are removed; the signed-in account is in ProviderView.oauth.account.

    Session and route from the start. The gateway assigns a request to a session when the request starts (the same conversation, within 30 minutes of its previous request), and RequestStarted carries that session id, the one the request is stored under. The first routing phase now finishes before the start event, so RequestStarted also carries the route, the rule, the strategy group and the rules that rewrote the request. RequestRouted carries the final record, including rewrites from the second phase and denied_by for a rule that denied the request there. The stored routing record has the same fields and is written from the start event, so a request that ends before routing completes still shows the route and rule it matched.

    Requests a rule decided. A request that a rule denies in the first routing phase, or that none of the upstreams a rule chose can serve, used to produce no events and no stored row. It now produces start, routed and failed events and a row, and counts among the requests and failures in /summary; a denial in the second phase now also produces the routed event. Requests that fail before a rule decides (authentication, model admission, no matching rule) are still not recorded.

    Live views. GET /in-flight returns core’s clock and, for each running request, the events seen so far in order, so that a client connecting while requests are running can replay them and compute elapsed time from core’s timestamps. The running requests in GET /live add elapsed_ms, session, route, rule, group and upstream.

    Rule hit counts. GET /summary/routes reports, for a time window (today by default), each route’s requests, failures and last request, and for each rule the requests it decided, the requests it applied to, their failures and the last one. The counts come from the stored routing records, not from the current configuration.

    Cost groups and the security log. Every group returned by GET /summary/buckets/by reports unpriced_requests and no_usage_requests, so a group with no cost can be told apart as unpriced, missing usage or free. GET /security/events adds total and by_outcome (recorded, replaced, cut, blocked): the counts for everything the query’s guard and time window match, the same on every page.

    ChatGPT account upstreams. ProviderView.oauth.account gives the email and plan of the account signed in on a ChatGPT account upstream. Both are read from the access token core already keeps with the credential: nothing is requested from the network, nothing new is stored, and the account and user ids in the token are not exposed. The plan is reported as the backend names it, including plans that core does not list.

    Releases. A release is now published by a single job once every platform has been built and each file matches its SHA-256. Previously the first platform to finish created the release, and made it the latest, while the others were still building, and each build job added another copy of the release notes. The server guide, docs/server.md, states that ThinkWatch Lite keeps a remote connection’s key in a private file in its data directory, and applies to the app on macOS, Windows and Linux. The crates carry homepage and documentation links, and twcore --help begins with the binary’s description.

    Tests. Tests no longer hand core a port that was just released, which could make tests that bind ports, such as those of the remote control port, fail intermittently on Linux.

  26. ThinkWatch Lite v2026.9.16 GitHub ↗

    Upgrade note: the proxy configuration now rejects misspelled field names, such as typ written for type. Such fields used to be ignored; after the upgrade, a configuration that contains one does not pass validation and the gateway does not start. A proxy configuration that was edited by hand should be checked before upgrading.

    Remote connections: the app can now connect to a core running on another machine, such as a Linux server. A remote connection is added in Settings → Connection with its address, port and key, and once the connection test succeeds the app can switch to it. Installing and configuring the server side is described in docs/server.md in the ThinkWatch-Core repository. The app connects to one core at a time: while it is connected to a remote core, the local core is stopped and its data is kept. When switching, the clients already connected on this computer can be pointed at the server in the same step.

    When a connection fails, the interface remains usable. At launch the app waits up to about 8 seconds, then shows the Not connected page, where the connection can be retried, edited or switched to this computer. If the connection drops while the app is running, pages become read-only and a notice is shown once; it is removed automatically when the connection returns. The menu in the menu bar, or in the tray on Windows and Linux, has a new Connection submenu. When Option is held at launch (Alt on Windows), or after two launches in a row did not finish, the connection choice is shown first.

    In remote mode: the Clients and MCP pages check and change the client configuration on this computer; Settings is divided into two groups, the app on this computer and the configuration on the server; ChatGPT accounts sign in with a device code only; the diagnostic bundle is not offered.

    The control channel between the app and core now begins with an encrypted handshake, and local and remote connections use the same key, which is written in the configuration file. Connecting clients, MCP management and the client configuration scan are now carried out by the app itself.

  27. ThinkWatch Core v0.47.0 GitHub ↗

    This release adds a remote control port, through which ThinkWatch Lite on another machine can manage a core running on a server, together with an install script, a systemd unit and twcore upgrade for running twcore as a service on Linux. ThinkWatch Lite 2026.9.16 includes this version.

    Upgrade notes

    • A misspelled field under proxies (for example typ: for type:) is now a configuration error; it used to be ignored. A configuration that contains one does not pass twcore check, and core does not start with it.
    • Every connection to the control plane now begins with a Noise handshake, and control requests no longer use a bearer token. HTTP clients such as curl cannot call the control plane directly; twcore call sends a request through the handshake.
    • Client adoption, MCP editing and the client configuration scan moved to ThinkWatch Lite, which runs on the machine where those files are. The twcore scan and twcore clients subcommands and the related control-plane endpoints and events are removed; POST /clients/{id}/key remains.

    Control key. listen.control.key in config.yaml holds the key, and twcore serve adds one to a configuration that has none. Every control transport (the unix socket, the loopback port on Windows and the remote port) starts with a Noise_NNpsk0_25519_ChaChaPoly_BLAKE2s handshake keyed by it, and a wrong key is refused before any HTTP. twcore control-key prints the key; --rotate replaces it and closes the connections made with the previous one. The key is masked in GET /config, in the configuration history and in diagnostic bundles, and it cannot be changed through the control plane.

    Remote control port. listen.control.remote opens a TCP port for ThinkWatch Lite on another machine, in addition to the local channel. twcore init writes the section disabled, with a random port from 20000–32000. twcore remote enable [--bind B] [--port N] [--allow CIDR], twcore remote disable and twcore remote show change and report it, and a running core applies a change within a second. A source outside allow_from (the private ranges by default; loopback is not added automatically) is closed before the handshake, and a source with five failed handshakes within 60 seconds is ignored for 60 seconds. Remote connections cannot shut core down, download diagnostics or change listen.control.

    Server deployment. scripts/install.sh installs twcore on Linux (x86_64, aarch64) as a systemd service, with a thinkwatch system user, data in /var/lib/thinkwatch and environment variables in /etc/thinkwatch/env; it checks the SHA-256 and never starts or restarts the service. Releases now include twcore-<target>.tar.gz for both Linux targets, holding the binary, twcore.service and LICENSE. twcore upgrade [--check] [--restart] [--version X.Y.Z] replaces a separately installed twcore with a release from GitHub after checking its SHA-256 and running it once; it does not replace the copy inside ThinkWatch Lite, which the app updates itself. The guide is docs/server.md.

    Configuration reference. docs/config.md (and docs/config.zh-CN.md) describes every section and field of config.yaml. The field tables and the lists of built-in rules are generated from the code, and a test fails when the manual and the code disagree.

    Fix. On Linux, changing listen.gateway.bind between all and a single address on the same port now takes effect without a restart; it used to be reported as a port in use by another program.

  28. ThinkWatch Lite v2026.9.15 中文 GitHub ↗

    升级须知:请求记录会清空。上游请求头中如写有 {{client}},升级前请先删除,该写法已不再支持,否则配置无法通过校验、网关不会启动。

    新增 Linux 版本(x86_64 与 aarch64 的 AppImage),支持应用内更新。可用一行命令安装:curl -fsSL https://github.com/ThinkWatchProject/ThinkWatch-Lite/releases/latest/download/install.sh | sh

    安全页新增三项防护:请求中的隐藏字符(标签字符、双向控制符)、内容规则(内置规则可逐条启停和改处置,另可添加「包含」或「正则」自定义规则)、输出长度上限。三项均可选择关闭、仅记录或拦截,命中会记入安全日志、请求详情和概览。

    安全加固:转发时不再跟随上游的重定向,避免上游凭据被带到其他地址;WebSocket 请求中的网关密钥不再发往上游;配置、请求记录等含密钥的文件在创建时即只允许本人读写;工具调用审查可识别更多 SSE 写法;窗口权限按需收紧。

    提醒改为按当前状态核对:应用启动或重新连接网关时,已存在的凭据失效、配置未通过校验等问题会补发提醒,已解决的会自动收起,同一问题不会重复提醒。系统通知和界面中的错误原因均显示为中文。

    其他:重放请求与正常转发使用同一出站设置;菜单栏的进行中请求和生成速率改由网关统计;网关已在运行时重新打开窗口不再播放启动动画;上游表单的请求头列表第一行显示密钥所用的认证头。

  29. ThinkWatch Enterprise v2.0.0 GitHub ↗

    Callers now get errors in their own API’s format, and an upstream that refuses a request no longer takes a model’s other routes down with it. The gateway also speaks two more client protocols: Gemini, and the Responses API over a WebSocket. Cached input is billed at cache prices, and a request with no usage report is billed on an estimate instead of at zero. The TOTP requirement, which never took effect before, is now enforced. This is a major release because error bodies, the content-filter preset ids and the TOTP behaviour all change in ways a client or a script can notice.

    Read before upgrading#

    • Check security.totp_required before you upgrade. In 1.x this setting never took effect: it is stored as a boolean and was read as a string, so it always read as off. From 2.0.0 it is enforced by the server. Find out what it is set to:

      SELECT value FROM system_settings WHERE key = 'security.totp_required';

      If it is true, every console user without TOTP, super admins included, is held at a TOTP setup screen on their next request, and sessions that are already open are held too. Until they set up TOTP, every console and admin endpoint answers 403 totp_enrollment_required, except /api/auth/me, logout, register-key and the TOTP status/setup/verify-setup calls. Setting up TOTP releases the session straight away. API keys are not affected: gateway, MCP and console tw- key traffic keeps working. While the setting is on, POST /api/auth/totp/disable is refused with 400. The setting must now be a JSON boolean; a string such as "true" is refused on save.

    • Gateway error bodies follow each client API’s own format. Status codes and Retry-After are unchanged. Code that reads the error type needs updating:

      • Chat Completions and Responses keep the {"error": {"message", "type", …}} shape, but type is now OpenAI’s value for the status, not a ThinkWatch tag: authentication_error (401), permission_error (403), not_found_error (404), rate_limit_error (429), invalid_request_error (other 4xx) and server_error (5xx). The old tags are gone: rate_limited, policy_blocked, provider_http_error, provider_error, provider_timeout, transform_error, network_error and auth_error. A policy block is now permission_error with status 403.
      • Anthropic Messages clients get Anthropic’s body, {"type": "error", "error": {"type", "message"}}, with Anthropic’s type names (rate_limit_error, overloaded_error, api_error, …).
      • Gemini clients get Google’s body, {"error": {"code", "message", "status"}}.
      • Once a stream has started, a Responses client gets a response.failed event, where before it got a Chat-style error frame that SDKs skip, so the stream just stopped. An Anthropic client gets an error event whose type follows the status.
      • The error_type field in gateway_logs and the metric labels are unchanged.
    • Only upstream failures fail over or count against a route’s breaker. In 1.x every non-2xx except 401, 403 and 429 was retried on the model’s other routes and counted as a failure on each of them, so one malformed request could open the breakers on all of a model’s routes.

      • Tried on another route and counted against this one: 5xx, 408, 429, 401 and 403 (the upstream refused the gateway’s own credential), timeouts, network errors and unreadable responses.
      • Returned to the caller straight away, and counted as the upstream working: every other 4xx.
      • What the caller sees: such a 4xx comes back with its own status and the upstream’s reason. In 1.x it came back as a 502 after every route had been tried.
      • An upstream 5xx comes back with the upstream’s status (500, 503, …) rather than a blanket 502.
      • An upstream timeout is now 504.
      • Streams follow the same rule. A stream cut by tool-call inspection no longer counts against the route.
    • Cached input is billed at cache prices. In 1.x cache reads and writes were billed, and debited from budgets and weighted rate limits, as full-price input. models gains three weights, cache_read_weight, cache_write_weight and cache_write_1h_weight. When a weight is unset, it is input_weight times Anthropic’s ratio: 0.1× for a read, 1.25× for a write and 2× for a one-hour write. What this changes:

      • Traffic with many cache reads (Claude Code, for instance) costs much less than it did.
      • Traffic that writes to the cache costs a little more.
      • Older OpenAI models discount cache reads less (0.5× or 0.25×). Set the weights on those models yourself.
      • input_tokens in the log is still the whole input. The log detail gains cache_read_tokens, cache_write_tokens and cache_write_1h.
    • Output length limits now apply to streams. In 1.x, max_length output guardrails checked only whole responses, so streamed answers were never checked. Now the frame that would cross the limit is not sent, and the stream ends with an error in the caller’s format. A response served from the cache is also checked against the limit in force. If you set a limit, streamed answers that used to go through can now be cut off.

    • Content-filter preset groups are renamed. The groups are now injection, persona and chinese; they used to be basic, strict and chinese. This matters only if you call the presets endpoint by group id. Rules you have already added are copies and are not affected. Other changes to the filter:

      • A rule with an empty pattern is now refused on save.
      • Each text part of a message is scanned separately, so a pattern no longer matches across two parts.
      • The engine is now shared with ThinkWatch-Core’s tw-guard. The stored format and the admin API are unchanged.
    • Requests with no usage report are billed on an estimate. In 1.x such a request was billed at zero. This happens when an upstream ignores the request for usage, or when the caller leaves before the final chunk arrives. The estimate is:

      • input: about four bytes of the request per token, not counting images and files;
      • output: the answer that actually arrived.

      Estimated rows carry usage_estimated: true in their detail and count in gateway_usage_estimated_total. A request with no answer at all is still billed at zero.

    • Clients that leave early are now logged. In 1.x a client that disconnected before its response existed left no gateway_logs row at all. That covers leaving during auth, limits or routing, or while waiting for a whole (not streamed) answer. Such a request now writes one row: status 499, stream_outcome: client_cancelled, cancelled_before: response, no tokens and no cost. Expect more 499 rows in dashboards and log forwarders. A new counter, gateway_cancelled_before_response_total, counts them.

    Database changes#

    Both apply on their own at startup, as every schema change does, and both are additive:

    • models gains three nullable columns: cache_read_weight, cache_write_weight and cache_write_1h_weight (ALTER TABLE … ADD COLUMN IF NOT EXISTS, CHECK (>= 0)).
    • system_settings gets an auth.default_role row, seeded empty (no role). Existing rows are left alone (ON CONFLICT DO NOTHING).

    Neither is irreversible. A 1.1.0 server runs against the upgraded database: it ignores the new columns and the new setting. What a 1.1.0 server cannot do is price cache tokens from the weights.

    Added#

    • Gemini clients. New endpoints:

      • POST /v1beta/models/{model}:generateContent and :streamGenerateContent, also served under /v1/models/…;
      • GET /v1beta/models, which lists models in Gemini’s format.

      These requests get the same limits, budgets, filters, routing with failover, format conversion, inspection, billing and audit as every other endpoint. A Gemini upstream gets the request as it was sent. A stream comes back as SSE with alt=sse, and as Gemini’s JSON array without it. :countTokens and :embedContent are refused with 400.

    • The Responses API over a WebSocket. Connect to GET /v1/responses with Upgrade: websocket.

      • Each response.create frame is handled like a streamed POST /v1/responses, with its own limits, routing, billing and audit row.
      • Turns on one connection run in order. A refused turn fails with response.failed, and the connection stays open.
      • The connection keeps its latest response. That lets a turn continue from it with previous_response_id, even with store: false, which is how Codex works, and against any upstream format.
      • A new counter, gateway_responses_ws_connections_total, counts connections.
    • More places to put an API key. Gateway keys are also accepted in x-api-key (Anthropic SDKs), x-goog-api-key and ?key= (Gemini SDKs), as well as Authorization: Bearer. Headers are checked first. A key given in the query string is never sent upstream.

    • Hidden-text audit events show what the text says. Each item in found gains revealed, the ASCII that the hidden tag characters spell.

    • auth.default_role can be set. It is the role that newly registered users and SSO users get. In 1.x, setting it through the admin API reported success but changed nothing, because the setting row did not exist.

    • Model editor fields for the three cache weights. Each placeholder shows the value used when the field is left empty.

    Changed#

    • Default output length for upstreams that require max_tokens. When the caller sets none, the gateway now sends 32000 for Claude models and 8192 for other models. It used to send 4096, which cut Claude answers short.
    • More upstreams count as the vendor’s own endpoint. DeepSeek, Moonshot, Zhipu/Z.ai, DashScope, xAI and *.amazonaws.com are now recognised, and the check reads the parsed host. A relay URL such as https://relay/api.openai.com no longer passes as official. Official endpoints are stricter about request parameters, so the gateway drops or renames some parameters before sending to them.
    • Hidden-text scanning uses ThinkWatch-Core’s tw-guard. Same scope, same actions, and nothing is stripped.
    • Requests forwarded in their own format lose ThinkWatch’s reasoning signatures. A tw1. signature written by an earlier format conversion is removed, because Anthropic rejects it. The upstream’s own signatures are kept.
    • Core crates: tw-dialect, tw-guard and tw-breaker at ThinkWatch-Core v0.43.0. The code only this edition used (at-rest crypto, SigV4 signing, the gateway error type) moved into this repository. It works the same, and stored secrets decrypt as before.
    • The server’s SQL moved from the request handlers into repository modules (catalog, dashboard, limits, log forwarding, identity, access, MCP). Every statement is unchanged. New integration tests cover these endpoints and pass on both the old and the new code.
    • CI runs on pull requests into dev, including the whole integration suite against Postgres, Redis and ClickHouse.

    Fixed#

    • Revoking a user’s default MCP connection always failed with a 500, and the account could not be revoked. The newest remaining account is now made the default.
  30. ThinkWatch Core v0.46.0 GitHub ↗
  31. ThinkWatch Core v0.45.0 GitHub ↗
  32. ThinkWatch Core v0.44.0 GitHub ↗
  33. ThinkWatch Core v0.43.0 GitHub ↗
  34. ThinkWatch Core v0.42.0 GitHub ↗
  35. ThinkWatch Enterprise v1.1.0 GitHub ↗

    The gateway stops rebuilding every request as a chat-shaped message. A request whose route speaks the caller’s own format goes out as the caller sent it; one that crosses formats is converted by ThinkWatch-Core, the same layer the desktop edition uses. Tools, tool choice, system prompt blocks, metadata and cache_control now reach the upstream, where they used to be dropped. Two new checks guard what goes in and out: tool calls an upstream returns, and invisible characters in what a caller sends.

    Read before upgrading#

    • Anthropic routes record more prompt tokens for the same work. Prompt tokens now count the same for every upstream: plain input plus cache reads and cache writes, which is OpenAI’s definition. Anthropic’s own input_tokens leaves the cached part out, so on those routes prompt tokens, cost and budget use all go up. The price model still charges every prompt token at one rate.
    • Requests in the upstream’s own format are forwarded as sent. Every field the caller sends reaches the upstream, along with the caller’s anthropic-beta and anthropic-version headers. Only the model name changes, and PII is swapped for placeholders. An OpenAI-compatible upstream that rejects fields it does not know may now refuse requests that used to succeed, because those fields were stripped before. Try your upstreams with the clients you actually run.
    • Two checks are on by default, and neither blocks anything yet.
      • Tool-call inspection starts in observe mode.
      • The hidden-character check starts in warn mode.
      • Both write audit events, so expect new entries in the audit log and in anything subscribed to it.
      • Nothing is refused until you switch to enforce or block.
    • The response cache starts cold. The cache key now covers the whole request, so entries written by 1.0.2 are never hit again. They expire on their own.
    • One PII value gets one placeholder. Within a request, the same e-mail address is {{EMAIL_1}} wherever it appears. It used to get a new number each time, so a model saw one person as several. Saving a PII pattern now also requires the placeholder prefix to be letters, digits or underscores.

    Added#

    • Tool-call inspection (security.tool_inspection).
      • Why: an upstream writes the response, so it can hand the caller a tool call the model never made, such as bash("curl … | sh") appended to an ordinary answer. An agent set to auto-approve then runs it.
      • Rules: a built-in set of dangerous-command rules. Each can be switched off or given a different action, and you can add your own.
      • Modes: off, observe (records hits and changes nothing on the wire) and enforce.
      • Enforce on a stream: the stream is cut at the frame that would complete a matching call, and the refusal arrives in the caller’s format.
      • Enforce on a whole response or a cache hit: refused with 403 (policy_blocked). A refused answer is neither cached nor billed.
      • Audit and metrics: every hit is an audit event (gateway.tool_call_flagged or gateway.tool_call_blocked) and counts in gateway_tool_call_flagged_total.
      • Admin API: GET /api/admin/settings/tool-inspection/rules and POST /api/admin/settings/tool-inspection/test.
      • Console: a card on the security page, plus a sandbox tab.
    • Hidden-character check (security.hidden_text: off, log, warn or block).
      • What it looks for:
        • Unicode tag characters, which carry an instruction invisibly into the model’s context;
        • bidirectional overrides, which make text read differently on screen than it is.
      • Where: the caller’s messages and the tool results inside them. The system prompt and the model’s own turns are not checked.
      • Not flagged: zero-width joiners (emoji), the zero-width non-joiner (Persian) and Cyrillic.
      • Actions: warn writes gateway.hidden_text_flagged; block refuses with 403 and writes gateway.hidden_text_blocked.
      • Console: a card on the security page.
    • Tool-call arguments get their PII back. A model asked to e-mail a@example.com used to call the tool with {{EMAIL_1}} as the address.

    Changed#

    • One pipeline for /v1/chat/completions, /v1/messages and /v1/responses. A same-format request is forwarded as sent. A cross-format request is converted, and the gateway logs what the target format cannot carry.
    • Content filtering and PII detection read tool results too. That is where an injected instruction, or customer data pulled in by a tool, usually sits.
    • Usage is read off the upstream’s own bytes. A streamed response is no longer held in memory for an accounting pass at the end.
    • Chat streams are billed on the upstream’s actual usage. They are always sent asking for it. A caller who did not ask for the usage chunk still does not get one. 1.0.2 estimated the count for these.
    • Streams send their headers at once. A caller who leaves while the upstream is still thinking is recorded as cancelled.
    • Connectivity tests use the live encoder. A route’s test request is built by the same encoder as real traffic, so a passing test means forwarding works.
    • The web console loads data through TanStack Query.
      • Signing out, including from another tab, clears everything cached.
      • After a change, screens refresh in place.
      • Polling pauses while the tab is hidden.
    • Core crates come from one pinned tag (ThinkWatch-Core v0.40.0), declared once at the workspace root.

    Fixed#

    • Requests lost their tools, tool choice and non-text content on the way upstream (ThinkWatch-Core#50). Claude Code’s system prompt, sent as an array, was dropped whole, and so was every cache_control breakpoint. Each cached prefix was billed as full-price input.
    • The response cache could serve the wrong answer. Its key covered only model, messages and max_tokens, so two requests that differed only in tools, response_format, seed and so on shared one entry.
    • A tripped route stayed out until its Redis key expired, roughly four cooldowns. It now gets probed once the cooldown is over, and a success closes it.
    • The dashboard showed every AI provider’s breaker as Closed. It now shows the real state of the provider’s routes, reporting the worst one.
    • The PII “try patterns” endpoint misreported labels. A pattern named CUSTOM_EMAIL was reported as CUSTOM.
    • :latest could point at a main build rather than the release, which is what happened for v1.0.2’s server image. Only the release workflow sets :latest now.
    • The web image hung for six hours. Its frontend was built under QEMU for arm64, where Node crashed and the step never returned. It is now built once, natively. The image contents are unchanged.

    Security#

    • Refreshed the web console’s lockfile to clear 56 Dependabot alerts (1 critical, 23 high). All were transitive, and none of them reached the shipped bundle.
  36. ThinkWatch Core v0.41.0 GitHub ↗
  37. ThinkWatch Lite v2026.9.14 中文 GitHub ↗

    上游的 API 密钥和请求头中的 ${变量名} 现在可以读取用户配置的环境变量。macOS 上读取登录 shell 中的变量,~/.zshrc 等文件中 export 的变量都会生效,此前从程序坞或访达打开时读不到;Windows 上每次启动网关都重新读取系统设置中的环境变量。代理相关的变量和 PATH 不会被读取。修改变量后重新打开应用即可生效。

    上游表单中,API 密钥与请求头的说明改为写明读取系统环境变量,请求头可一键插入网关密钥名称和 Access Token。

    熔断改用与企业版共用的状态机,行为不变:连续失败三次后熔断 60 秒,再放行一次探测;所有上游都不可用时照常放行,只有一个上游时不熔断。

  38. ThinkWatch Lite v2026.9.13 中文 GitHub ↗

    自动检查发现新版本时改为发送系统通知,点按通知后再打开更新窗口,不再自行弹出。更新窗口只显示版本号,不再显示更新说明。

  39. ThinkWatch Core v0.40.0 GitHub ↗
  40. ThinkWatch Lite v2026.9.12 中文 GitHub ↗

    Windows:网关监听选择「局域网」时,可以选用系统中的网卡(如「以太网」),此前保存会失败。快捷键改用 Ctrl(Ctrl+F 搜索、Ctrl+B 收起源列表),字号与字体按 Windows 调整,「访达」「菜单栏」等说法改为 Windows 对应的名称。安装版只运行安装目录中的网关,「关于」中显示的路径也已修正。

    开机启动的说明改为一句话,直接写明开启后的效果。

  41. ThinkWatch Core v0.39.0 GitHub ↗
  42. ThinkWatch Core v0.38.0 GitHub ↗
  43. ThinkWatch Core v0.37.0 GitHub ↗
  44. ThinkWatch Lite v2026.9.11 中文 GitHub ↗

    修复 2026.9.10 中所有修改操作都会失败的问题:保存设置、新建或编辑上游、接管客户端等操作会提示控制面需要凭据,本版恢复正常。

    Windows:启动时不再弹出网关的命令行窗口;客户端页中的配置文件路径改为 Windows 的写法,Zed 的配置位置以及 Continue、Gemini CLI 的手动配置步骤按 Windows 给出。

    报错信息改为按界面语言显示,系统或外部服务返回的原文除外。

  45. ThinkWatch Core v0.36.0 GitHub ↗
  46. ThinkWatch Core v0.35.0 GitHub ↗
  47. ThinkWatch Core v0.34.0 GitHub ↗
  48. ThinkWatch Lite v2026.9.10 中文 GitHub ↗

    首次提供 Windows 版,分 x64 与 arm64 两个安装程序,在发布页下载。Windows 版未做代码签名,首次运行时 SmartScreen 会拦截,选择「更多信息」→「仍要运行」即可。

    新增 Z.ai / BigModel 账号登录:在上游的服务类型中选择「Z.ai / BigModel 账号(登录)」,在浏览器中授权后,账号下会生成一把 API 密钥并写入这个上游。

    客户端页中的 Codex 一行同时对应 Codex 命令行和 ChatGPT 桌面应用内置的 Codex,二者共用同一份配置;接管时一并说明,ChatGPT 桌面应用需要重启才会生效。

    修复:更新窗口底部的按钮被窗口边缘截断;在接管对话框中展开「完整改动」后,内容超出对话框。

  49. ThinkWatch Core v0.33.0 GitHub ↗
  50. ThinkWatch Core v0.32.0 GitHub ↗
  51. ThinkWatch Core v0.31.0 GitHub ↗
  52. ThinkWatch Core v0.30.0 GitHub ↗
  53. ThinkWatch Core v0.29.0 GitHub ↗
  54. ThinkWatch Lite v2026.9.9 中文 GitHub ↗

    升级须知:2026.9.8 及更早的版本无法在应用内安装这一版,需从发布页下载磁盘映像(DMG)手动安装一次,此后的版本可照常在应用内更新;通过 Homebrew 安装的执行 brew upgrade 即可。

    推理测速支持全部类型的上游。ChatGPT 账号上游此前无法测速,只会返回一句英文提示,现在与其余上游一样发送探测请求并测量首 token 时间,失败原因按界面语言显示。ChatGPT 账号不接受输出上限,费用预估中的上限显示为「不限」、金额显示为「按实际用量」,合计随之标为无法计算;会推理的模型在探测时留出推理所需的输出额度,避免上限过小导致没有可见的输出。

    发布内容只保留磁盘映像及其校验值,应用内更新改为下载与手动安装同一份磁盘映像。

  55. ThinkWatch Core v0.28.0 GitHub ↗
  56. ThinkWatch Lite v2026.9.8 中文 GitHub ↗

    升级须知:升级后本机的请求记录会清空。部分旧设置项已删除,配置文件中仍有这些项时网关无法启动,应用会进入安全模式;删除这些项后重新启动即可恢复。用 ChatGPT 账号登录的上游写入的「billing: subscription」属于这一类。

    界面支持英文,默认跟随系统语言,也可在「设置」中选择;菜单栏、系统通知和错误信息使用同一种语言。外观可选浅色、深色或跟随系统。

    导航重新整理为概览、流量、客户端、密钥、上游、路由、安全、MCP 和设置。

    流量页将请求与会话合为一张表,可按会话归组;每条请求标出发出它的密钥、应用和来源机器,请求详情中的 JSON 会格式化显示。概览的实时图显示最近十分钟,切换页面后保留所选的时间范围。

    路由页以密钥为中心重新设计,规则按顺序编辑;试算会给出请求使用的路由,不经过任何规则的请求也会如实说明。

    密钥单独成页,显示完整密钥并可复制,标出由哪个客户端接管生成,可安全地更换;每把密钥可设置可见模型和并发上限,全局并发限制已删除。客户端页用一张表列出各客户端的状态和所用密钥,未检测到的客户端也给出手动配置方法;「全部还原」此前只还原 opencode,现在会还原所有已接管的客户端。

    安全页只保留出站脱敏和工具调用审查两项全局防护,匹配规则可查看和自定义,并新增安全日志;MCP 扫描单独成页。

    计费方式只分按量计费和不计费,ChatGPT 等订阅账号也按价目表计算费用。上游行可直接展开模型列表,缺失的列表会自动获取;ChatGPT 账号上游显示邮箱和套餐,也可在另一台已登录的设备上输入代码完成登录。

    菜单栏改为原生样式,显示今日 token 和今日费用,菜单中可查看额度、进行中的请求和提醒。提醒只保留一个总开关,可标为已读;不再监测磁盘空间。

    网关监听的访问范围、网卡、端口与放行网段移入「设置」,保存后生效;日志保留时长也可在「设置」中调整。启动时显示启动画面,网关就绪后进入主界面,各页面随数据变化实时更新。

  57. ThinkWatch Core v0.27.0 GitHub ↗
  58. ThinkWatch Core v0.26.0 GitHub ↗
  59. ThinkWatch Core v0.25.0 GitHub ↗
  60. ThinkWatch Core v0.24.1 GitHub ↗
  61. ThinkWatch Core v0.24.0 GitHub ↗
  62. ThinkWatch Core v0.23.0 GitHub ↗
  63. ThinkWatch Core v0.22.0 GitHub ↗
  64. ThinkWatch Core v0.21.0 GitHub ↗
  65. ThinkWatch Core v0.20.0 GitHub ↗
  66. ThinkWatch Core v0.19.0 GitHub ↗
  67. ThinkWatch Core v0.18.0 GitHub ↗
  68. ThinkWatch Core v0.17.2 GitHub ↗
  69. ThinkWatch Core v0.17.1 GitHub ↗
  70. ThinkWatch Core v0.17.0 GitHub ↗
  71. ThinkWatch Core v0.16.0 GitHub ↗
  72. ThinkWatch Core v0.15.0 GitHub ↗
  73. ThinkWatch Core v0.14.1 GitHub ↗
  74. ThinkWatch Core v0.14.0 GitHub ↗
  75. ThinkWatch Core v0.13.0 GitHub ↗
  76. ThinkWatch Core v0.12.0 GitHub ↗
  77. ThinkWatch Core v0.11.1 GitHub ↗
  78. ThinkWatch Core v0.11.0 GitHub ↗
  79. ThinkWatch Core v0.10.0 GitHub ↗
  80. ThinkWatch Core v0.9.1 GitHub ↗
  81. ThinkWatch Core v0.9.0 GitHub ↗
  82. ThinkWatch Core v0.8.2 GitHub ↗
  83. ThinkWatch Core v0.8.1 GitHub ↗
  84. ThinkWatch Core v0.8.0 GitHub ↗
  85. ThinkWatch Core v0.7.0 GitHub ↗
  86. ThinkWatch Lite v2026.9.7 中文 GitHub ↗

    通知改为原生 macOS 通知:同一条会就地更新,过期的会自己撤回,点开直接跳到对应的页面。

  87. ThinkWatch Lite v2026.9.6 中文 GitHub ↗

    可以用 ChatGPT 账号新建上游:在浏览器中完成登录后,订阅额度即可供各类客户端使用。登录前会说明相关风险;账号的额度与额度重置卡在该上游的菜单中查看和使用。

    不同格式之间的请求可以互相转换:使用 Anthropic、OpenAI Chat、OpenAI Responses 或 Gemini 格式的客户端,可以使用其中任何一种格式的上游。请求详情中会显示转换情况与未能保留的字段。

    新增提醒:订阅额度用完、登录失效、上游拒绝凭据、代理不通、磁盘空间不足等情况记录在工具栏的提醒列表中,需要处理的会弹出系统通知。每一类提醒可在「设置」中选择弹出通知、仅在应用内显示或关闭。

    上游页重新设计:上游、代理与价目表分为三个标签,新建与编辑均在对话框中完成;上游可以设置自定义价目表与额外的请求头。

    请求列表会标出被取消的请求、部分失败的请求与缺少用量的请求。

  88. ThinkWatch Core v0.6.0 GitHub ↗
  89. ThinkWatch Core v0.5.0 GitHub ↗
  90. ThinkWatch Core v0.4.0 GitHub ↗
  91. ThinkWatch Core v0.3.0 GitHub ↗
  92. ThinkWatch Core v0.2.2 GitHub ↗
  93. ThinkWatch Core v0.2.1 GitHub ↗
  94. ThinkWatch Core v0.2.0 GitHub ↗
  95. ThinkWatch Core v0.1.2 GitHub ↗
  96. ThinkWatch Lite v2026.9.5 中文 GitHub ↗

    设置里的「自动检查新版本」只保留开关;检查改为每天一次。

  97. ThinkWatch Lite v2026.9.4 中文 GitHub ↗

    安装包的窗口重新设计:打开磁盘映像即可看到拖放方向和首次打开的说明。签名改为固定的证书,此后通过 Homebrew 升级不再提示签名者变化。

  98. ThinkWatch Lite v2026.9.3 中文 GitHub ↗

    有新版本时弹出更新窗口。网页下载安装的,按一次即可完成更新,并会等正在进行的请求结束后再重启;Homebrew 安装的,窗口给出更新命令,可一键复制。更新检查默认开启,可在「设置」中关闭。

  99. ThinkWatch Lite v2026.9.2 中文 GitHub ↗

    安装包改为磁盘映像:打开后把应用拖进「应用程序」即可。

  100. ThinkWatch Core v0.1.1 GitHub ↗
  101. ThinkWatch Lite v2026.9.1 中文 GitHub ↗

    这一版会在有新版本时告诉你。检查默认关闭,在「设置」里打开。

  102. ThinkWatch Lite v2026.9.0 GitHub ↗

    The first build. An unsigned arm64 .app with the gateway inside it.

  103. ThinkWatch Core v0.1.0 GitHub ↗
  104. ThinkWatch Enterprise v1.0.2 GitHub ↗

    The shared gateway layer moves out into its own repository, and OIDC learns to accept an ID token whose aud carries more than the client ID. Mostly a fix release otherwise — the deploy and auth items below are the ones worth reading before you upgrade.

    Added#

    • OIDC_ADDITIONAL_TRUSTED_AUDIENCES — a comma-separated allowlist of extra ID-token audiences to trust. Some IdPs put something besides the client ID in aud; Zitadel includes the parent project ID, and openidconnect’s stock verifier rejects every non-client-ID audience, so login failed outright. Leave it unset and verification is byte-for-byte what it was. Set, it widens exactly one check: the extra audiences on a multi-audience token are matched against this list as exact strings — no wildcards, no prefixes. The client ID must still appear in aud regardless. Documented in .env.example. Thanks to @DaniW42 for the report and the patch (#12).
    • Automatic upstream protocol detection — the gateway learns an upstream’s dialect and relearns it when the upstream changes, instead of relying on a static guess. Protocol is now a property of the route.
    • Reverse-proxy identification — the server recognizes its own reverse proxy, which is what makes the dashboard WebSocket and the API docs work behind nginx.

    Changed#

    • azp is validated whenever it is present, per OIDC Core 3.1.3.7 step 5, which has no audience-count precondition. Previously it was only checked when a multi-audience token made it mandatory, so a single-audience token naming this client with azp pointing at a different client was accepted — the IdP stating plainly that the token was authorized for somebody else. This can reject a token that 1.0.1 accepted. If your IdP sets azp to something other than your client ID, logins will start failing and the message will say so; that is the IdP to fix, not this setting.
    • The shared gateway layer now comes from ThinkWatch-Core as a git dependency rather than living in this tree. Six crates — types, protocol, provider, resilience, crypto, and the JSON secret envelope — are one implementation shared with the desktop edition instead of two copies drifting apart. No runtime or API change; it affects you only if you build from source, where the build now needs network access to resolve that dependency. The published images are unaffected.
    • Stored secrets are redacted on read and decrypted on use. Provider credentials no longer travel in plaintext through code paths that merely display or list them.

    Fixed#

    • Deploy — nginx broke the API docs three separate ways; the dashboard WebSocket wasn’t proxied; the prod stack’s healthchecks were wrong; and the dev stack could overwrite production’s ClickHouse credentials. The last one is the reason to read this list.
    • Auth — the console logged itself out after every token refresh. Rate-limit counters could be created without an expiry and sit in Redis forever.
    • Models — the gateway no longer offers models the upstream refuses, no longer imports models it won’t serve, and reports every dialect that was refused rather than only the last. Provider base_url trailing slashes are normalized.
    • Analytics — spend is attributed to a person, not a UUID.
    • Audit — a syslog forwarder formatting fix that had CI red on clippy 1.98 (#11, thanks @DaniW42). Two more instances of the same lint, plus a result_large_err false positive in the MCP lifecycle stages, measured rather than boxed: the Ok variant is 296 bytes against the Err’s 152, so boxing saves nothing and adds an allocation.
    • UI — the brand mark stays inside the collapsed sidebar rail.

    Contributing#

    CONTRIBUTING.md and a PR template now state the branch contract that had only lived in docs/operations/release.md: routine work targets dev, and main is the release line. A workflow comments on PRs opened against main from anything other than dev or a hotfix/* branch, because GitHub pre-fills the base with the default branch and walks contributors into it.

    Added#

    • (nothing yet)

    Changed#

    • (nothing yet)

    Fixed#

    • (nothing yet)

    Removed#

    • (nothing yet)

    Security#

    • (nothing yet)

May 2026

  1. ThinkWatch Enterprise v1.0.1 GitHub ↗

    Release-pipeline validation. No product change — the published binary, REST surface, MCP wire shapes, and audit semantics are identical to v1.0.0. Operators pinning 1.0.0 have no reason to bump; those tracking :latest move forward.

    Changed#

    • Release workflow — arm64 image builds now run on a native arm64 runner (ubuntu-24.04-arm) instead of QEMU emulation. v1.0.0’s server image build took 1h24m; this should drop to ~10 min. Multi-arch manifest assembled by a new merge job via docker buildx imagetools create.
    • Node 24 opt-in — workflow sets FORCE_JAVASCRIPT_ACTIONS_TO_NODE24=true so all actions/* + docker/* run on Node 24 ahead of GitHub’s 2026-06-02 forced cutover.
  2. ThinkWatch Enterprise v1.0.0 GitHub ↗

    Stability commitment. No code delta since 0.5.0 — this tag marks the point at which the API surface becomes a SemVer commitment.

    Changed#

    • Versioning policy — from this tag onwards every breaking change (REST routes, MCP wire shapes, audit-row JSON keys, database schema, public Rust APIs in published crates) requires a major bump. Operators chasing the :latest tag on the GHCR images can do so without surprise.

    Notes#

    • Docker images cut at this tag receive :latest for the first time — the release workflow suppresses :latest on 0.x and pre-release tags. Pin the version in production rather than tracking :latest unless you have a controlled rollback path.
  3. ThinkWatch Enterprise v0.4.0

    See every byte

    • Request + response bodies captured for every gateway and MCP call
    • PII redaction at capture; per-column TTL splits hot bodies from cold metadata
    • S3-compatible body offload — AWS S3, MinIO, Ceph, or the bundled RustFS
    • Bloom-filter substring search across captured bodies
    • audit:read_bodies permission gates the body viewer
    • MCP downstream SSE — clients can subscribe to tools/call event timelines
    • Per-model routing with Auto / Manual modes + drag-to-redistribute traffic bar
    • Model-level kill switch — pause one model without the whole provider
    • Pre-call budget check rejects requests when caps are already exhausted
    • allowed_models enforced uniformly on Anthropic Messages + OpenAI Responses

    Full-body audit capture#

    ThinkWatch positioned itself as a bastion for AI traffic — but every audit row was metadata-only. Token counts, latency, cost. No record of what the user actually asked or what the model actually answered. Compliance teams couldn’t replay incidents, attest “PII X was sent”, or even tell two near-identical requests apart.

    v0.4.0 closes the gap. gateway_logs now persists request_body + response_body (with byte counts and a body_capture_status dimension); mcp_logs persists the parallel tool_arguments + tool_result. ZSTD(6) compression keeps cold storage tractable; six dynamic config toggles (capture toggles per field + a global body_max_bytes) let operators dial scope without a redeploy. Cache-hit responses now emit a full audit row instead of disappearing into a hole.

    Per-column TTL + S3 offload#

    Hot evidence and cold metadata have very different retention shapes. Per-column TTL lets you keep the last 30 days of bodies and several years of token/cost rows in the same table, without paying ClickHouse for the cross.

    For payloads that exceed body_max_bytes — 200k-context prompts, MCP file-read tools returning whole files, base64 images — bodies offload to S3 instead of being truncated. Backend is config-driven via six env vars (S3_ENDPOINT_URL, S3_BUCKET, etc.) and works against AWS S3, MinIO, Ceph, or the bundled RustFS container shipped in the dev and prod Docker Compose files. Object keys carry a date prefix (bodies/<table>/yyyy/mm/dd/...) so bucket lifecycle rules can age them off independently. The body-viewer endpoint transparently dereferences s3:// cells, so the console doesn’t know offload exists.

    Captured-body columns ship with a tokenbf_v1(512, 3, 0) index. Auditors hunting for “which conversations mentioned API key XYZ” or “which tool calls touched /etc/passwd” no longer scan every granule — the bloom filter rejects the vast majority of rows up front. Indexes are sized at ~50 KB per granule, paying for themselves the first time someone needs to answer a real compliance question.

    A new audit:read_bodies permission gates the body viewer in the console and the underlying /api/admin/{gateway,mcp}/logs/{id}/body endpoint, so the raw payload is preserved as legal record of intent while still requiring a separate grant to read.

    MCP downstream streaming#

    tools/call was buffered-only on the client side: even when the upstream emitted progress notifications, the gateway swallowed the timeline and returned a single JSON. v0.4.0 negotiates on the Accept header — clients including text/event-stream receive each upstream event (notifications/progress, partials, final response) as a discrete SSE event downstream. JSON-only clients keep the buffered shape for backward compat. The audit pipeline still gets the full event timeline either way.

    Per-model routing#

    Model-route configuration is now organised around the admin’s actual mental model: Auto (latency-cost, balanced, or latency-only) or Manual (a drag-to-redistribute traffic bar across peers). The priority-tier failover concept is gone; replaced with nginx-style flat peers + circuit-breaker bypass. Each model can override the global strategy; the dashboard surfaces routing-projection so the UI can show “your manual config will cost $X; auto would cost $Y” before you commit.

    Three new endpoints back the UI:

    • PATCH /api/admin/model-routes/batch-weights — the drag bar’s commit endpoint
    • GET /api/admin/models/{id}/routing-projection — current vs. auto split + projected $/1M tokens
    • GET /api/admin/models/{id}/route-history — 60 one-minute p50/p95 buckets for the latency sparkline

    The decision log captures provider chosen + reason + any fallback for every request, and every gateway_logs row stamps the actual upstream_model so post-mortems are unambiguous.

    Pre-call budget enforcement#

    Budget caps used to fire post-flight: a runaway client whose monthly cap was already at zero could still fire requests, and only the post-call debit would notice — by which point the upstream had already burned tokens. v0.4.0 adds a check_budget lifecycle stage that runs before the upstream call. Steady-state “your cap is exhausted, stop sending” is now blocked at the gate; the existing crossing alerts continue to cover the concurrent-burst race.

    allowed_models on every API surface#

    The per-API-key allowed_models allowlist is now enforced uniformly on OpenAI Chat Completions, Anthropic Messages, and OpenAI Responses. The same enforcement is wired into the shared lifecycle pipeline, so future surfaces inherit it automatically — no API surface bypasses the model allowlist.

    Under the hood#

    A shared lifecycle pipeline unifies the AI gateway and MCP gateway request paths behind a common Surface trait. check_limits, check_access, check_budget, record_usage are now stages applied symmetrically on both surfaces, eliminating the historical drift between buffered, streaming, and short-circuit branches. The MCP proxy.rs and gateway proxy.rs both got split into submodules, and the 1667-line dashboard component on the frontend was carved into focused subcomponents.

  4. ThinkWatch Enterprise v0.3.0

    MCP, the per-user way

    • Per-user upstream credentials — every developer authenticates as themselves to GitHub, Linear, Slack
    • One-paste OAuth onboarding via Dynamic Client Registration
    • MCP Store with bilingual templates (Linear OAuth seeded out of the box)
    • Step-by-step registration wizard + auth-mode-aware edit form
    • Three-tier upstream subject resolution (JWT + userinfo + discovery)
    • Per-credential Test Connection on /connections
    • Per-user tool catalogs — different users see different tools based on upstream permissions
    • MCP response cache scoped per (user, account_label) — no cross-user leakage
    • Security hardening from review — SSRF, cache, rate limits, audit

    Per-user upstream credentials#

    The MCP gateway no longer pretends every user is the same upstream service account. Each developer authenticates as themselves — via OAuth or PAT — to GitHub, Linear, Slack, and any other MCP-enabled service. The upstream audit trail finally works end-to-end: tickets are assigned to real people, GitHub issues are created by the engineer who actually filed them. Per-key account overrides let one user maintain multiple identities (personal + work GitHub) and pick which one a given API key uses.

    One-paste OAuth onboarding#

    Paste an MCP server URL into the registration wizard. ThinkWatch handles Dynamic Client Registration with the upstream OAuth server, runs the auth probe, captures the OAuth client credentials at install time, and walks you through the consent screen. 401/403 from anonymous probes is correctly classified as auth_required (with an amber status indicator on the catalog tile) rather than a hard failure, so partially-protected servers register cleanly.

    MCP Store#

    A bilingual template registry is now built into the gateway. Templates ship with sensible defaults, the necessary OAuth scopes, and end-user-facing notes. The Linear OAuth template is seeded out of the box; more popular services follow. Display labels disambiguate multi-install templates so a personal GitHub install and a work GitHub install appear as distinct tiles instead of two indistinguishable entries.

    Step-by-step registration wizard#

    MCP server registration was reworked into a guided wizard with auth-mode-aware screens. The edit form refuses to save fields that don’t belong to the current auth mode, and tool-call / install errors are now surfaced in plain language with actionable next steps instead of raw JSON-RPC error codes.

    Three-tier upstream subject resolution#

    When an upstream MCP server uses OAuth, ThinkWatch resolves the upstream user identity through a three-tier strategy: parse the JWT if present, fall back to the OAuth userinfo endpoint, fall back again to issuer discovery. The resolved subject becomes the cache and audit key, so two users who share an MCP server stay strictly separated downstream.

    Per-credential Test Connection#

    The /connections page surfaces a per-credential Test Connection button. Verify your OAuth/PAT actually works against the upstream server before committing to it; the result is structured (auth_ok / auth_required / unreachable / tool_call_ok) rather than a single green/red dot, so you know exactly which step is broken.

    Security hardening#

    Following an internal security review, the MCP path now has SSRF protection on probe URLs, the response cache is scoped per (user, account_label) so OAuth/PAT data never leaks across users, rate limits are applied at every gateway hop, audit records are emitted on tool-call boundaries, and admin foot-gun guards prevent the most common misconfigurations (verifying static tokens at paste time, blocking obviously-wrong credential combinations).

April 2026

  1. ThinkWatch Enterprise v0.2.0

    Teams, spending controls, and a programmable API

    • Teams — isolate keys, budgets, and analytics per business unit
    • Rate limits & budget caps with fail-closed enforcement and spend alerts
    • RBAC v2 — custom roles, permission history, one-click import/export
    • Tokens moved to HttpOnly cookies — XSS can no longer steal sessions
    • Full management API with OpenAPI docs and API-key auth
    • Real-time dashboard powered by WebSocket

    Teams#

    Users and API keys can now be organized into teams. Each team gets its own isolated view of the analytics dashboard, its own user list, and its own scope when assigning API keys. A structured scope picker makes member assignment straightforward.

    Spending controls#

    A new limits engine enforces rate limits and budget caps at every layer — per user, per API key, per provider, and per MCP server. Limits are configured through a panel embedded directly in each edit dialog. When a budget is exhausted the gateway fails closed — no silent overruns. Budget alert thresholds let you set a warning before the hard cap is hit. Streaming requests are metered correctly; weighted token costs are supported for models with asymmetric pricing.

    RBAC v2#

    The role system has been completely reworked. Built-in and custom roles share a unified table. Custom roles are fully editable in a CodeMirror editor, support cloning from an existing role as a starting point, and track a full permission history so you can see who changed what and when. Roles can be exported and imported as JSON for environment parity. Role assignments are now scope-aware — a role granted at the team level applies only within that team.

    Security hardening#

    Access and refresh tokens have been migrated from localStorage to HttpOnly cookies — the JavaScript-visible session surface is now zero. The refresh endpoint binds each token to the originating client IP, so a stolen cookie cannot be replayed from a different network. The WebSocket dashboard connection now uses a one-shot ticket rather than a long-lived credential.

    Management API#

    The gateway now exposes a programmable management API authenticated by API key. A full OpenAPI specification is served alongside the gateway so you can integrate key lifecycle and provisioning into your own tooling or CI pipelines without touching the console.

  2. ThinkWatch Enterprise v0.1.0

    Public preview

    • Multi-format AI API gateway (OpenAI, Anthropic, Gemini, Bedrock)
    • MCP gateway with namespace isolation and tool-level RBAC
    • ClickHouse-powered audit logs with multi-channel forwarding
    • First-run setup wizard and built-in configuration guide

    ThinkWatch is now in public preview. This first tagged release ships the dual-port architecture (gateway :3000, console :3001), virtual API key lifecycle management, sliding-window rate limiting, circuit breakers, and the unified log explorer.

    The Helm chart and distroless container images are available alongside Docker Compose for self-hosted deployments.