Skip to content

Tags: ThinkWatchProject/ThinkWatch

Tags

v2.2.0

Toggle v2.2.0's commit message
ThinkWatch 2.2.0

This release fixes authorization. The gateways never checked
`ai_gateway:use` or `mcp_gateway:use`, so a user with no role, or with
only roles that grant neither, could call every model and MCP tool; a
role granted at the scope of one team widened model and tool access
platform-wide; and an API key whose allow-list came out empty could
call any model. The gateways now require the permission and count only
the roles that grant it. The seeded `team_manager` role now works when
granted at team scope, and seven permissions that nothing ever checked
are retired. Some users lose gateway access on upgrade: read the first
section before deploying. There is one manual database script and no
schema, setting, environment variable or Helm value change.

### Read before upgrading

- **Users without a role that grants gateway use lose gateway access.**
  A request to the AI gateway is refused with `403` unless a role held
  by the key's owner grants `ai_gateway:use`, and a request to the MCP
  gateway unless one grants `mcp_gateway:use`. The body names the
  missing permission, in the error format of the caller's API. The
  built-in `developer`, `admin`, `super_admin` and `team_manager` roles
  grant both; `viewer` grants neither. Three groups of users are
  refused after the upgrade where 2.1.0 let them through:
  - users with no role at all;
  - users whose roles are only `viewer`, or custom roles without these
    actions;
  - users whose only gateway role is assigned at team scope (see the
    third item below).

  Find them before upgrading. This query lists every active user with a
  live key for a gateway that none of their roles will grant (it counts
  global assignments and roles attached to the user's teams, which is
  what the gateways now read; it does not account for `Deny`
  statements):

  ```sql
  WITH grants AS (
    SELECT id AS role_id,
           jsonb_path_exists(policy_document,
             '$.Statement[*] ? (@.Effect == "Allow").Action[*] ? (@ == "*" || @ == "ai_gateway:*" || @ == "ai_gateway:use")') AS ai,
           jsonb_path_exists(policy_document,
             '$.Statement[*] ? (@.Effect == "Allow").Action[*] ? (@ == "*" || @ == "mcp_gateway:*" || @ == "mcp_gateway:use")') AS mcp
      FROM rbac_roles
  ),
  held AS (
    SELECT user_id, role_id FROM rbac_role_assignments WHERE scope_kind = 'global'
    UNION
    SELECT tm.user_id, tra.role_id
      FROM team_members tm JOIN team_role_assignments tra USING (team_id)
  ),
  access AS (
    SELECT h.user_id, bool_or(g.ai) AS ai, bool_or(g.mcp) AS mcp
      FROM held h JOIN grants g USING (role_id) GROUP BY h.user_id
  )
  SELECT u.email, s.surface
    FROM api_keys k
    JOIN users u ON u.id = k.user_id
    CROSS JOIN LATERAL unnest(k.surfaces) AS s(surface)
    LEFT JOIN access a ON a.user_id = u.id
   WHERE k.is_active AND k.deleted_at IS NULL
     AND u.is_active AND u.deleted_at IS NULL
     AND NOT COALESCE(CASE s.surface WHEN 'ai_gateway' THEN a.ai
                                     WHEN 'mcp_gateway' THEN a.mcp END, false)
   GROUP BY u.email, s.surface
   ORDER BY u.email, s.surface;
  ```

  Give each of them `developer` (or a custom role with the actions)
  either at global scope (Users → edit the user's roles, scope
  *Global*), or by attaching the role to a team they belong to (Teams →
  the team → Roles), which every member of that team inherits. The
  change takes effect within a minute; each user's permissions are
  cached for 60 seconds.

  If SSO users are meant to use the gateway from their first sign-in,
  set *Default Role for New Users* (`auth.default_role`) in Settings to
  `developer`. It is empty by default, and it applies only to accounts
  created after it is set; existing users need the grant above.
- **A role that does not grant gateway use no longer widens model or
  tool access.** Model and MCP tool scopes are now the union over the
  roles that grant `ai_gateway:use` / `mcp_gateway:use` only. A role
  without those actions (such as `viewer`) used to count as
  "unrestricted", so adding it to a user limited to some models opened
  every model to them. Users relying on that lose the extra models.
- **A role assigned at team scope no longer grants gateway access.** An
  assignment with scope `team:<id>` administers that team from the
  console (team roster, team limits); it no longer contributes models,
  MCP tools or gateway use, which it used to do for every request the
  user made, member of the team or not. Roles attached to a team itself
  (Teams → Roles), which every member inherits, still count. A user
  whose only gateway role was a team-scoped `developer` or
  `team_manager` needs that role at global scope, or attached to their
  team. The role assignment editor now says this under the scope
  picker.
- **An empty model allow-list allows nothing.** A key whose
  `allowed_models` is `[]`, or whose list shares no model with what its
  owner's roles grant, used to call any model; it now calls none, as
  `allowed_mcp_tools: []` already did on the MCP gateway. The console
  never saves `[]` (clearing the picker sends `null`), so only keys
  written through the API are affected. To find them:
  `SELECT id, name, user_id FROM api_keys WHERE allowed_models = '{}' AND is_active AND deleted_at IS NULL;`
  Set such a key's list to `null` to leave it bounded by its owner's
  roles only. Model entries still match by prefix, and a key narrowed to
  `gpt-4o-mini` under a role granting `gpt-4o` keeps `gpt-4o-mini`.
- **Run `db/release_migrations/2026-09-30_retire_unchecked_permissions.sql`
  once, after deploying.** The server does not apply it. It adds
  `teams:read` to the `team_manager` role and removes the seven retired
  permissions (see *Changed*) from every role, system and custom. It is
  idempotent and changes no access for any other role, since nothing
  checked the removed keys. Without it, the server still starts and
  works, but:
  - `team_manager` keeps the old `team:read` it cannot use, so a team
    manager who is not a member of the team they manage still gets
    `403` opening it, its roster or its roles;
  - every start logs a warning listing the roles that still name
    retired permissions.

  *Reset to defaults* on a system role in the console has the same
  effect for that one role, but does not clean custom roles.
- **A team-scoped `rate_limits:write` now takes effect.** It used to
  require global scope for every subject, so the seeded `team_manager`
  granted at team scope could not touch any limit. It now covers the
  rate limits and budget caps of users in that team and of the API keys
  they own, never the holder's own user or keys; role subjects still
  need global scope. Anyone holding `team_manager` (or a custom role
  with `rate_limits:write`) at team scope can now change, lift or
  delete the limits of every member of that team, administrators
  included. Review team-scoped assignments of these roles before
  upgrading.

### Security

- **Gateway use is checked, and a missing grant means no access.** See
  the first three items above: a user with no role, or only roles
  without gateway use, could call every model and MCP tool; a
  `viewer`-style role widened a restricted user to every model; a role
  granted at the scope of any team, even one the user was not a member
  of, widened model and tool access platform-wide. An explicit `Deny`
  on `ai_gateway:use` or `mcp_gateway:use` now closes that gateway.
  Action wildcards (`ai_gateway:*`, `*`) grant it, as they already did
  for console permissions.
- **A key narrowed inside a prefix grant is no longer unrestricted.**
  The key's allow-list was intersected with its owner's role grants
  entry by entry, as literal strings. A key limited to `gpt-4o-mini`
  under a role granting `gpt-4o` (a prefix) came out with an empty list,
  which the gateway read as "no restriction", so the key could call
  every model. The intersection now keeps an entry of either side that
  the other side covers, by prefix for models and by `<server>__*`
  pattern for MCP tools, and an empty result allows nothing.

### Fixed

- **The seeded `team_manager` role works at team scope.** It granted
  `team:read`, but the team handlers check `teams:read`, so a team
  manager got `403` listing teams and opening the team, its roster or
  its roles. It now grants `teams:read`, and opening a team accepts
  `teams:read` scoped to that team, where it used to require global
  scope. Existing installations need the release migration above.
- **Bulk disable and delete work on API keys' limits.** The bulk
  disable and delete routes for rate-limit rules and budget caps
  rejected every row stored for an API key (`api_key_lineage`, the kind
  a key's limits are stored under) as an unknown subject kind.
- **The console's effective-permissions preview shows real gateway
  access.** It now counts only roles that grant gateway use and skips
  team-scoped assignments, as the gateways do, where it used to count
  every assigned role.

### Changed

- **Seven permissions that nothing checked are retired**: `team:read`,
  `team:write`, `logs:read_own`, `logs:read_team`,
  `audit_logs:read_own`, `audit_logs:read_team` and
  `audit_logs:read_all`. Every log endpoint, audit logs included, is
  gated on `logs:read_all` at global scope, and no own- or team-filtered
  log view exists, so these grants never did anything. They are gone
  from the permission catalog, the role editor and the seeded roles.
  Roles that still name them load, and the server logs a warning at
  start instead of refusing to boot.
- **Deleting one rate-limit rule or budget cap is bound to the subject
  in the path.** `DELETE` on
  `/api/admin/limits/{kind}/{id}/rules/{rule_id}` and
  `/api/admin/limits/{kind}/{id}/budgets/{cap_id}` now answers `404` when the rule or cap belongs to a different
  subject; it used to delete any row id once the path's subject was
  authorized.
- **Creating or editing a key with an explicit allow-list checks it
  against gateway-granting roles only.** An `allowed_models` or
  `allowed_mcp_tools` entry is refused with `400` when no role of the
  owner that grants the gateway covers it, so a user without a gateway
  role can no longer save a non-empty list.

v2.1.0

Toggle v2.1.0's commit message
ThinkWatch 2.1.0

Amazon Bedrock becomes a provider you can run from the console. It
authenticates with a Bedrock API key as well as access keys or the
instance role, lists its models for import and for the route editor,
and has a working Test Connection. Requests converted for Bedrock now
keep their prompt-cache breakpoints and Claude's thinking settings. The
route editor takes any upstream model name, and a route the provider
refused can be created anyway. ThinkWatch-Core moves to v0.55.0, whose
Bedrock layer this edition now shares with the desktop gateway. No
database, setting, environment variable or Helm value changes.

### Read before upgrading

- **Test Connection with a saved provider's secrets needs
  `providers:update`.** A test that names a saved provider
  (`provider_id`, which is what the Edit dialog sends) used to take only
  `providers:create`. A custom role that has `providers:create` without
  `providers:update` can no longer test existing providers; the built-in
  admin role has both. A test with values typed into the request still
  takes `providers:create`. See *Security* below.
- **Refusing a route the provider does not serve answers
  `model_not_served`.** `POST /api/admin/models/{model_id}/routes` still
  answers `400` when the import probe found no API the provider serves
  the model on, but `error.type` is now `model_not_served`, where it was
  `bad_request`. Scripts that matched on `bad_request` for this case need
  the new value. The request takes a new optional `"force": true` to
  create the route anyway.
- **`error_type` in `gateway_logs` for failed requests is always the
  error's tag.** An upstream HTTP error used to log its debug text, such
  as `ProviderHttpError { status: 502, message: "…" }`, and an upstream
  rate limit a cut-off `UpstreamRateLimited { retry_after_secs: Some`.
  They now log `ProviderHttpError` and `UpstreamRateLimited`, as the
  metric labels and streamed requests already did. Dashboards or log
  forwarder queries that grouped by those strings see them merge into
  one value each.
- **The server log now carries the upstream's reply to a 401 or 403**,
  as it already did for other upstream errors. For Bedrock that reply
  names the AWS account and the IAM principal. Callers still see only
  "Authentication failed with upstream", and failover, breakers and
  `gateway_logs` treat these errors as before.
- **Prompt caching now works on Bedrock routes, and is billed at cache
  prices.** Requests converted for Bedrock used to lose their
  `cache_control` breakpoints, so Claude on Bedrock re-read the whole
  prompt at full price every turn. Breakpoints now go out as Converse
  `cachePoint` blocks, to Claude and Nova models only (others reject
  them), and the cache reads and writes Bedrock reports come back into
  usage, one-hour writes included. Traffic with repeated prompts, such as
  Claude Code, costs much less on Bedrock routes than it did; writes cost
  a little more, at `cache_write_weight` / `cache_write_1h_weight`.
- **Claude's thinking reaches Bedrock.** A request converted for Claude
  on Bedrock now carries its thinking and effort settings, in the form
  Anthropic's API uses, along with the sampling limits thinking imposes
  (`temperature` 1, no `top_k`, `top_p` at least 0.95). Thinking used to
  be dropped on the way to Bedrock. Other Bedrock models still get no
  thinking.

### Added

- **Bedrock providers can authenticate with a Bedrock API key.** Pick
  *Bedrock API Key* as the authentication mode and paste the key AWS
  generated. It is kept like any other provider's API key, as an
  encrypted `Authorization: Bearer` header, and replaced in the
  provider's Edit dialog. A Bedrock provider that sends an
  `Authorization` header is not SigV4-signed; one without it is signed
  as before, with access keys or the instance role. Use a long-term key:
  a short-term one expires within 12 hours.
- **Bedrock models can be imported, and are suggested in the route
  editor.** A Bedrock provider now lists what it can be routed to: the
  region's foundation models that can be invoked on demand and answer in
  text, and the inference profiles AWS defines, such as
  `us.anthropic.claude-sonnet-4-5-20250929-v1:0`. A model that is only
  served through an inference profile, as most current models are, is
  listed under its profiles' ids and not its own. Whether the account
  may call a model is still checked one model at a time when it is
  imported. The provider's credential needs
  `bedrock:ListFoundationModels` and `bedrock:ListInferenceProfiles`;
  the `AmazonBedrockLimitedAccess` policy a long-term API key is created
  with allows both.
- **Test Connection for Bedrock providers**, however they authenticate:
  with an API key, access keys or the instance role. It lists the models
  above, so a wrong credential or a missing permission shows up before
  any traffic does.
- **Create anyway, for a route the provider refused.** When the route
  editor's save is refused because the provider does not serve the
  model, the error now offers *Create anyway*: the refusal can be
  stale, or be about the probe's request rather than the model. The
  route is created, and its `model_route.created` audit row records the
  refusal it overrode in a new `refusal_overridden` key.

### Changed

- **The route editor's upstream model field takes any model name.** For
  a provider that lists its models, the field used to turn into a list
  to pick from, so a model the listing leaves out could not be routed to
  from the console. It now suggests the provider's models as you type,
  and takes whatever is typed.
- **Bedrock instance-role credentials are cached.** They took three
  IMDSv2 round trips per request; they are now kept until five minutes
  before they expire, and a burst of requests at expiry makes one trip.
- **Chat Completions requests can switch reasoning with `thinking`.**
  When a request is converted for another format, `thinking.type`
  (`enabled` / `disabled`, as DeepSeek, GLM and Kimi write it) now turns
  reasoning on or off; `disabled` wins over `reasoning_effort`.
- **Core crates at ThinkWatch-Core v0.55.0.** `tw-dialect`, `tw-guard`
  and `tw-breaker` move from v0.43.0, and `tw-bedrock` joins them:
  SigV4 signing, eventstream unframing, Bedrock's addresses, region
  checks and model catalog now come from core. Where credentials come
  from (provider keys, the instance role) stays in this repository.
  `tw-breaker` is unchanged.

### Fixed

- **Bedrock providers could not be created or edited in the console.**
  The region was checked as a URL, so saving failed with
  `400 Invalid URL`. A Bedrock provider's `base_url` is now checked as an
  AWS region such as `us-east-1`, and anything else is refused, since the
  host is built from it. A provider saved with a URL there never reached
  Bedrock; set its region in the Edit dialog.
- **Bedrock models the account may not call were imported anyway.**
  Bedrock refuses such a model (model access not granted, or an IAM or
  organization policy that denies it) with a 403, and the import probe
  read every 403 as a credential problem that says nothing about the
  model. The route was created and failed on first use. Now, when the
  region's control plane accepts the same credential, the refusal is
  recorded as the model's, with AWS's reason, and the model is skipped
  on import like any other refused one. Once access is granted,
  re-check the provider's models.
- **A Bedrock provider whose stored secret key would not decrypt signed
  with an empty secret**, and every request failed with AWS's
  `SignatureDoesNotMatch`. It now refuses its requests with the reason
  until the keys are saved again, rather than falling back to the
  instance role, which would call AWS as a different identity. An
  access key ID saved without a secret is refused the same way.
- **A Bedrock model id that is an ARN could not be routed.** Its `/`
  went into the request path unescaped, adding a path segment AWS could
  not route. It is now escaped.
- **A base URL ending in its version segment doubled it.** A provider
  written as `https://api.openai.com/v1` sent requests, and the model
  listing, to `/v1/v1/…` and got 404s. A trailing version segment
  (`v1`, `v1beta`, …) that the request path starts with is now written
  once. Base URLs without one are unchanged.
- **Typing into a provider's API key field and clearing it again wiped
  the saved key on save.** The field sent `Bearer ` with nothing after
  it. A cleared field now keeps the saved key, as a blank one always did.

### Security

- **Testing a connection with a saved provider's secrets takes
  `providers:update`.** The test fills each header left blank with the
  saved provider's secret and sends it to the URL in the request, and
  `providers:create` was enough to ask for it. So a user allowed only to
  create providers could send any saved API key to a server of their
  own. A test that names a saved provider (`provider_id`) now takes
  `providers:update`, which lets a user point that provider elsewhere
  anyway; a test without one still takes `providers:create`. A user with
  `providers:update` alone can now use Test Connection in the Edit
  dialog, which used to answer `403`.

v2.0.0

Toggle v2.0.0's commit message
ThinkWatch 2.0.0

Callers now get errors in their own API's format, and an upstream that
refuses a request no longer takes a model's other routes down with it.
The gateway also speaks two more client protocols: Gemini, and the
Responses API over a WebSocket. Cached input is billed at cache prices,
and a request with no usage report is billed on an estimate instead of
at zero. The TOTP requirement, which never took effect before, is now
enforced. This is a major release because error bodies, the
content-filter preset ids and the TOTP behaviour all change in ways a
client or a script can notice.

### Read before upgrading

- **Check `security.totp_required` before you upgrade.** In 1.x this
  setting never took effect: it is stored as a boolean and was read as a
  string, so it always read as off. From 2.0.0 it is enforced by the
  server. Find out what it is set to:

  ```sql
  SELECT value FROM system_settings WHERE key = 'security.totp_required';
  ```

  If it is `true`, every console user without TOTP, super admins
  included, is held at a TOTP setup screen on their next request, and
  sessions that are already open are held too. Until they set up TOTP,
  every console and admin endpoint answers 403
  `totp_enrollment_required`, except `/api/auth/me`, logout,
  `register-key` and the TOTP status/setup/verify-setup calls. Setting
  up TOTP releases the session straight away. API keys are not affected:
  gateway, MCP and console `tw-` key traffic keeps working. While the
  setting is on, `POST /api/auth/totp/disable` is refused with 400. The
  setting must now be a JSON boolean; a string such as `"true"` is
  refused on save.
- **Gateway error bodies follow each client API's own format.** Status
  codes and `Retry-After` are unchanged. Code that reads the error
  `type` needs updating:
  - **Chat Completions and Responses** keep the
    `{"error": {"message", "type", …}}` shape, but `type` is now
    OpenAI's value for the status, not a ThinkWatch tag:
    `authentication_error` (401), `permission_error` (403),
    `not_found_error` (404), `rate_limit_error` (429),
    `invalid_request_error` (other 4xx) and `server_error` (5xx). The
    old tags are gone: `rate_limited`, `policy_blocked`,
    `provider_http_error`, `provider_error`, `provider_timeout`,
    `transform_error`, `network_error` and `auth_error`. A policy block
    is now `permission_error` with status 403.
  - **Anthropic Messages** clients get Anthropic's body,
    `{"type": "error", "error": {"type", "message"}}`, with Anthropic's
    type names (`rate_limit_error`, `overloaded_error`, `api_error`, …).
  - **Gemini** clients get Google's body,
    `{"error": {"code", "message", "status"}}`.
  - **Once a stream has started**, a Responses client gets a
    `response.failed` event, where before it got a Chat-style error
    frame that SDKs skip, so the stream just stopped. An Anthropic
    client gets an `error` event whose type follows the status.
  - The `error_type` field in `gateway_logs` and the metric labels are
    unchanged.
- **Only upstream failures fail over or count against a route's
  breaker.** In 1.x every non-2xx except 401, 403 and 429 was retried on
  the model's other routes and counted as a failure on each of them, so
  one malformed request could open the breakers on all of a model's
  routes.
  - **Tried on another route and counted against this one:** 5xx, 408,
    429, 401 and 403 (the upstream refused the gateway's own
    credential), timeouts, network errors and unreadable responses.
  - **Returned to the caller straight away, and counted as the upstream
    working:** every other 4xx.
  - **What the caller sees:** such a 4xx comes back with its own status
    and the upstream's reason. In 1.x it came back as a 502 after every
    route had been tried.
  - An upstream 5xx comes back with the upstream's status (500, 503, …)
    rather than a blanket 502.
  - An upstream timeout is now 504.
  - Streams follow the same rule. A stream cut by tool-call inspection
    no longer counts against the route.
- **Cached input is billed at cache prices.** In 1.x cache reads and
  writes were billed, and debited from budgets and weighted rate limits,
  as full-price input. `models` gains three weights, `cache_read_weight`,
  `cache_write_weight` and `cache_write_1h_weight`. When a weight is
  unset, it is `input_weight` times Anthropic's ratio: 0.1× for a read,
  1.25× for a write and 2× for a one-hour write. What this changes:
  - Traffic with many cache reads (Claude Code, for instance) costs much
    less than it did.
  - Traffic that writes to the cache costs a little more.
  - Older OpenAI models discount cache reads less (0.5× or 0.25×). Set
    the weights on those models yourself.
  - `input_tokens` in the log is still the whole input. The log detail
    gains `cache_read_tokens`, `cache_write_tokens` and `cache_write_1h`.
- **Output length limits now apply to streams.** In 1.x, `max_length`
  output guardrails checked only whole responses, so streamed answers
  were never checked. Now the frame that would cross the limit is not
  sent, and the stream ends with an error in the caller's format. A
  response served from the cache is also checked against the limit in
  force. If you set a limit, streamed answers that used to go through
  can now be cut off.
- **Content-filter preset groups are renamed.** The groups are now
  `injection`, `persona` and `chinese`; they used to be `basic`,
  `strict` and `chinese`. This matters only if you call the presets
  endpoint by group id. Rules you have already added are copies and are
  not affected. Other changes to the filter:
  - A rule with an empty pattern is now refused on save.
  - Each text part of a message is scanned separately, so a pattern no
    longer matches across two parts.
  - The engine is now shared with ThinkWatch-Core's `tw-guard`. The
    stored format and the admin API are unchanged.
- **Requests with no usage report are billed on an estimate.** In 1.x
  such a request was billed at zero. This happens when an upstream
  ignores the request for usage, or when the caller leaves before the
  final chunk arrives. The estimate is:
  - input: about four bytes of the request per token, not counting
    images and files;
  - output: the answer that actually arrived.

  Estimated rows carry `usage_estimated: true` in their detail and count
  in `gateway_usage_estimated_total`. A request with no answer at all is
  still billed at zero.
- **Clients that leave early are now logged.** In 1.x a client that
  disconnected before its response existed left no `gateway_logs` row at
  all. That covers leaving during auth, limits or routing, or while
  waiting for a whole (not streamed) answer. Such a request now writes
  one row: status 499, `stream_outcome: client_cancelled`,
  `cancelled_before: response`, no tokens and no cost. Expect more 499
  rows in dashboards and log forwarders. A new counter,
  `gateway_cancelled_before_response_total`, counts them.

### Database changes

Both apply on their own at startup, as every schema change does, and
both are additive:

- `models` gains three nullable columns: `cache_read_weight`,
  `cache_write_weight` and `cache_write_1h_weight`
  (`ALTER TABLE … ADD COLUMN IF NOT EXISTS`, `CHECK (>= 0)`).
- `system_settings` gets an `auth.default_role` row, seeded empty (no
  role). Existing rows are left alone (`ON CONFLICT DO NOTHING`).

Neither is irreversible. A 1.1.0 server runs against the upgraded
database: it ignores the new columns and the new setting. What a 1.1.0
server cannot do is price cache tokens from the weights.

### Added

- **Gemini clients.** New endpoints:
  - `POST /v1beta/models/{model}:generateContent` and
    `:streamGenerateContent`, also served under `/v1/models/…`;
  - `GET /v1beta/models`, which lists models in Gemini's format.

  These requests get the same limits, budgets, filters, routing with
  failover, format conversion, inspection, billing and audit as every
  other endpoint. A Gemini upstream gets the request as it was sent.
  A stream comes back as SSE with `alt=sse`, and as Gemini's JSON array
  without it. `:countTokens` and `:embedContent` are refused with 400.
- **The Responses API over a WebSocket.** Connect to `GET /v1/responses`
  with `Upgrade: websocket`.
  - Each `response.create` frame is handled like a streamed
    `POST /v1/responses`, with its own limits, routing, billing and
    audit row.
  - Turns on one connection run in order. A refused turn fails with
    `response.failed`, and the connection stays open.
  - The connection keeps its latest response. That lets a turn continue
    from it with `previous_response_id`, even with `store: false`,
    which is how Codex works, and against any upstream format.
  - A new counter, `gateway_responses_ws_connections_total`, counts
    connections.
- **More places to put an API key.** Gateway keys are also accepted in
  `x-api-key` (Anthropic SDKs), `x-goog-api-key` and `?key=` (Gemini
  SDKs), as well as `Authorization: Bearer`. Headers are checked first.
  A key given in the query string is never sent upstream.
- **Hidden-text audit events show what the text says.** Each item in
  `found` gains `revealed`, the ASCII that the hidden tag characters
  spell.
- **`auth.default_role` can be set.** It is the role that newly
  registered users and SSO users get. In 1.x, setting it through the
  admin API reported success but changed nothing, because the setting
  row did not exist.
- **Model editor** fields for the three cache weights. Each placeholder
  shows the value used when the field is left empty.

### Changed

- **Default output length for upstreams that require `max_tokens`.** When
  the caller sets none, the gateway now sends 32000 for Claude models and
  8192 for other models. It used to send 4096, which cut Claude answers
  short.
- **More upstreams count as the vendor's own endpoint.** DeepSeek,
  Moonshot, Zhipu/Z.ai, DashScope, xAI and `*.amazonaws.com` are now
  recognised, and the check reads the parsed host. A relay URL such as
  `https://relay/api.openai.com` no longer passes as official. Official
  endpoints are stricter about request parameters, so the gateway drops
  or renames some parameters before sending to them.
- **Hidden-text scanning uses ThinkWatch-Core's `tw-guard`.** Same
  scope, same actions, and nothing is stripped.
- **Requests forwarded in their own format lose ThinkWatch's reasoning
  signatures.** A `tw1.` signature written by an earlier format
  conversion is removed, because Anthropic rejects it. The upstream's
  own signatures are kept.
- **Core crates: `tw-dialect`, `tw-guard` and `tw-breaker` at
  ThinkWatch-Core v0.43.0.** The code only this edition used (at-rest
  crypto, SigV4 signing, the gateway error type) moved into this
  repository. It works the same, and stored secrets decrypt as before.
- **The server's SQL moved from the request handlers into repository
  modules** (catalog, dashboard, limits, log forwarding, identity,
  access, MCP). Every statement is unchanged. New integration tests cover
  these endpoints and pass on both the old and the new code.
- **CI runs on pull requests into `dev`**, including the whole
  integration suite against Postgres, Redis and ClickHouse.

### Fixed

- **Revoking a user's default MCP connection always failed with a 500**,
  and the account could not be revoked. The newest remaining account is
  now made the default.

v1.1.0

Toggle v1.1.0's commit message
ThinkWatch 1.1.0

The gateway stops rebuilding every request as a chat-shaped message. A
request whose route speaks the caller's own format goes out as the
caller sent it; one that crosses formats is converted by
[ThinkWatch-Core](https://github.com/ThinkWatchProject/ThinkWatch-Core),
the same layer the desktop edition uses. Tools, tool choice, system
prompt blocks, `metadata` and `cache_control` now reach the upstream,
where they used to be dropped. Two new checks guard what goes in and
out: tool calls an upstream returns, and invisible characters in what
a caller sends.

### Read before upgrading

- **Anthropic routes record more prompt tokens for the same work.**
  Prompt tokens now count the same for every upstream: plain input plus
  cache reads and cache writes, which is OpenAI's definition.
  Anthropic's own `input_tokens` leaves the cached part out, so on those
  routes prompt tokens, cost and budget use all go up. The price model
  still charges every prompt token at one rate.
- **Requests in the upstream's own format are forwarded as sent.** Every
  field the caller sends reaches the upstream, along with the caller's
  `anthropic-beta` and `anthropic-version` headers. Only the model name
  changes, and PII is swapped for placeholders. An OpenAI-compatible
  upstream that rejects fields it does not know may now refuse requests
  that used to succeed, because those fields were stripped before. Try
  your upstreams with the clients you actually run.
- **Two checks are on by default, and neither blocks anything yet.**
  - Tool-call inspection starts in `observe` mode.
  - The hidden-character check starts in `warn` mode.
  - Both write audit events, so expect new entries in the audit log and
    in anything subscribed to it.
  - Nothing is refused until you switch to `enforce` or `block`.
- **The response cache starts cold.** The cache key now covers the whole
  request, so entries written by 1.0.2 are never hit again. They expire
  on their own.
- **One PII value gets one placeholder.** Within a request, the same
  e-mail address is `{{EMAIL_1}}` wherever it appears. It used to get a
  new number each time, so a model saw one person as several. Saving a
  PII pattern now also requires the placeholder prefix to be letters,
  digits or underscores.

### Added

- **Tool-call inspection** (`security.tool_inspection`).
  - **Why:** an upstream writes the response, so it can hand the caller
    a tool call the model never made, such as `bash("curl … | sh")`
    appended to an ordinary answer. An agent set to auto-approve then
    runs it.
  - **Rules:** a built-in set of dangerous-command rules. Each can be
    switched off or given a different action, and you can add your own.
  - **Modes:** `off`, `observe` (records hits and changes nothing on the
    wire) and `enforce`.
  - **Enforce on a stream:** the stream is cut at the frame that would
    complete a matching call, and the refusal arrives in the caller's
    format.
  - **Enforce on a whole response or a cache hit:** refused with 403
    (`policy_blocked`). A refused answer is neither cached nor billed.
  - **Audit and metrics:** every hit is an audit event
    (`gateway.tool_call_flagged` or `gateway.tool_call_blocked`) and
    counts in `gateway_tool_call_flagged_total`.
  - **Admin API:** `GET /api/admin/settings/tool-inspection/rules` and
    `POST /api/admin/settings/tool-inspection/test`.
  - **Console:** a card on the security page, plus a sandbox tab.
- **Hidden-character check** (`security.hidden_text`: `off`, `log`,
  `warn` or `block`).
  - **What it looks for:**
    - Unicode tag characters, which carry an instruction invisibly into
      the model's context;
    - bidirectional overrides, which make text read differently on
      screen than it is.
  - **Where:** the caller's messages and the tool results inside them.
    The system prompt and the model's own turns are not checked.
  - **Not flagged:** zero-width joiners (emoji), the zero-width
    non-joiner (Persian) and Cyrillic.
  - **Actions:** `warn` writes `gateway.hidden_text_flagged`; `block`
    refuses with 403 and writes `gateway.hidden_text_blocked`.
  - **Console:** a card on the security page.
- **Tool-call arguments get their PII back.** A model asked to e-mail
  `a@example.com` used to call the tool with `{{EMAIL_1}}` as the
  address.

### Changed

- **One pipeline for `/v1/chat/completions`, `/v1/messages` and
  `/v1/responses`.** A same-format request is forwarded as sent. A
  cross-format request is converted, and the gateway logs what the
  target format cannot carry.
- **Content filtering and PII detection read tool results too.** That is
  where an injected instruction, or customer data pulled in by a tool,
  usually sits.
- **Usage is read off the upstream's own bytes.** A streamed response is
  no longer held in memory for an accounting pass at the end.
- **Chat streams are billed on the upstream's actual usage.** They are
  always sent asking for it. A caller who did not ask for the usage
  chunk still does not get one. 1.0.2 estimated the count for these.
- **Streams send their headers at once.** A caller who leaves while the
  upstream is still thinking is recorded as cancelled.
- **Connectivity tests use the live encoder.** A route's test request is
  built by the same encoder as real traffic, so a passing test means
  forwarding works.
- **The web console loads data through TanStack Query.**
  - Signing out, including from another tab, clears everything cached.
  - After a change, screens refresh in place.
  - Polling pauses while the tab is hidden.
- **Core crates come from one pinned tag** (ThinkWatch-Core v0.40.0),
  declared once at the workspace root.

### Fixed

- **Requests lost their tools, tool choice and non-text content** on the
  way upstream (ThinkWatch-Core#50). Claude Code's system prompt, sent
  as an array, was dropped whole, and so was every `cache_control`
  breakpoint. Each cached prefix was billed as full-price input.
- **The response cache could serve the wrong answer.** Its key covered
  only model, messages and `max_tokens`, so two requests that differed
  only in tools, `response_format`, `seed` and so on shared one entry.
- **A tripped route stayed out until its Redis key expired**, roughly
  four cooldowns. It now gets probed once the cooldown is over, and a
  success closes it.
- **The dashboard showed every AI provider's breaker as `Closed`.** It
  now shows the real state of the provider's routes, reporting the
  worst one.
- **The PII "try patterns" endpoint misreported labels.** A pattern
  named `CUSTOM_EMAIL` was reported as `CUSTOM`.
- **`:latest` could point at a `main` build rather than the release**,
  which is what happened for v1.0.2's server image. Only the release
  workflow sets `:latest` now.
- **The web image hung for six hours.** Its frontend was built under
  QEMU for arm64, where Node crashed and the step never returned. It is
  now built once, natively. The image contents are unchanged.

### Security

- Refreshed the web console's lockfile to clear 56 Dependabot alerts
  (1 critical, 23 high). All were transitive, and none of them reached
  the shipped bundle.

v1.0.2

Toggle v1.0.2's commit message
ThinkWatch 1.0.2

The shared gateway layer moves out into its own repository, and OIDC
learns to accept an ID token whose `aud` carries more than the client
ID. Mostly a fix release otherwise — the deploy and auth items below are
the ones worth reading before you upgrade.

- **`OIDC_ADDITIONAL_TRUSTED_AUDIENCES`** — a comma-separated allowlist
  of extra ID-token audiences to trust. Some IdPs put something besides
  the client ID in `aud`; Zitadel includes the parent project ID, and
  `openidconnect`'s stock verifier rejects every non-client-ID audience,
  so login failed outright. Leave it unset and verification is
  byte-for-byte what it was. Set, it widens exactly one check: the extra
  audiences on a multi-audience token are matched against this list as
  exact strings — no wildcards, no prefixes. The client ID must still
  appear in `aud` regardless. Documented in `.env.example`. Thanks to
  @DaniW42 for the report and the patch (#12).
- **Automatic upstream protocol detection** — the gateway learns an
  upstream's dialect and relearns it when the upstream changes, instead
  of relying on a static guess. Protocol is now a property of the route.
- **Reverse-proxy identification** — the server recognizes its own
  reverse proxy, which is what makes the dashboard WebSocket and the API
  docs work behind nginx.

- **`azp` is validated whenever it is present**, per OIDC Core 3.1.3.7
  step 5, which has no audience-count precondition. Previously it was
  only checked when a multi-audience token made it mandatory, so a
  single-audience token naming this client with `azp` pointing at a
  *different* client was accepted — the IdP stating plainly that the
  token was authorized for somebody else. **This can reject a token that
  1.0.1 accepted.** If your IdP sets `azp` to something other than your
  client ID, logins will start failing and the message will say so;
  that is the IdP to fix, not this setting.
- **The shared gateway layer now comes from
  [ThinkWatch-Core](https://github.com/ThinkWatchProject/ThinkWatch-Core)**
  as a git dependency rather than living in this tree. Six crates —
  types, protocol, provider, resilience, crypto, and the JSON secret
  envelope — are one implementation shared with the desktop edition
  instead of two copies drifting apart. No runtime or API change; it
  affects you only if you build from source, where the build now needs
  network access to resolve that dependency. The published images are
  unaffected.
- **Stored secrets are redacted on read and decrypted on use.** Provider
  credentials no longer travel in plaintext through code paths that
  merely display or list them.

- **Deploy** — nginx broke the API docs three separate ways; the
  dashboard WebSocket wasn't proxied; the prod stack's healthchecks were
  wrong; and the dev stack could overwrite production's ClickHouse
  credentials. The last one is the reason to read this list.
- **Auth** — the console logged itself out after every token refresh.
  Rate-limit counters could be created without an expiry and sit in
  Redis forever.
- **Models** — the gateway no longer offers models the upstream refuses,
  no longer imports models it won't serve, and reports every dialect
  that was refused rather than only the last. Provider `base_url`
  trailing slashes are normalized.
- **Analytics** — spend is attributed to a person, not a UUID.
- **Audit** — a syslog forwarder formatting fix that had CI red on
  clippy 1.98 (#11, thanks @DaniW42). Two more instances of the same
  lint, plus a `result_large_err` false positive in the MCP lifecycle
  stages, measured rather than boxed: the `Ok` variant is 296 bytes
  against the `Err`'s 152, so boxing saves nothing and adds an
  allocation.
- **UI** — the brand mark stays inside the collapsed sidebar rail.

`CONTRIBUTING.md` and a PR template now state the branch contract that
had only lived in `docs/operations/release.md`: routine work targets
`dev`, and `main` is the release line. A workflow comments on PRs opened
against `main` from anything other than `dev` or a `hotfix/*` branch,
because GitHub pre-fills the base with the default branch and walks
contributors into it.

- _(nothing yet)_

- _(nothing yet)_

- _(nothing yet)_

- _(nothing yet)_

- _(nothing yet)_

v1.0.1

Toggle v1.0.1's commit message
ThinkWatch 1.0.1 — release-pipeline validation

No product change vs 1.0.0. This tag exists to exercise the
release workflow's native arm64 + Node 24 path against a real
release, since the v1.0.0 build ran under the old QEMU path
and the tag itself is immutable.

See CHANGELOG.md `[1.0.1]` section.

v1.0.0

Toggle v1.0.0's commit message
ThinkWatch 1.0.0 — stability commitment release

No code delta vs 0.5.0. This tag is the SemVer commitment point:
every breaking change to the public surface (REST routes, MCP wire
shapes, audit-row JSON keys, DB schema, public Rust APIs) requires
a major bump from here on.

Operators:
- Pin chart values to `image.tag: "1.0.0"` for reproducible installs.
- The `:latest` tag now floats forward; bind it only with a tested
  rollback path.
- Read docs/operations/{secret-rotation,backup-restore}.md before
  first production deploy.

See CHANGELOG.md `[1.0.0]` and `[0.5.0]` sections for the full
product surface description.

v0.5.0

Toggle v0.5.0's commit message
ThinkWatch 0.5.0 — first public beta

Product surface stable enough to deploy against; API contract is NOT
committed (SemVer kicks in at 1.0.0). See CHANGELOG.md for the full
release notes, scope, and known limitations.

Highlights:
- Multi-provider AI gateway (OpenAI / Anthropic / Google / Azure /
  Bedrock) with failover, circuit breakers, semantic cache,
  streaming + PII restore.
- MCP gateway with per-user OAuth + admin-shared credentials, cache
  with prefix invalidation, UUID-keyed per-server breakers.
- Audit pipeline with PII-redacted bodies, ClickHouse for hot rows,
  S3 offload for oversize payloads (bucket lifecycle administered
  from the admin UI).
- RBAC + identity: JWT, refresh, OIDC SSO, TOTP, recovery codes,
  per-API-key limits + budgets.
- Observability: /metrics (bearer-protected) + opt-in ServiceMonitor,
  /api/health deep-probing PG / Redis / ClickHouse / S3.
- 675 tests, five rounds of systematic bug audits prior to tag.

Operators: read docs/operations/{secret-rotation,backup-restore}.md
before first production deploy.