ContextEditingMiddleware(
self,
*,
edits: Iterable[ContextEdit] | None = None,
token_count_method: Stream transformer factories registered by the middleware.
Logic to run before the agent execution starts.
Async logic to run before the agent execution starts.
Logic to run before the model is called.
Async logic to run before the model is called.
Logic to run after the model is called.
Async logic to run after the model is called.
Logic to run after the agent execution completes.
Async logic to run after the agent execution completes.
Intercept tool execution for retries, monitoring, or modification.
Intercept and control async tool execution via handler callback.
| Name | Type | Description |
|---|---|---|
edits | Iterable[ContextEdit] | None | Default: NoneSequence of edit strategies to apply. Defaults to a single |
token_count_method | Literal['approximate', 'model'] | Default: 'approximate'Whether to use approximate token counting (faster, less accurate) or exact counting implemented by the chat model (potentially slower, more accurate). Ignored when |
token_counter | TokenCounter | None | Default: None |
| Name | Type |
|---|---|
| edits | Iterable[ContextEdit] | None |
| token_count_method | Literal['approximate', 'model'] |
| token_counter | TokenCounter | None |
Automatically prune tool results to manage context size.
The middleware applies a sequence of edits when the total input token count exceeds configured thresholds.
Currently the ClearToolUsesEdit strategy is supported, aligning with Anthropic's
clear_tool_uses_20250919 behavior (read more).
Optional custom function counting tokens for a
sequence of messages. Takes precedence over
token_count_method when provided, mirroring
SummarizationMiddleware. Useful when the built-in
approximation is inaccurate for a workload (e.g. CJK-heavy
conversations) or when a provider-specific tokenizer should
be used without implementing
get_num_tokens_from_messages.