Skip to content

[WIP] Router Hint initiated P2P KV Cache Transfer Between TRT-LLM Workers - #18158

Draft
oandreeva-nv wants to merge 5 commits into
NVIDIA:mainfrom
oandreeva-nv:kv-cache-transfer-request-lifecycle
Draft

oandreeva-nv wants to merge 5 commits into
NVIDIA:mainfrom
oandreeva-nv:kv-cache-transfer-request-lifecycle

Conversation

@oandreeva-nv

Copy link
Copy Markdown

@coderabbitai summary

Description

This is a WIP implementation for RFC#18151

Implemented / validated locally

  • DP=1 G1->G1 full-prefix transfer

    • Source has full prompt KV.
    • Destination has no useful local prefix.
    • Destination pulls all matching blocks before prefill.
  • DP=1 G1->G1 local-prefix + remote-tail transfer

    • Destination already has a local prefix.
    • Source has the longer prefix/full prompt.
    • Destination skips local full blocks and receives the missing tail.
  • Overlapping partial local leaf handling

    • Destination has a partial local leaf.
    • Remote transfer needs that same logical block.
    • Destination allocates a fresh private receive block instead of overwriting the partial leaf.
  • Strict/all-or-fallback behavior

    • If source cannot satisfy requested blocks, transfer fails.
    • Request falls back to normal local prefill.
  • Async source metadata fetch

    • Hint endpoint fetch does not block admission path.
    • Request remains active but unschedulable while source state is resolving.
  • Context-prefetch request lifecycle

    • CONTEXT_INIT -> CONTEXT_PREFETCH_INIT -> CONTEXT_PREFETCH_IN_PROGRESS -> CONTEXT_PREFETCH_COMPLETE -> CONTEXT_INIT.
  • Fallback cleanup

    • Failed metadata fetch or failed receive returns request to CONTEXT_INIT.

WIP

  • DP>1 / attention-DP support

    • Need to verify matching source/destination DP-rank behavior.
    • Need explicit test where source KV is rank-specific.
  • Endpoint resolution

    • We can fetch source state from a URL.
    • Need final contract for exact URL vs host-only + default path.
    • Need native TRT-LLM serve and Dynamo endpoint compatibility decision.
  • Full-prefix routed through Dynamo

    • We validated with local Dynamo-style two-worker smoke.
    • Need decide whether this becomes a real integration test or remains local validation.

TODO [will be split into separate prs]

  • block_hashes validation

    • Router-provided hashes should validate source match/alignment.
    • Useful for stale source cache and shifted range detection.
  • Retention policy for received blocks

    • Ensure fetched blocks are not evicted before request reclaims/reuses them.
  • BestEffort policy

    • Transfer available prefix blocks.
    • Prefill missing suffix.
    • Report transferred vs skipped blocks.
  • Timeout policy

    • Explicit metadata fetch timeout and transfer timeout semantics.
    • Decide whether timeout is always fallback or can surface an error.
  • Security / SSRF handling

    • Validate allowed schemes/hosts.
    • Decide whether allowlisting belongs in TRT-LLM, serving, or orchestrator.
  • V2 KVCM / Python transceiver support

    • Current path is C++ CacheTransceiver / V1-style KVCM.
    • Python transceiver currently falls back.
  • Observability

    • Metrics/logs for hint accepted, ignored, fetch failed, transfer failed, transferred blocks, fallback reason.

Test Coverage

PR Checklist

Please review the following before submitting your PR:

  • PR description clearly explains what and why. If using CodeRabbit's summary, please make sure it makes sense.

  • PR Follows TRT-LLM CODING GUIDELINES to the best of your knowledge.

  • Test cases are provided for new code paths (see test instructions)

  • If PR introduces API changes, an appropriate PR label is added - either api-compatible or api-breaking. For api-breaking, include BREAKING in the PR title.

  • Any new dependencies have been scanned for license and vulnerabilities

  • CODEOWNERS updated if ownership changes

  • Documentation updated as needed

  • Update tava architecture diagram if there is a significant design change in PR.

  • The reviewers assigned automatically/manually are appropriate for the PR.

  • Please check this after reviewing the above items as appropriate for this PR.

GitHub Bot Help

To see a list of available CI bot commands, please comment /bot help.

Signed-off-by: Olga Andreeva <oandreeva@nvidia.com>
Signed-off-by: Olga Andreeva <oandreeva@nvidia.com>
Signed-off-by: Olga Andreeva <oandreeva@nvidia.com>
Signed-off-by: Olga Andreeva <oandreeva@nvidia.com>
Signed-off-by: Olga Andreeva <oandreeva@nvidia.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

1 participant