[Core] Preserve Marconi caching with selective hybrid cache retention - #47782
Merged
njhill merged 10 commits intoJul 13, 2026
Merged
Conversation
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
njhill
requested review from
ApostaC,
WoosukKwon,
alexm-redhat,
heheda12345,
ivanium,
orozery,
robertgshaw2-redhat and
ywang96
as code owners
July 6, 2026 22:25
ivanium
reviewed
Jul 6, 2026
ivanium
left a comment
Contributor
There was a problem hiding this comment.
Great idea! Overall makes sense. Left some comments.
A few high level thoughts:
- how does this work with chunked prefill? I feel somewhere we need to check if
request.shared_prefix_boundaryfalls in the current scheduled chunk. - We also need to update the mooncake store connector. Shouldn't be huge changes but maybe we can do it in a followup PR too
- As commented in scheduler, I found the previous handling of hybrid linear model is a bit hacky and silently breaks the marconi-style caching. Maybe we should discuss how to fix/refactor that as well.
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
4 tasks
Member
Author
|
Thanks @ivanium for the great comments. I've refactored a bit now and I think addressed all of them.
I think it should work naturally with chunked prefill. We do check within the current scheduled chunk. This change should also compose nicely with #46384. |
njhill
commented
Jul 8, 2026
Comment on lines
-730
to
-734
| num_uncached_common_prefix_tokens = getattr( | ||
| self.kv_cache_manager.coordinator, | ||
| "num_uncached_common_prefix_tokens", | ||
| 0, | ||
| ) |
Member
Author
There was a problem hiding this comment.
removed this side-channel which seems kind of hacky
ivanium
approved these changes
Jul 8, 2026
ivanium
left a comment
Contributor
There was a problem hiding this comment.
LGTM! Nice work. I left a small cosmetic comment but don't want to block the merge
ivanium
requested changes
Jul 8, 2026
Signed-off-by: Nick Hill <nickhill123@gmail.com>
…ix-retention Signed-off-by: Nick Hill <nickhill123@gmail.com>
Open
1 task done
MengqingCao
pushed a commit
to vllm-project/vllm-ascend
that referenced
this pull request
Jul 24, 2026
### What this PR does / why we need it? Adapt vllm-ascend to vLLM main commits up to July 17. ### Changes | Files | Upstream vLLM change | vllm-ascend adaptation | |-------|---------------------|------------------------| | `vllm_ascend/_310p/ops/fla/idex.py`<br>`vllm_ascend/_310p/ops/fla/l2norm.py`<br>`vllm_ascend/ops/bailing_moe_linear_attn.py`<br>`vllm_ascend/ops/gdn.py`<br>`vllm_ascend/ops/triton/fla/chunk.py`<br>`vllm_ascend/patch/worker/patch_idex_310.py`<br>`vllm_ascend/patch/worker/patch_triton.py`<br>`tests/e2e/nightly/.../test_fused_recurrent_gated_delta_rule.py`<br>`tests/e2e/nightly/.../test_fused_sigmoid_gating_delta_rule.py`<br>`tests/ut/ops/test_gdn_attn_builder.py` | [vllm#48500](vllm-project/vllm#48500) relocated flash-linear-attention from `vllm.third_party.flash_linear_attention` to `vllm.model_executor.layers.fla` | Version-gated all FLA imports and monkey-patch sites with `vllm_version_is("0.25.1")` | | `vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/config_data.py` | [vllm#45939](vllm-project/vllm#45939) replaced SHA-256 rehashed grouped block hashes with chained fine-grained hashes (terminal hash identifies complete block) | Removed `_rehash_block_hash_group` and associated constants. `get_block_hashes` returns `block_hashes[idx + scale_factor - 1]` instead of computing compound hash. | | `vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/coordinator.py`<br>`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/pool_scheduler.py`<br>`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/pool_worker.py` | [vllm#45939](vllm-project/vllm#45939) added `hash_block_size` to block pool API; [vllm#46384](vllm-project/vllm#46384) renamed `prefix_match_unit` -> `hash_block_size`, moved hash resolution into individual managers | `ExternalCachedBlockPool` takes `hash_block_size`. `_find_longest_cache_hit` return version-gated. `block_hashes_for_spec` no-op on main. `prefix_match_unit` fallback on main. | | `vllm_ascend/core/recompute_scheduler.py`<br>`vllm_ascend/core/scheduler_profiling_chunk.py`<br>`vllm_ascend/core/single_type_kv_cache_manager.py` | [vllm#46384](vllm-project/vllm#46384) / [vllm#47782](vllm-project/vllm#47782) changed `get_computed_blocks` 3-tuple -> 2-tuple, `get_num_blocks_to_allocate` params, `find_longest_cache_hit` return | Version-gated unpacking with `cast`. `num_tokens_main_model` made optional. `find_longest_cache_hit` returns blocks-only on v0.25.1 vs `(blocks, hit_length)` on main. | | `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` | [vllm#46384](vllm-project/vllm#46384) / [vllm#47782](vllm-project/vllm#47782) restructured return contracts; added partial Mamba hash hit support on main | Added `enable_partial_hash_hits` (main-only). Added `_cache_hit_alignment_tokens`. `find_longest_cache_hit` tracks per-group lengths, uses `cdiv`. `find_longest_cache_hit_per_group` simplified to single-pass; returns per-group `(blocks, lengths)`. Fixed v0.25.1 `hit_length` over-counting: use physical `spec.block_size` (not `_get_effective_block_size()` which includes `compress_ratio`) when multiplying by physical block count — affected both `find_longest_cache_hit` and `find_longest_cache_hit_per_group`; without this, MLA models (DS-V2-Lite/V3/V4) got 4x inflated hit length -> scheduler skipped non-cached tokens -> segfault. | | `vllm_ascend/patch/platform/patch_mamba_manager.py` | [vllm#46384](vllm-project/vllm#46384) / [vllm#47782](vllm-project/vllm#47782) changed method signatures | Version-gated `find_longest_cache_hit` (delegates to super on main) and `get_num_blocks_to_allocate` (optional params). | | `vllm_ascend/patch/worker/patch_qwen3_5.py` | [vllm#47006](vllm-project/vllm#47006) changed Qwen3-Next sequence-parallel gather contracts | Added `_ascend_all_gather_hidden_and_residual` (main-only). Gated `Qwen3_5DecoderLayer.forward` to v0.25.1. | | `vllm_ascend/ops/vocab_parallel_embedding.py` | [vllm#48390](vllm-project/vllm#48390) changed LM head apply contract | Added `_apply_head` routing to `quant_method.apply` on v0.25.1 and `super()._apply_head` on main. | | `vllm_ascend/patch/worker/__init__.py`<br>`vllm_ascend/patch/worker/patch_v2/patch_eagle_speculator.py`<br>`vllm_ascend/worker/v2/spec_decode/eagle/aclgraph.py`<br>`vllm_ascend/worker/v2/spec_decode/eagle/speculator.py`<br>`vllm_ascend/worker/v2/spec_decode/dflash/speculator.py`<br>`vllm_ascend/worker/v2/spec_decode/dspark/speculator.py` | [vllm#48261](vllm-project/vllm#48261) / [vllm#48167](vllm-project/vllm#48167) unified `PrefillSpeculatorCudaGraphManager` + `DecodeSpeculatorCudaGraphManager` -> `SpeculatorCudaGraphManager`; `capture()` parameterless; `set_attn` new params; `dflash_causal` -> `_group_causal` | Gated MRV2 spec decode imports to main-only. Consolidated Eagle managers into `EagleAclGraphManager`. Simplified DFlash/DSpark `capture()`. Added `target_input_buffers`/`target_attn_groups` to `set_attn`. Renamed `dflash_causal` -> `_group_causal` in DSpark `build_draft_attn_metadatas` (DFlash was already updated; DSpark was missed). | | `tests/ut/ops/test_gdn_layerwise_kv.py`<br>`tests/ut/ops/a2/test_gdn_layerwise_kv.py` | [vllm#46998](vllm-project/vllm#46998) changed GDN `forward` from `(hidden_states, output)` -> `(hidden_states) -> Tensor` | Added `_run_gdn_forward` wrapper handling both calling conventions. | | `tests/e2e/conftest.py` | [vllm#48549](vllm-project/vllm#48549) removed `swap_space` from `LLM` | Removed from `VllmRunner` and `DPVllmRunner`. | | `tests/e2e/pull_request/one_card/model_runner_v2/test_basic.py` | [vllm#48261](vllm-project/vllm#48261) MRV2 spec decode main-only | Added `_SKIP_V025_MRV2_SPEC_DECODE` skip marker. | | `tests/ut/core/test_recompute_scheduler.py` | Upstream main's `_free_request` accesses `self.ec_connector` (new attribute) | Added `scheduler.ec_connector = None` to test setup — scheduler is constructed via `__new__` (bypasses `__init__`), so the attribute must be mocked explicitly. | | `tests/ut/distributed/ascend_store/test_config_data.py`<br>`tests/ut/distributed/ascend_store/test_coordinator.py`<br>`tests/ut/distributed/ascend_store/test_pool_worker.py` | [vllm#45939](vllm-project/vllm#45939) / [vllm#46384](vllm-project/vllm#46384) | Updated block hash assertions to terminal hash. Added `hash_block_size`. Version-gated mock returns. | | `tests/ut/ops/test_vocab_parallel_embedding.py` | [vllm#48390](vllm-project/vllm#48390) | Replaced `patch()` with context manager for `set_current_vllm_config`. | | `tests/ut/patch/platform/test_prefix_cache_cp_patches.py` | [vllm#47782](vllm-project/vllm#47782) | Removed obsolete `num_prompt_tokens` kwarg. | | `tests/ut/patch/worker/test_patch_qwen3_5_mtp.py` | [vllm#48429](vllm-project/vllm#48429) | Set `use_attn_reduce_scatter_for_moe = False` on mock layer. | | `tests/ut/test_compressed_prefix_cache.py` | [vllm#46384](vllm-project/vllm#46384) / [vllm#47782](vllm-project/vllm#47782) | Version-gated unpacking. Added `test_find_longest_cache_hit_per_group`. | | `vllm_ascend/patch/__init__.py` | - | Updated HunyuanVL comment block. | - vLLM version: v0.25.1 - vLLM main: vllm-project/vllm@54503ec --------- Signed-off-by: wjunLu <wjunlu217@gmail.com> Signed-off-by: hfadzxy <starmoon_zhang@163.com> Co-authored-by: hfadzxy <starmoon_zhang@163.com>
Closed
7 tasks
This was referenced Jul 31, 2026
crb2nu
added a commit
to flexinfer/flexinfer
that referenced
this pull request
Aug 2, 2026
The ngram speculative-decoding re-canary (#66) is blocked on an engine-fatal assertion in `_prepare_inputs` at ~198k-token prefills. Upstream research (2026-08-01) traced the hybrid-GDN half of that crash to `_mamba_block_aligned_split`, which could floor `num_new_tokens` to 0 or amplify a negative to `-block_size` on every chunk of a chunked prefill. It was rewritten with an explicit clamp by vllm-project/vllm#47782 and first ships in v0.26.0. Bump the `gfx1100-native-w4a16` profile's base image, source ref, and published tag (`v023` -> `v026`), plus the matching build-args, BuildKit cache/output refs, and the pinned strings in the runtime-patch contract gate that assert them. The verify gate keeps its >= 0.23.0 floor: the RDNA3 W4A16 kernel and the `flexinfer_qwen35_text` plugin registration are both still required and compatible at v0.26.0. Runtime only. Digest promotion to the lanes and the speculation flip stay separate canaries per #66's plan, so `deploy/models/*` is deliberately untouched. The base-image change invalidates the BuildKit inline cache, so expect a full-length rebuild. Refs: services/flexinfer#66 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
ningjingbengxiaohai
pushed a commit
to vllm-project/vllm-ascend
that referenced
this pull request
Aug 3, 2026
…uple (#13275) ## Based on This PR is based on #13235 and re-submits the same production-code change on an independent branch. No additional functional changes are included. ## Why The copied `BalanceScheduler.schedule()` implementation still expects two return values from `KVCacheManager.get_computed_blocks()`. After the upstream vLLM API change, the method returns three values, causing the balance-scheduling path to fail at runtime. ## Upstream change Upstream PR: vllm-project/vllm#47782 The upstream change extends `get_computed_blocks()` to return: ```python (computed_blocks, num_computed_tokens, shared_prefix_boundary) ``` ## Downstream changes - Update `vllm_ascend/patch/platform/patch_balance_schedule.py` to unpack the third value into `request.shared_prefix_boundary`. ## Test plan - Full validation is delegated to PR CI. ## User-facing change No. - vLLM version: v0.26.0 - vLLM main: vllm-project/vllm@d02df74 --------- Signed-off-by: zhao-stack <80399320+zhao-stack@users.noreply.github.com>
3 of 4 tasks
xiayingqing
pushed a commit
to xiayingqing/vllm-ascend
that referenced
this pull request
Aug 3, 2026
…uple (vllm-project#13275) ## Based on This PR is based on vllm-project#13235 and re-submits the same production-code change on an independent branch. No additional functional changes are included. ## Why The copied `BalanceScheduler.schedule()` implementation still expects two return values from `KVCacheManager.get_computed_blocks()`. After the upstream vLLM API change, the method returns three values, causing the balance-scheduling path to fail at runtime. ## Upstream change Upstream PR: vllm-project/vllm#47782 The upstream change extends `get_computed_blocks()` to return: ```python (computed_blocks, num_computed_tokens, shared_prefix_boundary) ``` ## Downstream changes - Update `vllm_ascend/patch/platform/patch_balance_schedule.py` to unpack the third value into `request.shared_prefix_boundary`. ## Test plan - Full validation is delegated to PR CI. ## User-facing change No. - vLLM version: v0.26.0 - vLLM main: vllm-project/vllm@d02df74 --------- Signed-off-by: zhao-stack <80399320+zhao-stack@users.noreply.github.com>
HMCCMH
pushed a commit
to hotTea123/vllm-ascend
that referenced
this pull request
Aug 12, 2026
…uple (vllm-project#13275) ## Based on This PR is based on vllm-project#13235 and re-submits the same production-code change on an independent branch. No additional functional changes are included. ## Why The copied `BalanceScheduler.schedule()` implementation still expects two return values from `KVCacheManager.get_computed_blocks()`. After the upstream vLLM API change, the method returns three values, causing the balance-scheduling path to fail at runtime. ## Upstream change Upstream PR: vllm-project/vllm#47782 The upstream change extends `get_computed_blocks()` to return: ```python (computed_blocks, num_computed_tokens, shared_prefix_boundary) ``` ## Downstream changes - Update `vllm_ascend/patch/platform/patch_balance_schedule.py` to unpack the third value into `request.shared_prefix_boundary`. ## Test plan - Full validation is delegated to PR CI. ## User-facing change No. - vLLM version: v0.26.0 - vLLM main: vllm-project/vllm@d02df74 --------- Signed-off-by: zhao-stack <80399320+zhao-stack@users.noreply.github.com>
MmMmaru
pushed a commit
to jiaqi-lee/vllm-ascend
that referenced
this pull request
Aug 19, 2026
### What this PR does / why we need it? Adapt vllm-ascend to vLLM main commits up to July 17. ### Changes | Files | Upstream vLLM change | vllm-ascend adaptation | |-------|---------------------|------------------------| | `vllm_ascend/_310p/ops/fla/idex.py`<br>`vllm_ascend/_310p/ops/fla/l2norm.py`<br>`vllm_ascend/ops/bailing_moe_linear_attn.py`<br>`vllm_ascend/ops/gdn.py`<br>`vllm_ascend/ops/triton/fla/chunk.py`<br>`vllm_ascend/patch/worker/patch_idex_310.py`<br>`vllm_ascend/patch/worker/patch_triton.py`<br>`tests/e2e/nightly/.../test_fused_recurrent_gated_delta_rule.py`<br>`tests/e2e/nightly/.../test_fused_sigmoid_gating_delta_rule.py`<br>`tests/ut/ops/test_gdn_attn_builder.py` | [vllm#48500](vllm-project/vllm#48500) relocated flash-linear-attention from `vllm.third_party.flash_linear_attention` to `vllm.model_executor.layers.fla` | Version-gated all FLA imports and monkey-patch sites with `vllm_version_is("0.25.1")` | | `vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/config_data.py` | [vllm#45939](vllm-project/vllm#45939) replaced SHA-256 rehashed grouped block hashes with chained fine-grained hashes (terminal hash identifies complete block) | Removed `_rehash_block_hash_group` and associated constants. `get_block_hashes` returns `block_hashes[idx + scale_factor - 1]` instead of computing compound hash. | | `vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/coordinator.py`<br>`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/pool_scheduler.py`<br>`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/pool_worker.py` | [vllm#45939](vllm-project/vllm#45939) added `hash_block_size` to block pool API; [vllm#46384](vllm-project/vllm#46384) renamed `prefix_match_unit` -> `hash_block_size`, moved hash resolution into individual managers | `ExternalCachedBlockPool` takes `hash_block_size`. `_find_longest_cache_hit` return version-gated. `block_hashes_for_spec` no-op on main. `prefix_match_unit` fallback on main. | | `vllm_ascend/core/recompute_scheduler.py`<br>`vllm_ascend/core/scheduler_profiling_chunk.py`<br>`vllm_ascend/core/single_type_kv_cache_manager.py` | [vllm#46384](vllm-project/vllm#46384) / [vllm#47782](vllm-project/vllm#47782) changed `get_computed_blocks` 3-tuple -> 2-tuple, `get_num_blocks_to_allocate` params, `find_longest_cache_hit` return | Version-gated unpacking with `cast`. `num_tokens_main_model` made optional. `find_longest_cache_hit` returns blocks-only on v0.25.1 vs `(blocks, hit_length)` on main. | | `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` | [vllm#46384](vllm-project/vllm#46384) / [vllm#47782](vllm-project/vllm#47782) restructured return contracts; added partial Mamba hash hit support on main | Added `enable_partial_hash_hits` (main-only). Added `_cache_hit_alignment_tokens`. `find_longest_cache_hit` tracks per-group lengths, uses `cdiv`. `find_longest_cache_hit_per_group` simplified to single-pass; returns per-group `(blocks, lengths)`. Fixed v0.25.1 `hit_length` over-counting: use physical `spec.block_size` (not `_get_effective_block_size()` which includes `compress_ratio`) when multiplying by physical block count — affected both `find_longest_cache_hit` and `find_longest_cache_hit_per_group`; without this, MLA models (DS-V2-Lite/V3/V4) got 4x inflated hit length -> scheduler skipped non-cached tokens -> segfault. | | `vllm_ascend/patch/platform/patch_mamba_manager.py` | [vllm#46384](vllm-project/vllm#46384) / [vllm#47782](vllm-project/vllm#47782) changed method signatures | Version-gated `find_longest_cache_hit` (delegates to super on main) and `get_num_blocks_to_allocate` (optional params). | | `vllm_ascend/patch/worker/patch_qwen3_5.py` | [vllm#47006](vllm-project/vllm#47006) changed Qwen3-Next sequence-parallel gather contracts | Added `_ascend_all_gather_hidden_and_residual` (main-only). Gated `Qwen3_5DecoderLayer.forward` to v0.25.1. | | `vllm_ascend/ops/vocab_parallel_embedding.py` | [vllm#48390](vllm-project/vllm#48390) changed LM head apply contract | Added `_apply_head` routing to `quant_method.apply` on v0.25.1 and `super()._apply_head` on main. | | `vllm_ascend/patch/worker/__init__.py`<br>`vllm_ascend/patch/worker/patch_v2/patch_eagle_speculator.py`<br>`vllm_ascend/worker/v2/spec_decode/eagle/aclgraph.py`<br>`vllm_ascend/worker/v2/spec_decode/eagle/speculator.py`<br>`vllm_ascend/worker/v2/spec_decode/dflash/speculator.py`<br>`vllm_ascend/worker/v2/spec_decode/dspark/speculator.py` | [vllm#48261](vllm-project/vllm#48261) / [vllm#48167](vllm-project/vllm#48167) unified `PrefillSpeculatorCudaGraphManager` + `DecodeSpeculatorCudaGraphManager` -> `SpeculatorCudaGraphManager`; `capture()` parameterless; `set_attn` new params; `dflash_causal` -> `_group_causal` | Gated MRV2 spec decode imports to main-only. Consolidated Eagle managers into `EagleAclGraphManager`. Simplified DFlash/DSpark `capture()`. Added `target_input_buffers`/`target_attn_groups` to `set_attn`. Renamed `dflash_causal` -> `_group_causal` in DSpark `build_draft_attn_metadatas` (DFlash was already updated; DSpark was missed). | | `tests/ut/ops/test_gdn_layerwise_kv.py`<br>`tests/ut/ops/a2/test_gdn_layerwise_kv.py` | [vllm#46998](vllm-project/vllm#46998) changed GDN `forward` from `(hidden_states, output)` -> `(hidden_states) -> Tensor` | Added `_run_gdn_forward` wrapper handling both calling conventions. | | `tests/e2e/conftest.py` | [vllm#48549](vllm-project/vllm#48549) removed `swap_space` from `LLM` | Removed from `VllmRunner` and `DPVllmRunner`. | | `tests/e2e/pull_request/one_card/model_runner_v2/test_basic.py` | [vllm#48261](vllm-project/vllm#48261) MRV2 spec decode main-only | Added `_SKIP_V025_MRV2_SPEC_DECODE` skip marker. | | `tests/ut/core/test_recompute_scheduler.py` | Upstream main's `_free_request` accesses `self.ec_connector` (new attribute) | Added `scheduler.ec_connector = None` to test setup — scheduler is constructed via `__new__` (bypasses `__init__`), so the attribute must be mocked explicitly. | | `tests/ut/distributed/ascend_store/test_config_data.py`<br>`tests/ut/distributed/ascend_store/test_coordinator.py`<br>`tests/ut/distributed/ascend_store/test_pool_worker.py` | [vllm#45939](vllm-project/vllm#45939) / [vllm#46384](vllm-project/vllm#46384) | Updated block hash assertions to terminal hash. Added `hash_block_size`. Version-gated mock returns. | | `tests/ut/ops/test_vocab_parallel_embedding.py` | [vllm#48390](vllm-project/vllm#48390) | Replaced `patch()` with context manager for `set_current_vllm_config`. | | `tests/ut/patch/platform/test_prefix_cache_cp_patches.py` | [vllm#47782](vllm-project/vllm#47782) | Removed obsolete `num_prompt_tokens` kwarg. | | `tests/ut/patch/worker/test_patch_qwen3_5_mtp.py` | [vllm#48429](vllm-project/vllm#48429) | Set `use_attn_reduce_scatter_for_moe = False` on mock layer. | | `tests/ut/test_compressed_prefix_cache.py` | [vllm#46384](vllm-project/vllm#46384) / [vllm#47782](vllm-project/vllm#47782) | Version-gated unpacking. Added `test_find_longest_cache_hit_per_group`. | | `vllm_ascend/patch/__init__.py` | - | Updated HunyuanVL comment block. | - vLLM version: v0.25.1 - vLLM main: vllm-project/vllm@54503ec --------- Signed-off-by: wjunLu <wjunlu217@gmail.com> Signed-off-by: hfadzxy <starmoon_zhang@163.com> Co-authored-by: hfadzxy <starmoon_zhang@163.com>
MmMmaru
pushed a commit
to jiaqi-lee/vllm-ascend
that referenced
this pull request
Aug 19, 2026
…uple (vllm-project#13275) ## Based on This PR is based on vllm-project#13235 and re-submits the same production-code change on an independent branch. No additional functional changes are included. ## Why The copied `BalanceScheduler.schedule()` implementation still expects two return values from `KVCacheManager.get_computed_blocks()`. After the upstream vLLM API change, the method returns three values, causing the balance-scheduling path to fail at runtime. ## Upstream change Upstream PR: vllm-project/vllm#47782 The upstream change extends `get_computed_blocks()` to return: ```python (computed_blocks, num_computed_tokens, shared_prefix_boundary) ``` ## Downstream changes - Update `vllm_ascend/patch/platform/patch_balance_schedule.py` to unpack the third value into `request.shared_prefix_boundary`. ## Test plan - Full validation is delegated to PR CI. ## User-facing change No. - vLLM version: v0.26.0 - vLLM main: vllm-project/vllm@d02df74 --------- Signed-off-by: zhao-stack <80399320+zhao-stack@users.noreply.github.com>
shiqiangA
pushed a commit
to shiqiangA/vllm-ascend
that referenced
this pull request
Aug 20, 2026
### What this PR does / why we need it? Adapt vllm-ascend to vLLM main commits up to July 17. ### Changes | Files | Upstream vLLM change | vllm-ascend adaptation | |-------|---------------------|------------------------| | `vllm_ascend/_310p/ops/fla/idex.py`<br>`vllm_ascend/_310p/ops/fla/l2norm.py`<br>`vllm_ascend/ops/bailing_moe_linear_attn.py`<br>`vllm_ascend/ops/gdn.py`<br>`vllm_ascend/ops/triton/fla/chunk.py`<br>`vllm_ascend/patch/worker/patch_idex_310.py`<br>`vllm_ascend/patch/worker/patch_triton.py`<br>`tests/e2e/nightly/.../test_fused_recurrent_gated_delta_rule.py`<br>`tests/e2e/nightly/.../test_fused_sigmoid_gating_delta_rule.py`<br>`tests/ut/ops/test_gdn_attn_builder.py` | [vllm#48500](vllm-project/vllm#48500) relocated flash-linear-attention from `vllm.third_party.flash_linear_attention` to `vllm.model_executor.layers.fla` | Version-gated all FLA imports and monkey-patch sites with `vllm_version_is("0.25.1")` | | `vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/config_data.py` | [vllm#45939](vllm-project/vllm#45939) replaced SHA-256 rehashed grouped block hashes with chained fine-grained hashes (terminal hash identifies complete block) | Removed `_rehash_block_hash_group` and associated constants. `get_block_hashes` returns `block_hashes[idx + scale_factor - 1]` instead of computing compound hash. | | `vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/coordinator.py`<br>`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/pool_scheduler.py`<br>`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/pool_worker.py` | [vllm#45939](vllm-project/vllm#45939) added `hash_block_size` to block pool API; [vllm#46384](vllm-project/vllm#46384) renamed `prefix_match_unit` -> `hash_block_size`, moved hash resolution into individual managers | `ExternalCachedBlockPool` takes `hash_block_size`. `_find_longest_cache_hit` return version-gated. `block_hashes_for_spec` no-op on main. `prefix_match_unit` fallback on main. | | `vllm_ascend/core/recompute_scheduler.py`<br>`vllm_ascend/core/scheduler_profiling_chunk.py`<br>`vllm_ascend/core/single_type_kv_cache_manager.py` | [vllm#46384](vllm-project/vllm#46384) / [vllm#47782](vllm-project/vllm#47782) changed `get_computed_blocks` 3-tuple -> 2-tuple, `get_num_blocks_to_allocate` params, `find_longest_cache_hit` return | Version-gated unpacking with `cast`. `num_tokens_main_model` made optional. `find_longest_cache_hit` returns blocks-only on v0.25.1 vs `(blocks, hit_length)` on main. | | `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` | [vllm#46384](vllm-project/vllm#46384) / [vllm#47782](vllm-project/vllm#47782) restructured return contracts; added partial Mamba hash hit support on main | Added `enable_partial_hash_hits` (main-only). Added `_cache_hit_alignment_tokens`. `find_longest_cache_hit` tracks per-group lengths, uses `cdiv`. `find_longest_cache_hit_per_group` simplified to single-pass; returns per-group `(blocks, lengths)`. Fixed v0.25.1 `hit_length` over-counting: use physical `spec.block_size` (not `_get_effective_block_size()` which includes `compress_ratio`) when multiplying by physical block count — affected both `find_longest_cache_hit` and `find_longest_cache_hit_per_group`; without this, MLA models (DS-V2-Lite/V3/V4) got 4x inflated hit length -> scheduler skipped non-cached tokens -> segfault. | | `vllm_ascend/patch/platform/patch_mamba_manager.py` | [vllm#46384](vllm-project/vllm#46384) / [vllm#47782](vllm-project/vllm#47782) changed method signatures | Version-gated `find_longest_cache_hit` (delegates to super on main) and `get_num_blocks_to_allocate` (optional params). | | `vllm_ascend/patch/worker/patch_qwen3_5.py` | [vllm#47006](vllm-project/vllm#47006) changed Qwen3-Next sequence-parallel gather contracts | Added `_ascend_all_gather_hidden_and_residual` (main-only). Gated `Qwen3_5DecoderLayer.forward` to v0.25.1. | | `vllm_ascend/ops/vocab_parallel_embedding.py` | [vllm#48390](vllm-project/vllm#48390) changed LM head apply contract | Added `_apply_head` routing to `quant_method.apply` on v0.25.1 and `super()._apply_head` on main. | | `vllm_ascend/patch/worker/__init__.py`<br>`vllm_ascend/patch/worker/patch_v2/patch_eagle_speculator.py`<br>`vllm_ascend/worker/v2/spec_decode/eagle/aclgraph.py`<br>`vllm_ascend/worker/v2/spec_decode/eagle/speculator.py`<br>`vllm_ascend/worker/v2/spec_decode/dflash/speculator.py`<br>`vllm_ascend/worker/v2/spec_decode/dspark/speculator.py` | [vllm#48261](vllm-project/vllm#48261) / [vllm#48167](vllm-project/vllm#48167) unified `PrefillSpeculatorCudaGraphManager` + `DecodeSpeculatorCudaGraphManager` -> `SpeculatorCudaGraphManager`; `capture()` parameterless; `set_attn` new params; `dflash_causal` -> `_group_causal` | Gated MRV2 spec decode imports to main-only. Consolidated Eagle managers into `EagleAclGraphManager`. Simplified DFlash/DSpark `capture()`. Added `target_input_buffers`/`target_attn_groups` to `set_attn`. Renamed `dflash_causal` -> `_group_causal` in DSpark `build_draft_attn_metadatas` (DFlash was already updated; DSpark was missed). | | `tests/ut/ops/test_gdn_layerwise_kv.py`<br>`tests/ut/ops/a2/test_gdn_layerwise_kv.py` | [vllm#46998](vllm-project/vllm#46998) changed GDN `forward` from `(hidden_states, output)` -> `(hidden_states) -> Tensor` | Added `_run_gdn_forward` wrapper handling both calling conventions. | | `tests/e2e/conftest.py` | [vllm#48549](vllm-project/vllm#48549) removed `swap_space` from `LLM` | Removed from `VllmRunner` and `DPVllmRunner`. | | `tests/e2e/pull_request/one_card/model_runner_v2/test_basic.py` | [vllm#48261](vllm-project/vllm#48261) MRV2 spec decode main-only | Added `_SKIP_V025_MRV2_SPEC_DECODE` skip marker. | | `tests/ut/core/test_recompute_scheduler.py` | Upstream main's `_free_request` accesses `self.ec_connector` (new attribute) | Added `scheduler.ec_connector = None` to test setup — scheduler is constructed via `__new__` (bypasses `__init__`), so the attribute must be mocked explicitly. | | `tests/ut/distributed/ascend_store/test_config_data.py`<br>`tests/ut/distributed/ascend_store/test_coordinator.py`<br>`tests/ut/distributed/ascend_store/test_pool_worker.py` | [vllm#45939](vllm-project/vllm#45939) / [vllm#46384](vllm-project/vllm#46384) | Updated block hash assertions to terminal hash. Added `hash_block_size`. Version-gated mock returns. | | `tests/ut/ops/test_vocab_parallel_embedding.py` | [vllm#48390](vllm-project/vllm#48390) | Replaced `patch()` with context manager for `set_current_vllm_config`. | | `tests/ut/patch/platform/test_prefix_cache_cp_patches.py` | [vllm#47782](vllm-project/vllm#47782) | Removed obsolete `num_prompt_tokens` kwarg. | | `tests/ut/patch/worker/test_patch_qwen3_5_mtp.py` | [vllm#48429](vllm-project/vllm#48429) | Set `use_attn_reduce_scatter_for_moe = False` on mock layer. | | `tests/ut/test_compressed_prefix_cache.py` | [vllm#46384](vllm-project/vllm#46384) / [vllm#47782](vllm-project/vllm#47782) | Version-gated unpacking. Added `test_find_longest_cache_hit_per_group`. | | `vllm_ascend/patch/__init__.py` | - | Updated HunyuanVL comment block. | - vLLM version: v0.25.1 - vLLM main: vllm-project/vllm@54503ec --------- Signed-off-by: wjunLu <wjunlu217@gmail.com> Signed-off-by: hfadzxy <starmoon_zhang@163.com> Co-authored-by: hfadzxy <starmoon_zhang@163.com>
shiqiangA
pushed a commit
to shiqiangA/vllm-ascend
that referenced
this pull request
Aug 20, 2026
…uple (vllm-project#13275) ## Based on This PR is based on vllm-project#13235 and re-submits the same production-code change on an independent branch. No additional functional changes are included. ## Why The copied `BalanceScheduler.schedule()` implementation still expects two return values from `KVCacheManager.get_computed_blocks()`. After the upstream vLLM API change, the method returns three values, causing the balance-scheduling path to fail at runtime. ## Upstream change Upstream PR: vllm-project/vllm#47782 The upstream change extends `get_computed_blocks()` to return: ```python (computed_blocks, num_computed_tokens, shared_prefix_boundary) ``` ## Downstream changes - Update `vllm_ascend/patch/platform/patch_balance_schedule.py` to unpack the third value into `request.shared_prefix_boundary`. ## Test plan - Full validation is delegated to PR CI. ## User-facing change No. - vLLM version: v0.26.0 - vLLM main: vllm-project/vllm@d02df74 --------- Signed-off-by: zhao-stack <80399320+zhao-stack@users.noreply.github.com>
puririshi98
pushed a commit
to puririshi98/vllm
that referenced
this pull request
Aug 31, 2026
…vllm-project#47782) Signed-off-by: Nick Hill <nickhill123@gmail.com> (cherry picked from commit 8ac8375) (cherry picked from commit f9fda36f49b1314b03031e9263f9a5be82b38598)
3 of 8 tasks
Leetrytry
pushed a commit
to Leetrytry/vllm-ascend
that referenced
this pull request
Sep 11, 2026
### What this PR does / why we need it? Adapt vllm-ascend to vLLM main commits up to July 17. ### Changes | Files | Upstream vLLM change | vllm-ascend adaptation | |-------|---------------------|------------------------| | `vllm_ascend/_310p/ops/fla/idex.py`<br>`vllm_ascend/_310p/ops/fla/l2norm.py`<br>`vllm_ascend/ops/bailing_moe_linear_attn.py`<br>`vllm_ascend/ops/gdn.py`<br>`vllm_ascend/ops/triton/fla/chunk.py`<br>`vllm_ascend/patch/worker/patch_idex_310.py`<br>`vllm_ascend/patch/worker/patch_triton.py`<br>`tests/e2e/nightly/.../test_fused_recurrent_gated_delta_rule.py`<br>`tests/e2e/nightly/.../test_fused_sigmoid_gating_delta_rule.py`<br>`tests/ut/ops/test_gdn_attn_builder.py` | [vllm#48500](vllm-project/vllm#48500) relocated flash-linear-attention from `vllm.third_party.flash_linear_attention` to `vllm.model_executor.layers.fla` | Version-gated all FLA imports and monkey-patch sites with `vllm_version_is("0.25.1")` | | `vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/config_data.py` | [vllm#45939](vllm-project/vllm#45939) replaced SHA-256 rehashed grouped block hashes with chained fine-grained hashes (terminal hash identifies complete block) | Removed `_rehash_block_hash_group` and associated constants. `get_block_hashes` returns `block_hashes[idx + scale_factor - 1]` instead of computing compound hash. | | `vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/coordinator.py`<br>`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/pool_scheduler.py`<br>`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/pool_worker.py` | [vllm#45939](vllm-project/vllm#45939) added `hash_block_size` to block pool API; [vllm#46384](vllm-project/vllm#46384) renamed `prefix_match_unit` -> `hash_block_size`, moved hash resolution into individual managers | `ExternalCachedBlockPool` takes `hash_block_size`. `_find_longest_cache_hit` return version-gated. `block_hashes_for_spec` no-op on main. `prefix_match_unit` fallback on main. | | `vllm_ascend/core/recompute_scheduler.py`<br>`vllm_ascend/core/scheduler_profiling_chunk.py`<br>`vllm_ascend/core/single_type_kv_cache_manager.py` | [vllm#46384](vllm-project/vllm#46384) / [vllm#47782](vllm-project/vllm#47782) changed `get_computed_blocks` 3-tuple -> 2-tuple, `get_num_blocks_to_allocate` params, `find_longest_cache_hit` return | Version-gated unpacking with `cast`. `num_tokens_main_model` made optional. `find_longest_cache_hit` returns blocks-only on v0.25.1 vs `(blocks, hit_length)` on main. | | `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` | [vllm#46384](vllm-project/vllm#46384) / [vllm#47782](vllm-project/vllm#47782) restructured return contracts; added partial Mamba hash hit support on main | Added `enable_partial_hash_hits` (main-only). Added `_cache_hit_alignment_tokens`. `find_longest_cache_hit` tracks per-group lengths, uses `cdiv`. `find_longest_cache_hit_per_group` simplified to single-pass; returns per-group `(blocks, lengths)`. Fixed v0.25.1 `hit_length` over-counting: use physical `spec.block_size` (not `_get_effective_block_size()` which includes `compress_ratio`) when multiplying by physical block count — affected both `find_longest_cache_hit` and `find_longest_cache_hit_per_group`; without this, MLA models (DS-V2-Lite/V3/V4) got 4x inflated hit length -> scheduler skipped non-cached tokens -> segfault. | | `vllm_ascend/patch/platform/patch_mamba_manager.py` | [vllm#46384](vllm-project/vllm#46384) / [vllm#47782](vllm-project/vllm#47782) changed method signatures | Version-gated `find_longest_cache_hit` (delegates to super on main) and `get_num_blocks_to_allocate` (optional params). | | `vllm_ascend/patch/worker/patch_qwen3_5.py` | [vllm#47006](vllm-project/vllm#47006) changed Qwen3-Next sequence-parallel gather contracts | Added `_ascend_all_gather_hidden_and_residual` (main-only). Gated `Qwen3_5DecoderLayer.forward` to v0.25.1. | | `vllm_ascend/ops/vocab_parallel_embedding.py` | [vllm#48390](vllm-project/vllm#48390) changed LM head apply contract | Added `_apply_head` routing to `quant_method.apply` on v0.25.1 and `super()._apply_head` on main. | | `vllm_ascend/patch/worker/__init__.py`<br>`vllm_ascend/patch/worker/patch_v2/patch_eagle_speculator.py`<br>`vllm_ascend/worker/v2/spec_decode/eagle/aclgraph.py`<br>`vllm_ascend/worker/v2/spec_decode/eagle/speculator.py`<br>`vllm_ascend/worker/v2/spec_decode/dflash/speculator.py`<br>`vllm_ascend/worker/v2/spec_decode/dspark/speculator.py` | [vllm#48261](vllm-project/vllm#48261) / [vllm#48167](vllm-project/vllm#48167) unified `PrefillSpeculatorCudaGraphManager` + `DecodeSpeculatorCudaGraphManager` -> `SpeculatorCudaGraphManager`; `capture()` parameterless; `set_attn` new params; `dflash_causal` -> `_group_causal` | Gated MRV2 spec decode imports to main-only. Consolidated Eagle managers into `EagleAclGraphManager`. Simplified DFlash/DSpark `capture()`. Added `target_input_buffers`/`target_attn_groups` to `set_attn`. Renamed `dflash_causal` -> `_group_causal` in DSpark `build_draft_attn_metadatas` (DFlash was already updated; DSpark was missed). | | `tests/ut/ops/test_gdn_layerwise_kv.py`<br>`tests/ut/ops/a2/test_gdn_layerwise_kv.py` | [vllm#46998](vllm-project/vllm#46998) changed GDN `forward` from `(hidden_states, output)` -> `(hidden_states) -> Tensor` | Added `_run_gdn_forward` wrapper handling both calling conventions. | | `tests/e2e/conftest.py` | [vllm#48549](vllm-project/vllm#48549) removed `swap_space` from `LLM` | Removed from `VllmRunner` and `DPVllmRunner`. | | `tests/e2e/pull_request/one_card/model_runner_v2/test_basic.py` | [vllm#48261](vllm-project/vllm#48261) MRV2 spec decode main-only | Added `_SKIP_V025_MRV2_SPEC_DECODE` skip marker. | | `tests/ut/core/test_recompute_scheduler.py` | Upstream main's `_free_request` accesses `self.ec_connector` (new attribute) | Added `scheduler.ec_connector = None` to test setup — scheduler is constructed via `__new__` (bypasses `__init__`), so the attribute must be mocked explicitly. | | `tests/ut/distributed/ascend_store/test_config_data.py`<br>`tests/ut/distributed/ascend_store/test_coordinator.py`<br>`tests/ut/distributed/ascend_store/test_pool_worker.py` | [vllm#45939](vllm-project/vllm#45939) / [vllm#46384](vllm-project/vllm#46384) | Updated block hash assertions to terminal hash. Added `hash_block_size`. Version-gated mock returns. | | `tests/ut/ops/test_vocab_parallel_embedding.py` | [vllm#48390](vllm-project/vllm#48390) | Replaced `patch()` with context manager for `set_current_vllm_config`. | | `tests/ut/patch/platform/test_prefix_cache_cp_patches.py` | [vllm#47782](vllm-project/vllm#47782) | Removed obsolete `num_prompt_tokens` kwarg. | | `tests/ut/patch/worker/test_patch_qwen3_5_mtp.py` | [vllm#48429](vllm-project/vllm#48429) | Set `use_attn_reduce_scatter_for_moe = False` on mock layer. | | `tests/ut/test_compressed_prefix_cache.py` | [vllm#46384](vllm-project/vllm#46384) / [vllm#47782](vllm-project/vllm#47782) | Version-gated unpacking. Added `test_find_longest_cache_hit_per_group`. | | `vllm_ascend/patch/__init__.py` | - | Updated HunyuanVL comment block. | - vLLM version: v0.25.1 - vLLM main: vllm-project/vllm@54503ec --------- Signed-off-by: wjunLu <wjunlu217@gmail.com> Signed-off-by: hfadzxy <starmoon_zhang@163.com> Co-authored-by: hfadzxy <starmoon_zhang@163.com>
This was referenced Sep 13, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The Marconi paper was implemented in #37898, which enables caching of mamba state for shared prefixes in "align" mode.
This PR extends the shared prefix boundary caching to the retention-interval sparse caching added in #43447 and #45845 for both mamba states and SWA.
In particular it means we can set retention interval = 0 for optimal long-context/agent session caching and also benefit from implicitly cached system prompt.