Skip to content

[Core] Preserve Marconi caching with selective hybrid cache retention - #47782

Merged
njhill merged 10 commits into
vllm-project:mainfrom
njhill:fix-mamba-shared-prefix-retention
Jul 13, 2026
Merged

njhill merged 10 commits into
vllm-project:mainfrom
njhill:fix-mamba-shared-prefix-retention

Conversation

@njhill

@njhill njhill commented Jul 6, 2026 •

Copy link
Copy Markdown
Member

The Marconi paper was implemented in #37898, which enables caching of mamba state for shared prefixes in "align" mode.

This PR extends the shared prefix boundary caching to the retention-interval sparse caching added in #43447 and #45845 for both mamba states and SWA.

In particular it means we can set retention interval = 0 for optimal long-context/agent session caching and also benefit from implicitly cached system prompt.

njhill added 3 commits July 6, 2026 19:01
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mergify mergify Bot added the v1 label Jul 6, 2026
@njhill
njhill requested a review from tdoublep July 6, 2026 22:29

@ivanium ivanium left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great idea! Overall makes sense. Left some comments.

A few high level thoughts:

  • how does this work with chunked prefill? I feel somewhere we need to check if request.shared_prefix_boundary falls in the current scheduled chunk.
  • We also need to update the mooncake store connector. Shouldn't be huge changes but maybe we can do it in a followup PR too
  • As commented in scheduler, I found the previous handling of hybrid linear model is a bit hacky and silently breaks the marconi-style caching. Maybe we should discuss how to fix/refactor that as well.
Comment thread vllm/v1/core/sched/scheduler.py
Comment thread vllm/v1/core/single_type_kv_cache_manager.py Outdated
Comment thread vllm/v1/core/sched/scheduler.py Outdated
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
@njhill

njhill commented Jul 8, 2026 •

Copy link
Copy Markdown
Member Author

Thanks @ivanium for the great comments. I've refactored a bit now and I think addressed all of them.

  • how does this work with chunked prefill? I feel somewhere we need to check if request.shared_prefix_boundary falls in the current scheduled chunk.

I think it should work naturally with chunked prefill. We do check within the current scheduled chunk.

This change should also compose nicely with #46384.

Comment on lines -730 to -734
num_uncached_common_prefix_tokens = getattr(
self.kv_cache_manager.coordinator,
"num_uncached_common_prefix_tokens",
0,
)

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

removed this side-channel which seems kind of hacky

@njhill njhill added the ready ONLY add when PR is ready to merge/full CI is needed label Jul 8, 2026

@ivanium ivanium left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM! Nice work. I left a small cosmetic comment but don't want to block the merge

Comment thread vllm/v1/core/sched/scheduler.py Outdated

@ivanium ivanium left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just realized this issue and wanted to double check. Sorry for the churn 🙇

Comment thread vllm/v1/core/sched/scheduler.py
njhill added 2 commits July 9, 2026 09:56
Signed-off-by: Nick Hill <nickhill123@gmail.com>
…ix-retention

Signed-off-by: Nick Hill <nickhill123@gmail.com>
MengqingCao pushed a commit to vllm-project/vllm-ascend that referenced this pull request Jul 24, 2026
### What this PR does / why we need it?

Adapt vllm-ascend to vLLM main commits up to July 17.

### Changes

| Files | Upstream vLLM change | vllm-ascend adaptation |
|-------|---------------------|------------------------|
|
`vllm_ascend/_310p/ops/fla/idex.py`<br>`vllm_ascend/_310p/ops/fla/l2norm.py`<br>`vllm_ascend/ops/bailing_moe_linear_attn.py`<br>`vllm_ascend/ops/gdn.py`<br>`vllm_ascend/ops/triton/fla/chunk.py`<br>`vllm_ascend/patch/worker/patch_idex_310.py`<br>`vllm_ascend/patch/worker/patch_triton.py`<br>`tests/e2e/nightly/.../test_fused_recurrent_gated_delta_rule.py`<br>`tests/e2e/nightly/.../test_fused_sigmoid_gating_delta_rule.py`<br>`tests/ut/ops/test_gdn_attn_builder.py`
| [vllm#48500](vllm-project/vllm#48500)
relocated flash-linear-attention from
`vllm.third_party.flash_linear_attention` to
`vllm.model_executor.layers.fla` | Version-gated all FLA imports and
monkey-patch sites with `vllm_version_is("0.25.1")` |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/config_data.py`
| [vllm#45939](vllm-project/vllm#45939) replaced
SHA-256 rehashed grouped block hashes with chained fine-grained hashes
(terminal hash identifies complete block) | Removed
`_rehash_block_hash_group` and associated constants. `get_block_hashes`
returns `block_hashes[idx + scale_factor - 1]` instead of computing
compound hash. |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/coordinator.py`<br>`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/pool_scheduler.py`<br>`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/pool_worker.py`
| [vllm#45939](vllm-project/vllm#45939) added
`hash_block_size` to block pool API;
[vllm#46384](vllm-project/vllm#46384) renamed
`prefix_match_unit` -> `hash_block_size`, moved hash resolution into
individual managers | `ExternalCachedBlockPool` takes `hash_block_size`.
`_find_longest_cache_hit` return version-gated. `block_hashes_for_spec`
no-op on main. `prefix_match_unit` fallback on main. |
|
`vllm_ascend/core/recompute_scheduler.py`<br>`vllm_ascend/core/scheduler_profiling_chunk.py`<br>`vllm_ascend/core/single_type_kv_cache_manager.py`
| [vllm#46384](vllm-project/vllm#46384) /
[vllm#47782](vllm-project/vllm#47782) changed
`get_computed_blocks` 3-tuple -> 2-tuple, `get_num_blocks_to_allocate`
params, `find_longest_cache_hit` return | Version-gated unpacking with
`cast`. `num_tokens_main_model` made optional. `find_longest_cache_hit`
returns blocks-only on v0.25.1 vs `(blocks, hit_length)` on main. |
| `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` |
[vllm#46384](vllm-project/vllm#46384) /
[vllm#47782](vllm-project/vllm#47782)
restructured return contracts; added partial Mamba hash hit support on
main | Added `enable_partial_hash_hits` (main-only). Added
`_cache_hit_alignment_tokens`. `find_longest_cache_hit` tracks per-group
lengths, uses `cdiv`. `find_longest_cache_hit_per_group` simplified to
single-pass; returns per-group `(blocks, lengths)`. Fixed v0.25.1
`hit_length` over-counting: use physical `spec.block_size` (not
`_get_effective_block_size()` which includes `compress_ratio`) when
multiplying by physical block count — affected both
`find_longest_cache_hit` and `find_longest_cache_hit_per_group`; without
this, MLA models (DS-V2-Lite/V3/V4) got 4x inflated hit length ->
scheduler skipped non-cached tokens -> segfault. |
| `vllm_ascend/patch/platform/patch_mamba_manager.py` |
[vllm#46384](vllm-project/vllm#46384) /
[vllm#47782](vllm-project/vllm#47782) changed
method signatures | Version-gated `find_longest_cache_hit` (delegates to
super on main) and `get_num_blocks_to_allocate` (optional params). |
| `vllm_ascend/patch/worker/patch_qwen3_5.py` |
[vllm#47006](vllm-project/vllm#47006) changed
Qwen3-Next sequence-parallel gather contracts | Added
`_ascend_all_gather_hidden_and_residual` (main-only). Gated
`Qwen3_5DecoderLayer.forward` to v0.25.1. |
| `vllm_ascend/ops/vocab_parallel_embedding.py` |
[vllm#48390](vllm-project/vllm#48390) changed LM
head apply contract | Added `_apply_head` routing to
`quant_method.apply` on v0.25.1 and `super()._apply_head` on main. |
|
`vllm_ascend/patch/worker/__init__.py`<br>`vllm_ascend/patch/worker/patch_v2/patch_eagle_speculator.py`<br>`vllm_ascend/worker/v2/spec_decode/eagle/aclgraph.py`<br>`vllm_ascend/worker/v2/spec_decode/eagle/speculator.py`<br>`vllm_ascend/worker/v2/spec_decode/dflash/speculator.py`<br>`vllm_ascend/worker/v2/spec_decode/dspark/speculator.py`
| [vllm#48261](vllm-project/vllm#48261) /
[vllm#48167](vllm-project/vllm#48167) unified
`PrefillSpeculatorCudaGraphManager` + `DecodeSpeculatorCudaGraphManager`
-> `SpeculatorCudaGraphManager`; `capture()` parameterless; `set_attn`
new params; `dflash_causal` -> `_group_causal` | Gated MRV2 spec decode
imports to main-only. Consolidated Eagle managers into
`EagleAclGraphManager`. Simplified DFlash/DSpark `capture()`. Added
`target_input_buffers`/`target_attn_groups` to `set_attn`. Renamed
`dflash_causal` -> `_group_causal` in DSpark
`build_draft_attn_metadatas` (DFlash was already updated; DSpark was
missed). |
|
`tests/ut/ops/test_gdn_layerwise_kv.py`<br>`tests/ut/ops/a2/test_gdn_layerwise_kv.py`
| [vllm#46998](vllm-project/vllm#46998) changed
GDN `forward` from `(hidden_states, output)` -> `(hidden_states) ->
Tensor` | Added `_run_gdn_forward` wrapper handling both calling
conventions. |
| `tests/e2e/conftest.py` |
[vllm#48549](vllm-project/vllm#48549) removed
`swap_space` from `LLM` | Removed from `VllmRunner` and `DPVllmRunner`.
|
| `tests/e2e/pull_request/one_card/model_runner_v2/test_basic.py` |
[vllm#48261](vllm-project/vllm#48261) MRV2 spec
decode main-only | Added `_SKIP_V025_MRV2_SPEC_DECODE` skip marker. |
| `tests/ut/core/test_recompute_scheduler.py` | Upstream main's
`_free_request` accesses `self.ec_connector` (new attribute) | Added
`scheduler.ec_connector = None` to test setup — scheduler is constructed
via `__new__` (bypasses `__init__`), so the attribute must be mocked
explicitly. |
|
`tests/ut/distributed/ascend_store/test_config_data.py`<br>`tests/ut/distributed/ascend_store/test_coordinator.py`<br>`tests/ut/distributed/ascend_store/test_pool_worker.py`
| [vllm#45939](vllm-project/vllm#45939) /
[vllm#46384](vllm-project/vllm#46384) | Updated
block hash assertions to terminal hash. Added `hash_block_size`.
Version-gated mock returns. |
| `tests/ut/ops/test_vocab_parallel_embedding.py` |
[vllm#48390](vllm-project/vllm#48390) | Replaced
`patch()` with context manager for `set_current_vllm_config`. |
| `tests/ut/patch/platform/test_prefix_cache_cp_patches.py` |
[vllm#47782](vllm-project/vllm#47782) | Removed
obsolete `num_prompt_tokens` kwarg. |
| `tests/ut/patch/worker/test_patch_qwen3_5_mtp.py` |
[vllm#48429](vllm-project/vllm#48429) | Set
`use_attn_reduce_scatter_for_moe = False` on mock layer. |
| `tests/ut/test_compressed_prefix_cache.py` |
[vllm#46384](vllm-project/vllm#46384) /
[vllm#47782](vllm-project/vllm#47782) |
Version-gated unpacking. Added `test_find_longest_cache_hit_per_group`.
|
| `vllm_ascend/patch/__init__.py` | - | Updated HunyuanVL comment block.
|

- vLLM version: v0.25.1
- vLLM main:
vllm-project/vllm@54503ec
---------
Signed-off-by: wjunLu <wjunlu217@gmail.com>
Signed-off-by: hfadzxy <starmoon_zhang@163.com>
Co-authored-by: hfadzxy <starmoon_zhang@163.com>
crb2nu added a commit to flexinfer/flexinfer that referenced this pull request Aug 2, 2026
The ngram speculative-decoding re-canary (#66) is blocked on an engine-fatal
assertion in `_prepare_inputs` at ~198k-token prefills. Upstream research
(2026-08-01) traced the hybrid-GDN half of that crash to
`_mamba_block_aligned_split`, which could floor `num_new_tokens` to 0 or
amplify a negative to `-block_size` on every chunk of a chunked prefill. It was
rewritten with an explicit clamp by vllm-project/vllm#47782 and first ships in
v0.26.0.

Bump the `gfx1100-native-w4a16` profile's base image, source ref, and published
tag (`v023` -> `v026`), plus the matching build-args, BuildKit cache/output
refs, and the pinned strings in the runtime-patch contract gate that assert
them. The verify gate keeps its >= 0.23.0 floor: the RDNA3 W4A16 kernel and the
`flexinfer_qwen35_text` plugin registration are both still required and
compatible at v0.26.0.

Runtime only. Digest promotion to the lanes and the speculation flip stay
separate canaries per #66's plan, so `deploy/models/*` is deliberately
untouched. The base-image change invalidates the BuildKit inline cache, so
expect a full-length rebuild.

Refs: services/flexinfer#66

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
ningjingbengxiaohai pushed a commit to vllm-project/vllm-ascend that referenced this pull request Aug 3, 2026
…uple (#13275)

## Based on

This PR is based on
#13235 and re-submits
the same production-code change on an independent branch. No additional
functional changes are included.

## Why

The copied `BalanceScheduler.schedule()` implementation still expects
two return values from `KVCacheManager.get_computed_blocks()`. After the
upstream vLLM API change, the method returns three values, causing the
balance-scheduling path to fail at runtime.

## Upstream change

Upstream PR: vllm-project/vllm#47782

The upstream change extends `get_computed_blocks()` to return:

```python
(computed_blocks, num_computed_tokens, shared_prefix_boundary)
```

## Downstream changes

- Update `vllm_ascend/patch/platform/patch_balance_schedule.py` to
unpack the third value into `request.shared_prefix_boundary`.

## Test plan

- Full validation is delegated to PR CI.

## User-facing change

No.

- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@d02df74

---------

Signed-off-by: zhao-stack <80399320+zhao-stack@users.noreply.github.com>
xiayingqing pushed a commit to xiayingqing/vllm-ascend that referenced this pull request Aug 3, 2026
…uple (vllm-project#13275)

## Based on

This PR is based on
vllm-project#13235 and re-submits
the same production-code change on an independent branch. No additional
functional changes are included.

## Why

The copied `BalanceScheduler.schedule()` implementation still expects
two return values from `KVCacheManager.get_computed_blocks()`. After the
upstream vLLM API change, the method returns three values, causing the
balance-scheduling path to fail at runtime.

## Upstream change

Upstream PR: vllm-project/vllm#47782

The upstream change extends `get_computed_blocks()` to return:

```python
(computed_blocks, num_computed_tokens, shared_prefix_boundary)
```

## Downstream changes

- Update `vllm_ascend/patch/platform/patch_balance_schedule.py` to
unpack the third value into `request.shared_prefix_boundary`.

## Test plan

- Full validation is delegated to PR CI.

## User-facing change

No.

- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@d02df74

---------

Signed-off-by: zhao-stack <80399320+zhao-stack@users.noreply.github.com>
HMCCMH pushed a commit to hotTea123/vllm-ascend that referenced this pull request Aug 12, 2026
…uple (vllm-project#13275)

## Based on

This PR is based on
vllm-project#13235 and re-submits
the same production-code change on an independent branch. No additional
functional changes are included.

## Why

The copied `BalanceScheduler.schedule()` implementation still expects
two return values from `KVCacheManager.get_computed_blocks()`. After the
upstream vLLM API change, the method returns three values, causing the
balance-scheduling path to fail at runtime.

## Upstream change

Upstream PR: vllm-project/vllm#47782

The upstream change extends `get_computed_blocks()` to return:

```python
(computed_blocks, num_computed_tokens, shared_prefix_boundary)
```

## Downstream changes

- Update `vllm_ascend/patch/platform/patch_balance_schedule.py` to
unpack the third value into `request.shared_prefix_boundary`.

## Test plan

- Full validation is delegated to PR CI.

## User-facing change

No.

- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@d02df74

---------

Signed-off-by: zhao-stack <80399320+zhao-stack@users.noreply.github.com>
MmMmaru pushed a commit to jiaqi-lee/vllm-ascend that referenced this pull request Aug 19, 2026
### What this PR does / why we need it?

Adapt vllm-ascend to vLLM main commits up to July 17.

### Changes

| Files | Upstream vLLM change | vllm-ascend adaptation |
|-------|---------------------|------------------------|
|
`vllm_ascend/_310p/ops/fla/idex.py`<br>`vllm_ascend/_310p/ops/fla/l2norm.py`<br>`vllm_ascend/ops/bailing_moe_linear_attn.py`<br>`vllm_ascend/ops/gdn.py`<br>`vllm_ascend/ops/triton/fla/chunk.py`<br>`vllm_ascend/patch/worker/patch_idex_310.py`<br>`vllm_ascend/patch/worker/patch_triton.py`<br>`tests/e2e/nightly/.../test_fused_recurrent_gated_delta_rule.py`<br>`tests/e2e/nightly/.../test_fused_sigmoid_gating_delta_rule.py`<br>`tests/ut/ops/test_gdn_attn_builder.py`
| [vllm#48500](vllm-project/vllm#48500)
relocated flash-linear-attention from
`vllm.third_party.flash_linear_attention` to
`vllm.model_executor.layers.fla` | Version-gated all FLA imports and
monkey-patch sites with `vllm_version_is("0.25.1")` |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/config_data.py`
| [vllm#45939](vllm-project/vllm#45939) replaced
SHA-256 rehashed grouped block hashes with chained fine-grained hashes
(terminal hash identifies complete block) | Removed
`_rehash_block_hash_group` and associated constants. `get_block_hashes`
returns `block_hashes[idx + scale_factor - 1]` instead of computing
compound hash. |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/coordinator.py`<br>`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/pool_scheduler.py`<br>`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/pool_worker.py`
| [vllm#45939](vllm-project/vllm#45939) added
`hash_block_size` to block pool API;
[vllm#46384](vllm-project/vllm#46384) renamed
`prefix_match_unit` -> `hash_block_size`, moved hash resolution into
individual managers | `ExternalCachedBlockPool` takes `hash_block_size`.
`_find_longest_cache_hit` return version-gated. `block_hashes_for_spec`
no-op on main. `prefix_match_unit` fallback on main. |
|
`vllm_ascend/core/recompute_scheduler.py`<br>`vllm_ascend/core/scheduler_profiling_chunk.py`<br>`vllm_ascend/core/single_type_kv_cache_manager.py`
| [vllm#46384](vllm-project/vllm#46384) /
[vllm#47782](vllm-project/vllm#47782) changed
`get_computed_blocks` 3-tuple -> 2-tuple, `get_num_blocks_to_allocate`
params, `find_longest_cache_hit` return | Version-gated unpacking with
`cast`. `num_tokens_main_model` made optional. `find_longest_cache_hit`
returns blocks-only on v0.25.1 vs `(blocks, hit_length)` on main. |
| `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` |
[vllm#46384](vllm-project/vllm#46384) /
[vllm#47782](vllm-project/vllm#47782)
restructured return contracts; added partial Mamba hash hit support on
main | Added `enable_partial_hash_hits` (main-only). Added
`_cache_hit_alignment_tokens`. `find_longest_cache_hit` tracks per-group
lengths, uses `cdiv`. `find_longest_cache_hit_per_group` simplified to
single-pass; returns per-group `(blocks, lengths)`. Fixed v0.25.1
`hit_length` over-counting: use physical `spec.block_size` (not
`_get_effective_block_size()` which includes `compress_ratio`) when
multiplying by physical block count — affected both
`find_longest_cache_hit` and `find_longest_cache_hit_per_group`; without
this, MLA models (DS-V2-Lite/V3/V4) got 4x inflated hit length ->
scheduler skipped non-cached tokens -> segfault. |
| `vllm_ascend/patch/platform/patch_mamba_manager.py` |
[vllm#46384](vllm-project/vllm#46384) /
[vllm#47782](vllm-project/vllm#47782) changed
method signatures | Version-gated `find_longest_cache_hit` (delegates to
super on main) and `get_num_blocks_to_allocate` (optional params). |
| `vllm_ascend/patch/worker/patch_qwen3_5.py` |
[vllm#47006](vllm-project/vllm#47006) changed
Qwen3-Next sequence-parallel gather contracts | Added
`_ascend_all_gather_hidden_and_residual` (main-only). Gated
`Qwen3_5DecoderLayer.forward` to v0.25.1. |
| `vllm_ascend/ops/vocab_parallel_embedding.py` |
[vllm#48390](vllm-project/vllm#48390) changed LM
head apply contract | Added `_apply_head` routing to
`quant_method.apply` on v0.25.1 and `super()._apply_head` on main. |
|
`vllm_ascend/patch/worker/__init__.py`<br>`vllm_ascend/patch/worker/patch_v2/patch_eagle_speculator.py`<br>`vllm_ascend/worker/v2/spec_decode/eagle/aclgraph.py`<br>`vllm_ascend/worker/v2/spec_decode/eagle/speculator.py`<br>`vllm_ascend/worker/v2/spec_decode/dflash/speculator.py`<br>`vllm_ascend/worker/v2/spec_decode/dspark/speculator.py`
| [vllm#48261](vllm-project/vllm#48261) /
[vllm#48167](vllm-project/vllm#48167) unified
`PrefillSpeculatorCudaGraphManager` + `DecodeSpeculatorCudaGraphManager`
-> `SpeculatorCudaGraphManager`; `capture()` parameterless; `set_attn`
new params; `dflash_causal` -> `_group_causal` | Gated MRV2 spec decode
imports to main-only. Consolidated Eagle managers into
`EagleAclGraphManager`. Simplified DFlash/DSpark `capture()`. Added
`target_input_buffers`/`target_attn_groups` to `set_attn`. Renamed
`dflash_causal` -> `_group_causal` in DSpark
`build_draft_attn_metadatas` (DFlash was already updated; DSpark was
missed). |
|
`tests/ut/ops/test_gdn_layerwise_kv.py`<br>`tests/ut/ops/a2/test_gdn_layerwise_kv.py`
| [vllm#46998](vllm-project/vllm#46998) changed
GDN `forward` from `(hidden_states, output)` -> `(hidden_states) ->
Tensor` | Added `_run_gdn_forward` wrapper handling both calling
conventions. |
| `tests/e2e/conftest.py` |
[vllm#48549](vllm-project/vllm#48549) removed
`swap_space` from `LLM` | Removed from `VllmRunner` and `DPVllmRunner`.
|
| `tests/e2e/pull_request/one_card/model_runner_v2/test_basic.py` |
[vllm#48261](vllm-project/vllm#48261) MRV2 spec
decode main-only | Added `_SKIP_V025_MRV2_SPEC_DECODE` skip marker. |
| `tests/ut/core/test_recompute_scheduler.py` | Upstream main's
`_free_request` accesses `self.ec_connector` (new attribute) | Added
`scheduler.ec_connector = None` to test setup — scheduler is constructed
via `__new__` (bypasses `__init__`), so the attribute must be mocked
explicitly. |
|
`tests/ut/distributed/ascend_store/test_config_data.py`<br>`tests/ut/distributed/ascend_store/test_coordinator.py`<br>`tests/ut/distributed/ascend_store/test_pool_worker.py`
| [vllm#45939](vllm-project/vllm#45939) /
[vllm#46384](vllm-project/vllm#46384) | Updated
block hash assertions to terminal hash. Added `hash_block_size`.
Version-gated mock returns. |
| `tests/ut/ops/test_vocab_parallel_embedding.py` |
[vllm#48390](vllm-project/vllm#48390) | Replaced
`patch()` with context manager for `set_current_vllm_config`. |
| `tests/ut/patch/platform/test_prefix_cache_cp_patches.py` |
[vllm#47782](vllm-project/vllm#47782) | Removed
obsolete `num_prompt_tokens` kwarg. |
| `tests/ut/patch/worker/test_patch_qwen3_5_mtp.py` |
[vllm#48429](vllm-project/vllm#48429) | Set
`use_attn_reduce_scatter_for_moe = False` on mock layer. |
| `tests/ut/test_compressed_prefix_cache.py` |
[vllm#46384](vllm-project/vllm#46384) /
[vllm#47782](vllm-project/vllm#47782) |
Version-gated unpacking. Added `test_find_longest_cache_hit_per_group`.
|
| `vllm_ascend/patch/__init__.py` | - | Updated HunyuanVL comment block.
|

- vLLM version: v0.25.1
- vLLM main:
vllm-project/vllm@54503ec
---------
Signed-off-by: wjunLu <wjunlu217@gmail.com>
Signed-off-by: hfadzxy <starmoon_zhang@163.com>
Co-authored-by: hfadzxy <starmoon_zhang@163.com>
MmMmaru pushed a commit to jiaqi-lee/vllm-ascend that referenced this pull request Aug 19, 2026
…uple (vllm-project#13275)

## Based on

This PR is based on
vllm-project#13235 and re-submits
the same production-code change on an independent branch. No additional
functional changes are included.

## Why

The copied `BalanceScheduler.schedule()` implementation still expects
two return values from `KVCacheManager.get_computed_blocks()`. After the
upstream vLLM API change, the method returns three values, causing the
balance-scheduling path to fail at runtime.

## Upstream change

Upstream PR: vllm-project/vllm#47782

The upstream change extends `get_computed_blocks()` to return:

```python
(computed_blocks, num_computed_tokens, shared_prefix_boundary)
```

## Downstream changes

- Update `vllm_ascend/patch/platform/patch_balance_schedule.py` to
unpack the third value into `request.shared_prefix_boundary`.

## Test plan

- Full validation is delegated to PR CI.

## User-facing change

No.

- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@d02df74

---------

Signed-off-by: zhao-stack <80399320+zhao-stack@users.noreply.github.com>
shiqiangA pushed a commit to shiqiangA/vllm-ascend that referenced this pull request Aug 20, 2026
### What this PR does / why we need it?

Adapt vllm-ascend to vLLM main commits up to July 17.

### Changes

| Files | Upstream vLLM change | vllm-ascend adaptation |
|-------|---------------------|------------------------|
|
`vllm_ascend/_310p/ops/fla/idex.py`<br>`vllm_ascend/_310p/ops/fla/l2norm.py`<br>`vllm_ascend/ops/bailing_moe_linear_attn.py`<br>`vllm_ascend/ops/gdn.py`<br>`vllm_ascend/ops/triton/fla/chunk.py`<br>`vllm_ascend/patch/worker/patch_idex_310.py`<br>`vllm_ascend/patch/worker/patch_triton.py`<br>`tests/e2e/nightly/.../test_fused_recurrent_gated_delta_rule.py`<br>`tests/e2e/nightly/.../test_fused_sigmoid_gating_delta_rule.py`<br>`tests/ut/ops/test_gdn_attn_builder.py`
| [vllm#48500](vllm-project/vllm#48500)
relocated flash-linear-attention from
`vllm.third_party.flash_linear_attention` to
`vllm.model_executor.layers.fla` | Version-gated all FLA imports and
monkey-patch sites with `vllm_version_is("0.25.1")` |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/config_data.py`
| [vllm#45939](vllm-project/vllm#45939) replaced
SHA-256 rehashed grouped block hashes with chained fine-grained hashes
(terminal hash identifies complete block) | Removed
`_rehash_block_hash_group` and associated constants. `get_block_hashes`
returns `block_hashes[idx + scale_factor - 1]` instead of computing
compound hash. |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/coordinator.py`<br>`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/pool_scheduler.py`<br>`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/pool_worker.py`
| [vllm#45939](vllm-project/vllm#45939) added
`hash_block_size` to block pool API;
[vllm#46384](vllm-project/vllm#46384) renamed
`prefix_match_unit` -> `hash_block_size`, moved hash resolution into
individual managers | `ExternalCachedBlockPool` takes `hash_block_size`.
`_find_longest_cache_hit` return version-gated. `block_hashes_for_spec`
no-op on main. `prefix_match_unit` fallback on main. |
|
`vllm_ascend/core/recompute_scheduler.py`<br>`vllm_ascend/core/scheduler_profiling_chunk.py`<br>`vllm_ascend/core/single_type_kv_cache_manager.py`
| [vllm#46384](vllm-project/vllm#46384) /
[vllm#47782](vllm-project/vllm#47782) changed
`get_computed_blocks` 3-tuple -> 2-tuple, `get_num_blocks_to_allocate`
params, `find_longest_cache_hit` return | Version-gated unpacking with
`cast`. `num_tokens_main_model` made optional. `find_longest_cache_hit`
returns blocks-only on v0.25.1 vs `(blocks, hit_length)` on main. |
| `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` |
[vllm#46384](vllm-project/vllm#46384) /
[vllm#47782](vllm-project/vllm#47782)
restructured return contracts; added partial Mamba hash hit support on
main | Added `enable_partial_hash_hits` (main-only). Added
`_cache_hit_alignment_tokens`. `find_longest_cache_hit` tracks per-group
lengths, uses `cdiv`. `find_longest_cache_hit_per_group` simplified to
single-pass; returns per-group `(blocks, lengths)`. Fixed v0.25.1
`hit_length` over-counting: use physical `spec.block_size` (not
`_get_effective_block_size()` which includes `compress_ratio`) when
multiplying by physical block count — affected both
`find_longest_cache_hit` and `find_longest_cache_hit_per_group`; without
this, MLA models (DS-V2-Lite/V3/V4) got 4x inflated hit length ->
scheduler skipped non-cached tokens -> segfault. |
| `vllm_ascend/patch/platform/patch_mamba_manager.py` |
[vllm#46384](vllm-project/vllm#46384) /
[vllm#47782](vllm-project/vllm#47782) changed
method signatures | Version-gated `find_longest_cache_hit` (delegates to
super on main) and `get_num_blocks_to_allocate` (optional params). |
| `vllm_ascend/patch/worker/patch_qwen3_5.py` |
[vllm#47006](vllm-project/vllm#47006) changed
Qwen3-Next sequence-parallel gather contracts | Added
`_ascend_all_gather_hidden_and_residual` (main-only). Gated
`Qwen3_5DecoderLayer.forward` to v0.25.1. |
| `vllm_ascend/ops/vocab_parallel_embedding.py` |
[vllm#48390](vllm-project/vllm#48390) changed LM
head apply contract | Added `_apply_head` routing to
`quant_method.apply` on v0.25.1 and `super()._apply_head` on main. |
|
`vllm_ascend/patch/worker/__init__.py`<br>`vllm_ascend/patch/worker/patch_v2/patch_eagle_speculator.py`<br>`vllm_ascend/worker/v2/spec_decode/eagle/aclgraph.py`<br>`vllm_ascend/worker/v2/spec_decode/eagle/speculator.py`<br>`vllm_ascend/worker/v2/spec_decode/dflash/speculator.py`<br>`vllm_ascend/worker/v2/spec_decode/dspark/speculator.py`
| [vllm#48261](vllm-project/vllm#48261) /
[vllm#48167](vllm-project/vllm#48167) unified
`PrefillSpeculatorCudaGraphManager` + `DecodeSpeculatorCudaGraphManager`
-> `SpeculatorCudaGraphManager`; `capture()` parameterless; `set_attn`
new params; `dflash_causal` -> `_group_causal` | Gated MRV2 spec decode
imports to main-only. Consolidated Eagle managers into
`EagleAclGraphManager`. Simplified DFlash/DSpark `capture()`. Added
`target_input_buffers`/`target_attn_groups` to `set_attn`. Renamed
`dflash_causal` -> `_group_causal` in DSpark
`build_draft_attn_metadatas` (DFlash was already updated; DSpark was
missed). |
|
`tests/ut/ops/test_gdn_layerwise_kv.py`<br>`tests/ut/ops/a2/test_gdn_layerwise_kv.py`
| [vllm#46998](vllm-project/vllm#46998) changed
GDN `forward` from `(hidden_states, output)` -> `(hidden_states) ->
Tensor` | Added `_run_gdn_forward` wrapper handling both calling
conventions. |
| `tests/e2e/conftest.py` |
[vllm#48549](vllm-project/vllm#48549) removed
`swap_space` from `LLM` | Removed from `VllmRunner` and `DPVllmRunner`.
|
| `tests/e2e/pull_request/one_card/model_runner_v2/test_basic.py` |
[vllm#48261](vllm-project/vllm#48261) MRV2 spec
decode main-only | Added `_SKIP_V025_MRV2_SPEC_DECODE` skip marker. |
| `tests/ut/core/test_recompute_scheduler.py` | Upstream main's
`_free_request` accesses `self.ec_connector` (new attribute) | Added
`scheduler.ec_connector = None` to test setup — scheduler is constructed
via `__new__` (bypasses `__init__`), so the attribute must be mocked
explicitly. |
|
`tests/ut/distributed/ascend_store/test_config_data.py`<br>`tests/ut/distributed/ascend_store/test_coordinator.py`<br>`tests/ut/distributed/ascend_store/test_pool_worker.py`
| [vllm#45939](vllm-project/vllm#45939) /
[vllm#46384](vllm-project/vllm#46384) | Updated
block hash assertions to terminal hash. Added `hash_block_size`.
Version-gated mock returns. |
| `tests/ut/ops/test_vocab_parallel_embedding.py` |
[vllm#48390](vllm-project/vllm#48390) | Replaced
`patch()` with context manager for `set_current_vllm_config`. |
| `tests/ut/patch/platform/test_prefix_cache_cp_patches.py` |
[vllm#47782](vllm-project/vllm#47782) | Removed
obsolete `num_prompt_tokens` kwarg. |
| `tests/ut/patch/worker/test_patch_qwen3_5_mtp.py` |
[vllm#48429](vllm-project/vllm#48429) | Set
`use_attn_reduce_scatter_for_moe = False` on mock layer. |
| `tests/ut/test_compressed_prefix_cache.py` |
[vllm#46384](vllm-project/vllm#46384) /
[vllm#47782](vllm-project/vllm#47782) |
Version-gated unpacking. Added `test_find_longest_cache_hit_per_group`.
|
| `vllm_ascend/patch/__init__.py` | - | Updated HunyuanVL comment block.
|

- vLLM version: v0.25.1
- vLLM main:
vllm-project/vllm@54503ec
---------
Signed-off-by: wjunLu <wjunlu217@gmail.com>
Signed-off-by: hfadzxy <starmoon_zhang@163.com>
Co-authored-by: hfadzxy <starmoon_zhang@163.com>
shiqiangA pushed a commit to shiqiangA/vllm-ascend that referenced this pull request Aug 20, 2026
…uple (vllm-project#13275)

## Based on

This PR is based on
vllm-project#13235 and re-submits
the same production-code change on an independent branch. No additional
functional changes are included.

## Why

The copied `BalanceScheduler.schedule()` implementation still expects
two return values from `KVCacheManager.get_computed_blocks()`. After the
upstream vLLM API change, the method returns three values, causing the
balance-scheduling path to fail at runtime.

## Upstream change

Upstream PR: vllm-project/vllm#47782

The upstream change extends `get_computed_blocks()` to return:

```python
(computed_blocks, num_computed_tokens, shared_prefix_boundary)
```

## Downstream changes

- Update `vllm_ascend/patch/platform/patch_balance_schedule.py` to
unpack the third value into `request.shared_prefix_boundary`.

## Test plan

- Full validation is delegated to PR CI.

## User-facing change

No.

- vLLM version: v0.26.0
- vLLM main:
vllm-project/vllm@d02df74

---------

Signed-off-by: zhao-stack <80399320+zhao-stack@users.noreply.github.com>
puririshi98 pushed a commit to puririshi98/vllm that referenced this pull request Aug 31, 2026
…vllm-project#47782)

Signed-off-by: Nick Hill <nickhill123@gmail.com>
(cherry picked from commit 8ac8375)
(cherry picked from commit f9fda36f49b1314b03031e9263f9a5be82b38598)
Leetrytry pushed a commit to Leetrytry/vllm-ascend that referenced this pull request Sep 11, 2026
### What this PR does / why we need it?

Adapt vllm-ascend to vLLM main commits up to July 17.

### Changes

| Files | Upstream vLLM change | vllm-ascend adaptation |
|-------|---------------------|------------------------|
|
`vllm_ascend/_310p/ops/fla/idex.py`<br>`vllm_ascend/_310p/ops/fla/l2norm.py`<br>`vllm_ascend/ops/bailing_moe_linear_attn.py`<br>`vllm_ascend/ops/gdn.py`<br>`vllm_ascend/ops/triton/fla/chunk.py`<br>`vllm_ascend/patch/worker/patch_idex_310.py`<br>`vllm_ascend/patch/worker/patch_triton.py`<br>`tests/e2e/nightly/.../test_fused_recurrent_gated_delta_rule.py`<br>`tests/e2e/nightly/.../test_fused_sigmoid_gating_delta_rule.py`<br>`tests/ut/ops/test_gdn_attn_builder.py`
| [vllm#48500](vllm-project/vllm#48500)
relocated flash-linear-attention from
`vllm.third_party.flash_linear_attention` to
`vllm.model_executor.layers.fla` | Version-gated all FLA imports and
monkey-patch sites with `vllm_version_is("0.25.1")` |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/config_data.py`
| [vllm#45939](vllm-project/vllm#45939) replaced
SHA-256 rehashed grouped block hashes with chained fine-grained hashes
(terminal hash identifies complete block) | Removed
`_rehash_block_hash_group` and associated constants. `get_block_hashes`
returns `block_hashes[idx + scale_factor - 1]` instead of computing
compound hash. |
|
`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/coordinator.py`<br>`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/pool_scheduler.py`<br>`vllm_ascend/distributed/kv_transfer/kv_pool/ascend_store/pool_worker.py`
| [vllm#45939](vllm-project/vllm#45939) added
`hash_block_size` to block pool API;
[vllm#46384](vllm-project/vllm#46384) renamed
`prefix_match_unit` -> `hash_block_size`, moved hash resolution into
individual managers | `ExternalCachedBlockPool` takes `hash_block_size`.
`_find_longest_cache_hit` return version-gated. `block_hashes_for_spec`
no-op on main. `prefix_match_unit` fallback on main. |
|
`vllm_ascend/core/recompute_scheduler.py`<br>`vllm_ascend/core/scheduler_profiling_chunk.py`<br>`vllm_ascend/core/single_type_kv_cache_manager.py`
| [vllm#46384](vllm-project/vllm#46384) /
[vllm#47782](vllm-project/vllm#47782) changed
`get_computed_blocks` 3-tuple -> 2-tuple, `get_num_blocks_to_allocate`
params, `find_longest_cache_hit` return | Version-gated unpacking with
`cast`. `num_tokens_main_model` made optional. `find_longest_cache_hit`
returns blocks-only on v0.25.1 vs `(blocks, hit_length)` on main. |
| `vllm_ascend/patch/platform/patch_kv_cache_coordinator.py` |
[vllm#46384](vllm-project/vllm#46384) /
[vllm#47782](vllm-project/vllm#47782)
restructured return contracts; added partial Mamba hash hit support on
main | Added `enable_partial_hash_hits` (main-only). Added
`_cache_hit_alignment_tokens`. `find_longest_cache_hit` tracks per-group
lengths, uses `cdiv`. `find_longest_cache_hit_per_group` simplified to
single-pass; returns per-group `(blocks, lengths)`. Fixed v0.25.1
`hit_length` over-counting: use physical `spec.block_size` (not
`_get_effective_block_size()` which includes `compress_ratio`) when
multiplying by physical block count — affected both
`find_longest_cache_hit` and `find_longest_cache_hit_per_group`; without
this, MLA models (DS-V2-Lite/V3/V4) got 4x inflated hit length ->
scheduler skipped non-cached tokens -> segfault. |
| `vllm_ascend/patch/platform/patch_mamba_manager.py` |
[vllm#46384](vllm-project/vllm#46384) /
[vllm#47782](vllm-project/vllm#47782) changed
method signatures | Version-gated `find_longest_cache_hit` (delegates to
super on main) and `get_num_blocks_to_allocate` (optional params). |
| `vllm_ascend/patch/worker/patch_qwen3_5.py` |
[vllm#47006](vllm-project/vllm#47006) changed
Qwen3-Next sequence-parallel gather contracts | Added
`_ascend_all_gather_hidden_and_residual` (main-only). Gated
`Qwen3_5DecoderLayer.forward` to v0.25.1. |
| `vllm_ascend/ops/vocab_parallel_embedding.py` |
[vllm#48390](vllm-project/vllm#48390) changed LM
head apply contract | Added `_apply_head` routing to
`quant_method.apply` on v0.25.1 and `super()._apply_head` on main. |
|
`vllm_ascend/patch/worker/__init__.py`<br>`vllm_ascend/patch/worker/patch_v2/patch_eagle_speculator.py`<br>`vllm_ascend/worker/v2/spec_decode/eagle/aclgraph.py`<br>`vllm_ascend/worker/v2/spec_decode/eagle/speculator.py`<br>`vllm_ascend/worker/v2/spec_decode/dflash/speculator.py`<br>`vllm_ascend/worker/v2/spec_decode/dspark/speculator.py`
| [vllm#48261](vllm-project/vllm#48261) /
[vllm#48167](vllm-project/vllm#48167) unified
`PrefillSpeculatorCudaGraphManager` + `DecodeSpeculatorCudaGraphManager`
-> `SpeculatorCudaGraphManager`; `capture()` parameterless; `set_attn`
new params; `dflash_causal` -> `_group_causal` | Gated MRV2 spec decode
imports to main-only. Consolidated Eagle managers into
`EagleAclGraphManager`. Simplified DFlash/DSpark `capture()`. Added
`target_input_buffers`/`target_attn_groups` to `set_attn`. Renamed
`dflash_causal` -> `_group_causal` in DSpark
`build_draft_attn_metadatas` (DFlash was already updated; DSpark was
missed). |
|
`tests/ut/ops/test_gdn_layerwise_kv.py`<br>`tests/ut/ops/a2/test_gdn_layerwise_kv.py`
| [vllm#46998](vllm-project/vllm#46998) changed
GDN `forward` from `(hidden_states, output)` -> `(hidden_states) ->
Tensor` | Added `_run_gdn_forward` wrapper handling both calling
conventions. |
| `tests/e2e/conftest.py` |
[vllm#48549](vllm-project/vllm#48549) removed
`swap_space` from `LLM` | Removed from `VllmRunner` and `DPVllmRunner`.
|
| `tests/e2e/pull_request/one_card/model_runner_v2/test_basic.py` |
[vllm#48261](vllm-project/vllm#48261) MRV2 spec
decode main-only | Added `_SKIP_V025_MRV2_SPEC_DECODE` skip marker. |
| `tests/ut/core/test_recompute_scheduler.py` | Upstream main's
`_free_request` accesses `self.ec_connector` (new attribute) | Added
`scheduler.ec_connector = None` to test setup — scheduler is constructed
via `__new__` (bypasses `__init__`), so the attribute must be mocked
explicitly. |
|
`tests/ut/distributed/ascend_store/test_config_data.py`<br>`tests/ut/distributed/ascend_store/test_coordinator.py`<br>`tests/ut/distributed/ascend_store/test_pool_worker.py`
| [vllm#45939](vllm-project/vllm#45939) /
[vllm#46384](vllm-project/vllm#46384) | Updated
block hash assertions to terminal hash. Added `hash_block_size`.
Version-gated mock returns. |
| `tests/ut/ops/test_vocab_parallel_embedding.py` |
[vllm#48390](vllm-project/vllm#48390) | Replaced
`patch()` with context manager for `set_current_vllm_config`. |
| `tests/ut/patch/platform/test_prefix_cache_cp_patches.py` |
[vllm#47782](vllm-project/vllm#47782) | Removed
obsolete `num_prompt_tokens` kwarg. |
| `tests/ut/patch/worker/test_patch_qwen3_5_mtp.py` |
[vllm#48429](vllm-project/vllm#48429) | Set
`use_attn_reduce_scatter_for_moe = False` on mock layer. |
| `tests/ut/test_compressed_prefix_cache.py` |
[vllm#46384](vllm-project/vllm#46384) /
[vllm#47782](vllm-project/vllm#47782) |
Version-gated unpacking. Added `test_find_longest_cache_hit_per_group`.
|
| `vllm_ascend/patch/__init__.py` | - | Updated HunyuanVL comment block.
|

- vLLM version: v0.25.1
- vLLM main:
vllm-project/vllm@54503ec
---------
Signed-off-by: wjunLu <wjunlu217@gmail.com>
Signed-off-by: hfadzxy <starmoon_zhang@163.com>
Co-authored-by: hfadzxy <starmoon_zhang@163.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

kv-connector ready ONLY add when PR is ready to merge/full CI is needed v1

4 participants