RFC: Prefix-Aware Delta KV Cache Transfer for Disaggregated Prefill/Decode Serving #7861
liuhuijiayou
started this conversation in
General
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Hi @nvda-mesharma @tedzhouhk and the Dynamo community,
We are the Kuaishou (快手) infrastructure team working on large-scale LLM serving. We've been following the Dynamo disagg roadmap (#2436) closely and would like to implement Enhancement 1 from that roadmap:
We've done a detailed design analysis and would love to contribute this to the project. This RFC describes the problem, root causes, and a phased implementation plan.
Summary
In disaggregated prefill/decode serving, when a Decode worker already has part of the request's prefix KV cache, the Prefill worker should transfer only the missing KV blocks — not all of them. This significantly reduces KV transfer latency and bandwidth, especially for long-context and shared-prefix workloads.
Motivation
The Problem
Currently, Dynamo's disagg mode transfers all KV cache for every request — even if the Decode worker already has an identical prefix cached.
A typical example:
Scale of Impact
Using Llama-70B (GQA, 8 KV heads, 128 head_dim, FP16) as an example:
On a 100Gbps RDMA network:
Root Cause Analysis: Three Broken Links
There are 3 broken links in the current codebase that prevent delta transfer:
Broken Link 1: Decode Router has no awareness of D worker's cache
In
lib/llm/src/kv_router/prefill_router/types.rs,build_decode_router_overridesets:overlap_score_weight = 0causes the Decode Router's Indexer to be set toNone(lib/llm/src/kv_router/indexer.rs):KV event subscription is also skipped (
lib/kv-router/src/scheduling/config.rs):Result: The Decode Router is completely unaware of D worker KV cache state.
Broken Link 2: KV events don't distinguish "locally computed" vs "received via transfer"
In
lib/kv-router/src/zmq_wire.rs, KV events only have 3 types:BlockStored { block_hashes, token_ids, ... }BlockRemoved { block_hashes }AllBlocksClearedThere's no field indicating whether a block was computed locally by prefill or received via transfer from a P worker.
Broken Link 3: PrefillRouter doesn't pass D's cache info to P
In
lib/llm/src/kv_router/prefill_router/mod.rs, thegenerate()flow:In step 3, the prefill request contains no information about "which blocks D already has cached." P has no knowledge of D's cache state and can only perform full transfer.
Proposed Design
Core Idea
Phase 1: Enable Decode Router Indexer (fix Broken Links 1 & 2)
Change 1a: Don't force
overlap_score_weight=0for Decode RouterChange 1b: Ensure D worker's transferred KV triggers
BlockStoredeventsprefix_cachingis enabled, transferred blocks enter prefix cache → events fired ✓Phase 2: Reorder scheduling — D before P (fix Broken Link 3)
Change 2a: PrefillRouter queries Decode Router overlap before selecting P
Change 2b: Attach D's cache info to the prefill request
Phase 3: P Worker Implements Delta Transfer
D worker stitches together local cache + transferred data:
Phase 4 (Optional): P reads from D's cache
This is Enhancement 2 from the roadmap: when D has more cached blocks than P, P can read from D to reduce its own compute. We propose this as a follow-up RFC.
Implementation Plan
lib/llm/src/kv_router/prefill_router/types.rsoverlap_score_weight=0for Decode Routerlib/kv-router/src/scheduling/config.rsdecode_overlap_score_weightconfig paramBlockStoredlib/kv-router/src/zmq_wire.rssource: KvSourcefield (Local/Transferred)lib/llm/src/kv_router/prefill_router/mod.rslib/llm/src/protocols/common/preprocessor.rsDecodeCacheInfostructlib/llm/src/kv_router/push_router.rscomponents/src/dynamo/backends/vllm/decode_cache_info, execute delta transfercomponents/src/dynamo/backends/sglang/Open Questions
Scheduling order latency: "D first, then P" adds one overlap query latency (~10-50μs). Can this be parallelized with P selection in the bootstrap path?
Cache eviction during transfer: D may evict blocks under memory pressure while P is transferring. We need a "lock" mechanism or at least detection.
Bootstrap path compatibility: Currently P and D start almost simultaneously. Delta transfer requires knowing D's cache state before P starts transferring — this may require adjusting bootstrap timing.
Multiple D worker candidates: If D0 has 50 blocks cached and D1 has 80, should we prefer D1 (more cache but possibly higher load)? With
overlap_score_weight > 0, the existing cost function already handles this balance.Block hash consistency: Are block hashes computed by P and D consistent for the same token sequence? In Dynamo, block hashes are computed by the Router side via XXH3 (independent of the backend engine), so consistency is guaranteed.
Expected Impact
Prerequisite: D worker prefix cache hit rate must be high enough. This depends on the Decode Router routing similar requests to the same D worker — which Phase 1 (enabling the Decode Router Indexer) directly addresses.
We are planning to start with Phase 1 (enabling the Decode Router Indexer) and would appreciate feedback on:
Happy to open individual issues for each phase once we get alignment on the overall direction. Looking forward to the discussion!
All reactions