ACM Transactions on Architecture and Code Optimization· 2026Q2
PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
- 0citations
- Q2SCImago
- 2026year
Short summary
PCR, a new system, reduces RAG serving latency by proactively prefetching KV cache data from SSDs to DRAM, hiding I/O bottlenecks.
AI-generated from the title and abstract; the full text is not read.
Key points
- PCR introduces a unified Multi-Tier Prefix-Tree Cache Manager for GPU, DRAM, and SSD.
- A Queue-Guided Runtime Scheduler proactively prefetches SSD data to DRAM using request queues.
- A Latency-Hiding KV Transfer Pipeline overlaps data movement with model computation.
- Evaluations show PCR achieves a 2.45x speedup over vLLM for Llama3.1-8B and improves tail latency.
AI-generated from the title and abstract; the full text is not read.
Abstract
Retrieval-Augmented Generation (RAG) significantly improves Large Language Models (LLMs) but introduces massive input sequences that severely bottleneck the prefill stage. While KV-cache reuse reduces redundant computation for shared document prefixes, the reusable KV working set in RAG serving can exceed GPU memory capacity, requiring KV chunks to be retained across host DRAM and SSDs. However, naive multi-tier storage extensions suffer from severe I/O bottlenecks, suboptimal eviction, and high CPU-GPU data transfer overheads, which often negate the latency benefits of cache reuse. In this paper, we propose PCR, a P refetch-enhanced C ache R euse system for low-latency RAG serving. PCR transforms passive SSD-backed storage into an active, latency-hiding memory hierarchy through three core modules: (1) a Multi-Tier Prefix-Tree Cache Manager that unifies the organization and tracking of reusable KV chunks across the entire memory hierarchy of GPU, host DRAM, and SSD; (2) a Queue-Guided Runtime Scheduler that leverages the post-retrieval waiting queue as a look-ahead signal to proactively protect hot chunks and prefetch SSD-resident data into host DRAM before execution; and (3) a Latency-Hiding KV Transfer Pipeline that overlaps fine-grained PCIe data movement and asynchronous SSD operations with model computation. Extensive evaluations across diverse models and RAG workloads show that PCR reduces TTFT over the evaluated KV-cache reuse systems in most settings, with a 2.45 × speedup over vLLM for Llama3.1-8B on A6000 under Workload 1 at 1.0 request/s, while maintaining lower tail latency under high load.
The authors' abstract, as published at the source. ACM Transactions on Architecture and Code Optimization, 2026 · DOI ↗
Continue with a free account
Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.
Continue free on the webSign in with Google or Apple; no card needed. You come back to this paper.
On your phone:
Field: Hardware and Architecture
Hardware and ArchitectureComputer Science