ACM Transactions on Architecture and Code Optimization· 2026Q2
PCR: Düşük Gecikmeli RAG Sunumu İçin Önbelleğe Almayı Geliştirilmiş Önbellek Yeniden Kullanım Sistemi
PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
- 0atıf
- Q2SCImago
- 2026yıl
Kısa özet
Yeni bir sistem olan PCR, SSD'lerden DRAM'e KV önbellek verilerini proaktif olarak önceden getirerek I/O darboğazlarını gizleyerek RAG sunum gecikmesini azaltır.
Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.
Ana noktalar
- PCR, GPU, DRAM ve SSD için birleşik bir Çok Katmanlı Önek-Ağaç Önbellek Yöneticisi sunar.
- Kuyruk Güdümlü Çalışma Zamanlayıcısı, istek kuyruklarını kullanarak SSD verilerini proaktif olarak DRAM'e önceden getirir.
- Gecikme Gizleyen KV Transfer Hattı, veri hareketini model hesaplamasıyla örtüştürür.
- Değerlendirmeler, PCR'nin Llama3.1-8B için vLLM'ye göre 2.45 kat hızlanma sağladığını ve kuyruk gecikmesini iyileştirdiğini göstermektedir.
Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.
Özet (abstract)
Retrieval-Augmented Generation (RAG) significantly improves Large Language Models (LLMs) but introduces massive input sequences that severely bottleneck the prefill stage. While KV-cache reuse reduces redundant computation for shared document prefixes, the reusable KV working set in RAG serving can exceed GPU memory capacity, requiring KV chunks to be retained across host DRAM and SSDs. However, naive multi-tier storage extensions suffer from severe I/O bottlenecks, suboptimal eviction, and high CPU-GPU data transfer overheads, which often negate the latency benefits of cache reuse. In this paper, we propose PCR, a P refetch-enhanced C ache R euse system for low-latency RAG serving. PCR transforms passive SSD-backed storage into an active, latency-hiding memory hierarchy through three core modules: (1) a Multi-Tier Prefix-Tree Cache Manager that unifies the organization and tracking of reusable KV chunks across the entire memory hierarchy of GPU, host DRAM, and SSD; (2) a Queue-Guided Runtime Scheduler that leverages the post-retrieval waiting queue as a look-ahead signal to proactively protect hot chunks and prefetch SSD-resident data into host DRAM before execution; and (3) a Latency-Hiding KV Transfer Pipeline that overlaps fine-grained PCIe data movement and asynchronous SSD operations with model computation. Extensive evaluations across diverse models and RAG workloads show that PCR reduces TTFT over the evaluated KV-cache reuse systems in most settings, with a 2.45 × speedup over vLLM for Llama3.1-8B on A6000 under Workload 1 at 1.0 request/s, while maintaining lower tail latency under high load.
Yazarların özeti; kaynağından alınmıştır. ACM Transactions on Architecture and Code Optimization, 2026 · DOI ↗
Ücretsiz hesapla devam et
Makaleye Sor ile bu makaleye günde 3 soru ücretsiz; makaleyi kaydet, kaynakçasını al, ilgi alanına göre her gün yeni özetler. Çıkarımlar Premium.
Web'de ücretsiz devam etGoogle ya da Apple hesabınla giriş; kart istemez. Bu makaleye geri dönersin.
Telefonda:
Alan: Donanım ve Mimari
Hardware and ArchitectureComputer Science