PofoliaPofolia ile paylaşıldı

ACM Transactions on Reconfigurable Technology and Systems· 2026Q2

OLA+: Tam Homomorfik Şifreleme için Çoklu-FPGA Yer Paylaşımı Hızlandırıcı Sistemi

OLA+: Multi-FPGA O ver l ay A ccelerator System for Fully Homomorphic Encryption

Yang Yang, Rajgopal Kannan, Viktor K. Prasanna

Kısa özet

OLA+, Tam Homomorfik Şifreleme (FHE) için tasarlanmış, CPU'lara kıyasla uçtan uca FHE Makine Öğrenmesi eğitim ve çıkarımında 8835 kata kadar hızlanma sağlayan ve Python arayüzü ile derleyicisi sayesinde manuel FPGA programlamasını ortadan kaldıran yeni bir çoklu-FPGA yer paylaşımı hızlandırıcı sistemidir.

Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.

Ana noktalar

  • OLA+, Tam Homomorfik Şifreleme (FHE) için tasarlanmış bir çoklu-FPGA yer paylaşımı hızlandırıcı sistemidir.
  • FHE için FPGA programlamasını otomatikleştirmek üzere bir Python arayüzü ve derleyici kullanır.
  • Temel optimizasyonlar arasında derleme zamanı bölümlemesi, asenkron veri akışı ve gecikmeye duyarlı komut zamanlaması bulunur.
  • 8 Alveo U280 FPGA üzerinde OLA+, FHE Makine Öğrenmesi görevleri için CPU'lara göre 8835 kata ve GPU'lara göre 14.6 kata kadar hızlanma sağlar.

Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.

Özet (abstract)

Fully Homomorphic Encryption (FHE) is a promising technique for privacy preserving computation. However, homomorphically encrypted operations are orders of magnitude slower than the corresponding unencrypted operations due to high computation and memory bandwidth requirements. FPGAs are attractive platforms for accelerating FHE workloads, but manually programming FPGAs is challenging, as different FHE parameters and operations require different mapping strategies. We propose OLA+, a scalable overlay accelerator system for FHE, designed to run efficiently on multi-FPGA platforms. OLA+ eliminates the need for manual and time-consuming FPGA programming by providing a Python-based interface and a compiler that efficiently maps OLA+ programs to hardware instructions for execution. The hardware architecture and instruction set are co-designed to accelerate common FHE primitives, while the compiler maps all FHE operations to these primitives for efficient execution. We propose a compile-time partitioning strategy that parallelizes computation both across multiple ciphertexts and within individual ciphertexts, enabling efficient execution of OLA+ programs on multiple FPGA accelerators. OLA+ features two optimizations to address memory bandwidth challenges in FHE computation. First, we propose an asynchronous dataflow execution model where the compiler manages data processing order and the architecture enforces it at run time. This approach enables guaranteed data reuse via on-chip SRAMs. Second, we design latency-aware instruction scheduling in the compiler to reduce data reuse distance and overlap data transfers with computation. We implement the overlay accelerator on AMD Alveo U280 FPGAs. We evaluate the effectiveness of the proposed overlay accelerator across a range of FHE parameters by executing various FHE operations, FHE linear algebra benchmarks, and end-to-end ML training and inference. Experimental results show that running end-to-end HE ML training and inference tasks on OLA+ with 8 overlay accelerators achieves speedups of up to \(8835\times\) and \(14.6\times\) over state-of-the-art CPU and GPU implementations, respectively.

Yazarların özeti; kaynağından alınmıştır. ACM Transactions on Reconfigurable Technology and Systems, 2026 · DOI ↗

ÇıkarımlarPremium
Makaleye SorÜcretsiz hesapla

Ücretsiz hesapla devam et

Makaleye Sor ile bu makaleye günde 3 soru ücretsiz; makaleyi kaydet, kaynakçasını al, ilgi alanına göre her gün yeni özetler. Çıkarımlar Premium.

Web'de ücretsiz devam et

Google ya da Apple hesabınla giriş; kart istemez. Bu makaleye geri dönersin.

Telefonda:

Alan: Donanım ve Mimari

Hardware and ArchitectureComputer Science