ACM Transactions on Reconfigurable Technology and Systems· 2026Q2
OLA+: Tam Homomorfik Şifreleme için Çoklu-FPGA Yer Paylaşımı Hızlandırıcı Sistemi
OLA+: Multi-FPGA O ver l ay A ccelerator System for Fully Homomorphic Encryption
- 0atıf
- Q2SCImago
- 2026yıl
Kısa özet
OLA+, Tam Homomorfik Şifreleme (FHE) için tasarlanmış, CPU'lara kıyasla uçtan uca FHE Makine Öğrenmesi eğitim ve çıkarımında 8835 kata kadar hızlanma sağlayan ve Python arayüzü ile derleyicisi sayesinde manuel FPGA programlamasını ortadan kaldıran yeni bir çoklu-FPGA yer paylaşımı hızlandırıcı sistemidir.
Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.
Ana noktalar
- OLA+, Tam Homomorfik Şifreleme (FHE) için tasarlanmış bir çoklu-FPGA yer paylaşımı hızlandırıcı sistemidir.
- FHE için FPGA programlamasını otomatikleştirmek üzere bir Python arayüzü ve derleyici kullanır.
- Temel optimizasyonlar arasında derleme zamanı bölümlemesi, asenkron veri akışı ve gecikmeye duyarlı komut zamanlaması bulunur.
- 8 Alveo U280 FPGA üzerinde OLA+, FHE Makine Öğrenmesi görevleri için CPU'lara göre 8835 kata ve GPU'lara göre 14.6 kata kadar hızlanma sağlar.
Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.
Özet (abstract)
Fully Homomorphic Encryption (FHE) is a promising technique for privacy preserving computation. However, homomorphically encrypted operations are orders of magnitude slower than the corresponding unencrypted operations due to high computation and memory bandwidth requirements. FPGAs are attractive platforms for accelerating FHE workloads, but manually programming FPGAs is challenging, as different FHE parameters and operations require different mapping strategies. We propose OLA+, a scalable overlay accelerator system for FHE, designed to run efficiently on multi-FPGA platforms. OLA+ eliminates the need for manual and time-consuming FPGA programming by providing a Python-based interface and a compiler that efficiently maps OLA+ programs to hardware instructions for execution. The hardware architecture and instruction set are co-designed to accelerate common FHE primitives, while the compiler maps all FHE operations to these primitives for efficient execution. We propose a compile-time partitioning strategy that parallelizes computation both across multiple ciphertexts and within individual ciphertexts, enabling efficient execution of OLA+ programs on multiple FPGA accelerators. OLA+ features two optimizations to address memory bandwidth challenges in FHE computation. First, we propose an asynchronous dataflow execution model where the compiler manages data processing order and the architecture enforces it at run time. This approach enables guaranteed data reuse via on-chip SRAMs. Second, we design latency-aware instruction scheduling in the compiler to reduce data reuse distance and overlap data transfers with computation. We implement the overlay accelerator on AMD Alveo U280 FPGAs. We evaluate the effectiveness of the proposed overlay accelerator across a range of FHE parameters by executing various FHE operations, FHE linear algebra benchmarks, and end-to-end ML training and inference. Experimental results show that running end-to-end HE ML training and inference tasks on OLA+ with 8 overlay accelerators achieves speedups of up to \(8835\times\) and \(14.6\times\) over state-of-the-art CPU and GPU implementations, respectively.
Yazarların özeti; kaynağından alınmıştır. ACM Transactions on Reconfigurable Technology and Systems, 2026 · DOI ↗
Ücretsiz hesapla devam et
Makaleye Sor ile bu makaleye günde 3 soru ücretsiz; makaleyi kaydet, kaynakçasını al, ilgi alanına göre her gün yeni özetler. Çıkarımlar Premium.
Web'de ücretsiz devam etGoogle ya da Apple hesabınla giriş; kart istemez. Bu makaleye geri dönersin.
Telefonda:
Alan: Donanım ve Mimari
Hardware and ArchitectureComputer Science