PofoliaShared via Pofolia

ACM Transactions on Reconfigurable Technology and Systems· 2026Q2

OLA+: Multi-FPGA O ver l ay A ccelerator System for Fully Homomorphic Encryption

Yang Yang, Rajgopal Kannan, Viktor K. Prasanna

Short summary

OLA+ is a novel multi-FPGA overlay accelerator system that achieves up to 8835x speedup for end-to-end HE ML training and inference compared to CPUs, eliminating manual FPGA programming via a Python interface and compiler.

AI-generated from the title and abstract; the full text is not read.

Key points

  • OLA+ is a multi-FPGA overlay accelerator system for Fully Homomorphic Encryption (FHE).
  • It uses a Python interface and compiler to automate FPGA programming for FHE.
  • Key optimizations include compile-time partitioning, asynchronous dataflow, and latency-aware instruction scheduling.
  • On 8 Alveo U280 FPGAs, OLA+ achieves up to 8835x speedup over CPUs and 14.6x over GPUs for HE ML tasks.

AI-generated from the title and abstract; the full text is not read.

Abstract

Fully Homomorphic Encryption (FHE) is a promising technique for privacy preserving computation. However, homomorphically encrypted operations are orders of magnitude slower than the corresponding unencrypted operations due to high computation and memory bandwidth requirements. FPGAs are attractive platforms for accelerating FHE workloads, but manually programming FPGAs is challenging, as different FHE parameters and operations require different mapping strategies. We propose OLA+, a scalable overlay accelerator system for FHE, designed to run efficiently on multi-FPGA platforms. OLA+ eliminates the need for manual and time-consuming FPGA programming by providing a Python-based interface and a compiler that efficiently maps OLA+ programs to hardware instructions for execution. The hardware architecture and instruction set are co-designed to accelerate common FHE primitives, while the compiler maps all FHE operations to these primitives for efficient execution. We propose a compile-time partitioning strategy that parallelizes computation both across multiple ciphertexts and within individual ciphertexts, enabling efficient execution of OLA+ programs on multiple FPGA accelerators. OLA+ features two optimizations to address memory bandwidth challenges in FHE computation. First, we propose an asynchronous dataflow execution model where the compiler manages data processing order and the architecture enforces it at run time. This approach enables guaranteed data reuse via on-chip SRAMs. Second, we design latency-aware instruction scheduling in the compiler to reduce data reuse distance and overlap data transfers with computation. We implement the overlay accelerator on AMD Alveo U280 FPGAs. We evaluate the effectiveness of the proposed overlay accelerator across a range of FHE parameters by executing various FHE operations, FHE linear algebra benchmarks, and end-to-end ML training and inference. Experimental results show that running end-to-end HE ML training and inference tasks on OLA+ with 8 overlay accelerators achieves speedups of up to \(8835\times\) and \(14.6\times\) over state-of-the-art CPU and GPU implementations, respectively.

The authors' abstract, as published at the source. ACM Transactions on Reconfigurable Technology and Systems, 2026 · DOI ↗

TakeawaysPremium
Ask the paperFree account

Continue with a free account

Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.

Continue free on the web

Sign in with Google or Apple; no card needed. You come back to this paper.

On your phone:

Field: Hardware and Architecture

Hardware and ArchitectureComputer Science