PofoliaShared via Pofolia

ACM Transactions on Architecture and Code Optimization· 2026Q2

LightZK: An Efficient GPU Acceleration Framework for Large-Scale Zero-Knowledge Proofs

Biran Lu, Qiang-Sheng Hua, Xuanhua Shi, Hai Jin

Short summary

LightZK, a new GPU framework, enables large-scale zero-knowledge proof generation on single, memory-constrained nodes by employing a tile-based execution model and addressing GPU underutilization.

AI-generated from the title and abstract; the full text is not read.

Key points

  • LightZK introduces a tile-based execution model for GPU acceleration of large-scale zero-knowledge proofs.
  • Micro-architectural diagnosis reveals GPU underutilization in MSM, leading to kernel optimizations yielding 1.5x-2.5x speedups.
  • LightZK enables large-scale NTT execution and employs a three-tier pipeline for efficient DRAM-VRAM and disk-DRAM-VRAM data transfer.
  • The framework successfully generates proofs at 2^30 scale on a single node, outperforming existing GPU frameworks that exceed memory limits.

AI-generated from the title and abstract; the full text is not read.

Abstract

Zero-Knowledge Proofs (ZKPs), particularly zk-SNARKs, have become a cornerstone for privacy-preserving protocols and blockchain scalability. However, the generation of proofs remains computationally prohibitive due to the immense overhead of underlying cryptographic primitives, namely Multi-Scalar Multiplication (MSM) and Fast Number Theoretic Transform (NTT). As applications scale up to billion-scale constraints, conventional GPU-accelerated frameworks tailored for smaller instances suffer from overwhelming memory usage, invalid architectural assumptions, and suboptimal end-to-end orchestration. Driven by the motivation to complete massive proof generation on a single memory-constrained node, this paper presents LightZK, an efficient GPU acceleration framework for large-scale zero-knowledge proofs based on a tile-based execution model. LightZK not only aims to support ultra-large-scale ZKP generation, but also strives to improve the practical performance of proof generation. LightZK contributes a micro-architectural diagnosis of MSM on GPUs: hardware performance counters expose the register-pressure and low instruction-level-parallelism mismatches that under-utilize the GPU during bucket-merge computation, and targeted kernel optimizations derived from this diagnosis yield a 1.5 × –2.5 × MSM speedup over the SOTA baseline. LightZK further makes large-scale NTT executable where the SOTA GPU baseline fails, via storage and structural optimizations that deliver nontrivial acceleration (2.2 × –2.5 × over a functional baseline). At the systemic level, LightZK builds on a data-dependency graph (DDG) analysis of the Groth16 proving flow to drive a three-tier pipeline orchestration: a hand-crafted DRAM-VRAM pipeline for Poly-phase that overlaps sparse matrix operations with data transfer, and a producer-consumer disk-DRAM-VRAM pipeline for Group-phase that treats each MSM tile as an independently schedulable unit. Experimental results demonstrate that LightZK achieves a 2.8 × end-to-end performance improvement over the leading executable GPU baseline at standard scales. Crucially, while existing GPU frameworks exceed available memory at billion-scale, LightZK successfully sustains 2 30 -scale proof generation with near-linear scalability on a single node.

The authors' abstract, as published at the source. ACM Transactions on Architecture and Code Optimization, 2026 · DOI ↗

TakeawaysPremium
Ask the paperFree account

Continue with a free account

Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.

Continue free on the web

Sign in with Google or Apple; no card needed. You come back to this paper.

On your phone:

Field: Hardware and Architecture

Hardware and ArchitectureComputer Science