ACM Transactions on Architecture and Code Optimization· 2026Q2
EcoMLC: An Energy-Efficient Machine Learning Compilation Framework via Cache Reuse and Frequency Scaling
- 0citations
- Q2SCImago
- 2026year
Short summary
EcoMLC, a new machine learning compilation framework, achieves 43% average energy savings over TensorRT and vLLM by optimizing cache reuse and frequency scaling.
AI-generated from the title and abstract; the full text is not read.
Key points
- EcoMLC framework introduces a reverse scheduling policy and explicit cache priority to mitigate cache thrashing.
- It employs end-to-end frequency-sweeping analysis for optimal energy operating points.
- Achieves 43% average energy savings (J/request) over TensorRT and vLLM on NVIDIA GPUs.
- Achieves 26% average energy savings (J/token) over TensorRT and vLLM on NVIDIA GPUs.
- Demonstrates 33.9% energy reduction for LLMs with a 7% performance trade-off.
AI-generated from the title and abstract; the full text is not read.
Abstract
While modern machine learning compilers (MLCs) are pushing deep neural network (DNN) inference latency to its theoretical limits, the accompanying energy overhead has become a critical bottleneck for large-scale deployment. Excessive energy consumption not only increases electricity costs, but also affects the sustainability of edge platforms. Traditional performance-first optimizations often lead to significant energy waste, as they overlook the complex interplay between kernel configurations, cache behaviors and frequency scaling. This paper presents EcoMLC, an energy-efficient machine learning compilation framework. EcoMLC introduces a reverse scheduling policy combined with an explicit cache priority mechanism, mitigating cache thrashing inherent in standard least recently used (LRU) policies. EcoMLC employs an end-to-end frequency-sweeping analysis to identify optimal energy operating points and enables fine-grained trade-offs between performance and energy under diverse service level objective (SLO) constraints. Experimental results demonstrate that EcoMLC effectively optimizes various workloads, achieving an average of 43% and 26% energy savings (in J/request and J/token) over NVIDIA’s TensorRT and vLLM, and 22% energy savings on an AMD GPU compared to vLLM. For certain large language model (LLM) configurations, EcoMLC achieves a 33.9% reduction in energy consumption at the cost of a marginal 7% decline in performance.
The authors' abstract, as published at the source. ACM Transactions on Architecture and Code Optimization, 2026 · DOI ↗
Continue with a free account
Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.
Continue free on the webSign in with Google or Apple; no card needed. You come back to this paper.
On your phone:
Field: Hardware and Architecture
Hardware and ArchitectureComputer Science