PofoliaShared via Pofolia

Systems· 2026Q2

From Reactive Response to Proactive Mitigation: The Value of Precursor Risk Signals in DRL Inventory Control

Hyuksoo Han, Yong Won Seo

Short summary

A new risk-sensing Deep Reinforcement Learning (DRL) framework, incorporating precursor supply disruption signals into its state, achieves substantial cost savings compared to reactive DRL and calibrated (s, Q) policies, even with imperfect signals.

AI-generated from the title and abstract; the full text is not read.

Abstract

Modern supply chains face abrupt supply disruptions, under which the stable lead time assumption behind conventional inventory control no longer holds. Most replenishment models, including existing deep reinforcement learning (DRL) formulations, act as feedback controllers. A disruption is treated as an unobservable event, and correction begins only after deliveries are already delayed. In practice, however, disruptions are often preceded by observable precursor signals. This paper develops a risk sensing DRL framework that incorporates such signals into the state of a proximal policy optimization (PPO) agent, introducing feedforward control into the replenishment decision. The signal quality (detection rate) and the predictive horizon are treated as explicit design parameters. In a two-echelon system with graded disruptions, the risk sensing policy achieves substantial cost savings against both a per-environment calibrated (s, Q) policy and an identically trained reactive DRL policy, even with an imperfect signal. The required horizon is short, as horizons longer than needed bring no further gain. In contrast, the savings depend critically on the signal quality, as a meaningful gain over the calibrated benchmark requires a signal that detects well over half of upcoming disruptions. The advantage is greatest where severe disruptions are infrequent. The learned policy operates a state-dependent reorder point that stays low in calm conditions and rises with the severity and proximity of a predicted threat. A transparent rule driven by the same signal captures about 83% of the gain, indicating that most of the value comes from the information itself rather than from the learning method. These results quantify the value of advance supply information in inventory control and show how it supports the transition from reactive response to proactive mitigation.

The authors' abstract, as published at the source. Systems, 2026 · DOI ↗

TakeawaysIn the app
Key pointsIn the app
Ask the paperIn the app

The rest is in the Pofolia app

Takeaways, key points and questions to the paper; new summaries every day for your field. Free.

Sign in on the web to open

Field: Strategy and Management

Strategy and ManagementBusiness, Management and Accounting