PofoliaPofolia ile paylaşıldı

Systems· 2026Q2

Reaktif Müdahaleden Proaktif Azaltmaya: DRL Envanter Kontrolünde Öncül Risk Sinyallerinin Değeri

From Reactive Response to Proactive Mitigation: The Value of Precursor Risk Signals in DRL Inventory Control

Hyuksoo Han, Yong Won Seo

Kısa özet

Tedarik zinciri aksaklıklarının gözlemlenebilir öncül sinyallerini durumuna dahil eden yeni bir risk algılama Derin Pekiştirmeli Öğrenme (DRL) çerçevesi, kusurlu sinyallerle bile reaktif DRL ve kalibre edilmiş (s, Q) politikalara kıyasla önemli maliyet tasarrufu sağlıyor.

Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.

Özet (abstract)

Modern supply chains face abrupt supply disruptions, under which the stable lead time assumption behind conventional inventory control no longer holds. Most replenishment models, including existing deep reinforcement learning (DRL) formulations, act as feedback controllers. A disruption is treated as an unobservable event, and correction begins only after deliveries are already delayed. In practice, however, disruptions are often preceded by observable precursor signals. This paper develops a risk sensing DRL framework that incorporates such signals into the state of a proximal policy optimization (PPO) agent, introducing feedforward control into the replenishment decision. The signal quality (detection rate) and the predictive horizon are treated as explicit design parameters. In a two-echelon system with graded disruptions, the risk sensing policy achieves substantial cost savings against both a per-environment calibrated (s, Q) policy and an identically trained reactive DRL policy, even with an imperfect signal. The required horizon is short, as horizons longer than needed bring no further gain. In contrast, the savings depend critically on the signal quality, as a meaningful gain over the calibrated benchmark requires a signal that detects well over half of upcoming disruptions. The advantage is greatest where severe disruptions are infrequent. The learned policy operates a state-dependent reorder point that stays low in calm conditions and rises with the severity and proximity of a predicted threat. A transparent rule driven by the same signal captures about 83% of the gain, indicating that most of the value comes from the information itself rather than from the learning method. These results quantify the value of advance supply information in inventory control and show how it supports the transition from reactive response to proactive mitigation.

Yazarların özeti; kaynağından alınmıştır. Systems, 2026 · DOI ↗

ÇıkarımlarUygulamada
Ana noktalarUygulamada
Makaleye SorUygulamada

Devamı Pofolia uygulamasında

Çıkarımlar, ana noktalar ve makaleye soru sorma; ilgi alanına göre her gün yeni özetler. Ücretsiz.

Web'de giriş yaparak aç

Alan: Strateji ve Yönetim

Strategy and ManagementBusiness, Management and Accounting