PofoliaShared via Pofolia

Smart Construction· 2026Q2

YOLO-SafeAttr: improving the ability to identify unsafe conditions of objects in construction scenes by using visual attribute representation

Yichuan Deng, Zile Deng, Hengyun Zhang, Hui Deng et al.

Short summary

A new deep learning method, YOLO-SafeAttr, improves construction safety by identifying unsafe object conditions (fall, strike, collapse, rollover) using visual attributes, achieving an F1-score of 0.82 on the VALD dataset.

AI-generated from the title and abstract; the full text is not read.

Key points

  • YOLO-SafeAttr integrates visual unsafe attributes (shape, state, context) to identify four risk types: fall, strike, collapse, and rollover.
  • The VALD dataset was created by extending the SODA dataset to include visual unsafe attributes for training and evaluation.
  • The YOLOv11-based framework achieved an mAP@50 of 84.8%, recall of 91%, and F1-score of 0.82.
  • Attention mechanisms improved performance by 21.1% (mAP@50), 33.6% (recall), and 0.175 (F1-score) over baseline models.

AI-generated from the title and abstract; the full text is not read.

Abstract

Unsafe conditions of objects are major contributors to construction accidents such as falls, strikes, and collapses. Although traditional inspections and existing vision methods can detect objects, they often fail to infer their potential risks. To overcome this limitation, this paper proposes a deep identification method integrating visual attribute representations. We introduce the high-level concept of visual unsafe attributes to describe hazards arising from an object’s shape, state, and environmental context. Based on accident text analysis, a visual attribute system covering four risk types—fall, strike, collapse, and rollover—is established, and the VALD dataset is built by extending SODA dataset. An integrated detection framework based on YOLOv11 is then developed to achieve end-to-end joint inference of object locations and unsafe attributes. The model is trained and evaluated on VALD. Experiments show mAP@50 of 84.8%, recall of 91%, and F1-score of 0.82, with improvements of 21.1%, 33.6%, and 0.175 over models without attention mechanisms. These results demonstrate that visual attribute representation and attention significantly enhance risk feature extraction. The resulting dynamic risk assessment system quantifies object-strike scenarios spatiotemporally, providing a real-time and interpretable safety monitoring solution for construction sites.

The authors' abstract, as published at the source. Smart Construction, 2026 · DOI ↗

TakeawaysPremium
Ask the paperFree account

Continue with a free account

Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.

Continue free on the web

Sign in with Google or Apple; no card needed. You come back to this paper.

On your phone:

Field: Radiological and Ultrasound Technology

Radiological and Ultrasound TechnologyHealth Professions