Automation in Construction· 2026Q1
Vision-encoder-integrated lifting tracker for cost-effective crane operations in modular construction
- 0citations
- Q1SCImago
- 2026year
Short summary
A new vision-encoder-integrated lifting tracker (VE-LIFT) uses a monocular camera and crane encoders to track modules during lifting, achieving LiDAR-comparable accuracy (sub-0.5m tracking, 98.65%/98.93% mAP) at lower cost.
AI-generated from the title and abstract; the full text is not read.
Abstract
Cost-effective tracking of modules during lifting remains challenging in modular construction (MC), where existing solutions rely on densely deployed, costly sensors such as Light Detection and Ranging (LiDAR). This paper presents a vision-encoder-integrated lifting tracker (VE-LIFT) that reuses two low-cost on-site assets: a monocular camera and the crane's factory-installed encoders. The camera estimates a cylindrical bounding volume (CBV) enveloping the module through a three-stage pipeline: a You Only Look Once (YOLO)v8-Pose model detects the module, lifting frame, and keypoints; Segment Anything Model (SAM) 3 refines keypoints via box and text prompts; and a hierarchical Perspective-n-Point solver derives the CBV dimensions. Encoders update the CBV position via Modbus-over-Ethernet. On a real-life MC project, VE-LIFT achieved 98.65%/98.93% mAP@0.5:0.95 for box/pose, over 22% RMSE reduction by SAM 3, and sub-0.5 m tracking during module lifting, suggesting LiDAR-comparable accuracy at lower cost. VE-LIFT delivers AI-based high-accuracy and cost-effective lifting tracking without additional dedicated sensors.
The authors' abstract, as published at the source. Automation in Construction, 2026 · DOI ↗
The rest is in the Pofolia app
Takeaways, key points and questions to the paper; new summaries every day for your field. Free.
Sign in on the web to openField: Building and Construction
Building and ConstructionEngineering