Journal of Computing in Civil Engineering· 2026Q1
Vision-Language Model-Based Demonstrators for Imitation Learning in Construction Robotic Timber Assembly
- 0citations
- Q1SCImago
- 2026year
Short summary
A framework using Segment Anything Model (SAM) and vision-language models (VLMs) autonomously generates demonstration data for robot imitation learning, achieving 100% success in timber assembly tasks after distillation.
AI-generated from the title and abstract; the full text is not read.
Key points
- A framework uses SAM for object segmentation and VLMs for action suggestion to autonomously generate robot demonstration data.
- The system automatically filters successful episodes to create datasets for imitation learning.
- Distilled policies achieved 100% success in timber assembly tasks across various perturbation levels.
- The number of demonstrations required for 100% success varies with task difficulty, needing 70% for low variation and 90% for larger perturbations.
AI-generated from the title and abstract; the full text is not read.
Abstract
Abstract Intelligent construction robots are deemed the future of on-site construction for improved productivity and safety. To automate the construction process, imitation learning (IL) has been adopted to train construction robots in a repertoire of tasks. However, collecting demonstrations for robots to imitate from usually requires teleoperation setup and devices, such as virtual reality (VR) and gloves. To autonomously and efficiently generate demonstration data for imitation learning in construction tasks, we propose a large foundation model-based demonstration-generation framework, in which the Segment Anything Model (SAM) is used to extract geometric representations of objects of interest and a vision-language model (VLM), conditioned on mark-based visual prompting and languages, produces high-level action suggestions. As the demonstrator framework is imperfect and is computation-intensive to be fine-tuned, only successful episodes are retained automatically and structured into demonstrations for distilling a lightweight text-conditioned robot policy via behavioral cloning (BC). We evaluate the framework on a UR10 robot arm in Isaac Sim for a language-conditioned timber assembly task with frame pose randomization. The autonomous demonstrator achieves success rates ranging from 76.9% (small perturbations) to 33.7% (largest perturbations) but can be rolled out to curate balanced demonstration data sets. Policies distilled from these demonstrations attain 100% success across all three perturbation ranges and two placement scenarios, while retaining up to 70% success under modest out-of-distribution (OOD) frame poses. An ablation on the number of demonstrations shows that the distilled policies reach 100% success with only 70% of the data in low-variation settings but require about 90% of the data to achieve 100% success under larger workspace perturbations.
The authors' abstract, as published at the source. Journal of Computing in Civil Engineering, 2026 · DOI ↗
Continue with a free account
Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.
Continue free on the webSign in with Google or Apple; no card needed. You come back to this paper.
On your phone:
Field: Building and Construction
Building and ConstructionEngineering