PofoliaPofolia ile paylaşıldı

Journal of Computing in Civil Engineering· 2026Q1

Görsel-Dil Modelleri Robot Montaj Verisi Üretiyor

Vision-Language Model-Based Demonstrators for Imitation Learning in Construction Robotic Timber Assembly

Lei Huang, Xinhe Yang, Qingyu Yan, Zhengbo Zou

Kısa özet

Segment Anything Model (SAM) ve görsel-dil modellerini (VLM'ler) kullanan bir çerçeve, robot taklit öğrenimi için otonom olarak gösterim verileri üreterek, damıtmadan sonra ahşap montaj görevlerinde %100 başarı elde ediyor.

Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.

Ana noktalar

  • Bir çerçeve, nesne segmentasyonu için SAM'i ve eylem önerisi için VLM'leri kullanarak robot gösterim verilerini otonom olarak üretir.
  • Sistem, taklit öğrenimi için veri kümeleri oluşturmak üzere başarılı bölümleri otomatik olarak filtreler.
  • Damıtılmış politikalar, çeşitli pertürbasyon seviyelerinde ahşap montaj görevlerinde %100 başarı elde etti.
  • 100% başarı için gereken gösterim sayısı görev zorluğuna göre değişir; düşük varyasyon için %70, daha büyük pertürbasyonlar için ise %90 gerektirir.

Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.

Özet (abstract)

Abstract Intelligent construction robots are deemed the future of on-site construction for improved productivity and safety. To automate the construction process, imitation learning (IL) has been adopted to train construction robots in a repertoire of tasks. However, collecting demonstrations for robots to imitate from usually requires teleoperation setup and devices, such as virtual reality (VR) and gloves. To autonomously and efficiently generate demonstration data for imitation learning in construction tasks, we propose a large foundation model-based demonstration-generation framework, in which the Segment Anything Model (SAM) is used to extract geometric representations of objects of interest and a vision-language model (VLM), conditioned on mark-based visual prompting and languages, produces high-level action suggestions. As the demonstrator framework is imperfect and is computation-intensive to be fine-tuned, only successful episodes are retained automatically and structured into demonstrations for distilling a lightweight text-conditioned robot policy via behavioral cloning (BC). We evaluate the framework on a UR10 robot arm in Isaac Sim for a language-conditioned timber assembly task with frame pose randomization. The autonomous demonstrator achieves success rates ranging from 76.9% (small perturbations) to 33.7% (largest perturbations) but can be rolled out to curate balanced demonstration data sets. Policies distilled from these demonstrations attain 100% success across all three perturbation ranges and two placement scenarios, while retaining up to 70% success under modest out-of-distribution (OOD) frame poses. An ablation on the number of demonstrations shows that the distilled policies reach 100% success with only 70% of the data in low-variation settings but require about 90% of the data to achieve 100% success under larger workspace perturbations.

Yazarların özeti; kaynağından alınmıştır. Journal of Computing in Civil Engineering, 2026 · DOI ↗

ÇıkarımlarPremium
Makaleye SorÜcretsiz hesapla

Ücretsiz hesapla devam et

Makaleye Sor ile bu makaleye günde 3 soru ücretsiz; makaleyi kaydet, kaynakçasını al, ilgi alanına göre her gün yeni özetler. Çıkarımlar Premium.

Web'de ücretsiz devam et

Google ya da Apple hesabınla giriş; kart istemez. Bu makaleye geri dönersin.

Telefonda:

Alan: Yapı ve İnşaat

Building and ConstructionEngineering