A Vision-Language-Driven UxV System for Safety Monitoring

  • Abidi, Syed Murtaza Hussain; 
  • Raza, Syed Muhammad; 
  • Park, Jinsu; 
  • Shin, Soo Young

초록

This paper present a Vision-Language Model (VLM)-driven unmanned vehicle (UxV) system designed for realtime safety monitoring in high-risk industrial environments such as construction and oil and gas operations. The proposed system integrates advanced computer vision and natural language processing (NLP) to generate context-aware textual descriptions of safety conditions. It employs unmanned aerial vehicles (UAVs) for wide-area surveillance and unmanned ground vehicles (UGVs) for close-range inspection, thereby enhancing spatial coverage and operational flexibility. To support VLM fine-tuning, a custom dataset comprising 2,000 safety-related images was developed, enabling dynamic hazard detection and assessment of regulatory compliance. Unlike conventional object detection methods, the system leverages zero-shot learning to facilitate rapid adaptation to new environments with minimal retraining. Preliminary evaluations indicate that the system-generated captions closely align with human annotations, demonstrating high interpretability and real-time applicability. The framework exhibits strong scalability, offering a proactive approach to hazard detection and safety compliance monitoring in complex industrial settings.

제목
A Vision-Language-Driven UxV System for Safety Monitoring
저자
Abidi, Syed Murtaza Hussain; Raza, Syed Muhammad; Park, Jinsu; Shin, Soo Young
DOI
10.1109/ICUFN65838.2025.11169839
발행일
2025-07-11
학회명
16th International Conference on Ubiquitous and Future Networks-ICUFN-Annual
개최지
Lisbon, PORTUGAL
개최국가
미국
학회 개최일
2025-07-08 ~ 2025-07-11

파일 다운로드

Thumbnail