Multimodal 3D Gaze–Object Localization for Shared Human–Machine Situational Awareness in Tactical Environments
Report Number:
ARL-TR-10183
September 9, 2025
Approved for public release: distribution is unlimited.
Author(s):
Soorya Saravanan and Russell Cohen Hoffing
Abstract:Modern military operations require seamless human–machine integration to enhance situational awareness. This research, part of the Distributed Information for Enhanced Squad Lethality (DIESL) program, introduces a novel 3D gaze–object localization system that creates machine-readable human-generated environmental information for shared squad awareness without manual input from users. The system fuses multimodal data, from including Soldier eye-tracking, first-person 2D video feeds, camera pose, and visual embeddings from the CLIP vision-language model, to uniquely identify objects observed by different team members across time and space. The central hypothesis is that combining geometric data with visual features resolves ambiguities inherent in computer vision-only approaches. To test this, the system was evaluated using data from field experiments with Soldiers. Results demonstrate that fusing geometric and visual features significantly improves gaze–object localization accuracy compared to geometric-only baselines. The resulting 3D localized objects enable human operators and autonomous systems to share a synchronized understanding of the environment, thus improving coordination and mission effectiveness. This work validates multimodal sensor fusion from opportunistically sensed human-centric data as foundational for future advances in human–machine teaming.
