A Comparative Benchmark of Open-Source and Proprietary Vision–Language Models for Military Image Assessment Tasks
Report Number:
ARL-TR-10174
September 3, 2025
Approved for public release: distribution is unlimited.
Author(s):
Jesse Barkley and Derrik Asher
Abstract:As open-source vision–language models (VLMs) gain traction in robotics, it remains unclear which models are suitable for resourceconstrained military platforms. This study evaluates the performance of seven VLMs from the Ollama framework, ranging from 3 billion (3b) to 12 billion (12b) parameters, on a two-part visual reasoning task involving modern Russian military vehicles. Each model was first prompted to produce a scene summary and then generate a doctrinal battle damage assessment (BDA) across 10 operational and 10 destroyed vehicles. The BDA prompt included a structured report with physical and functional damage assessments, task success determination, and a threat-level recommendation. To establish a benchmark, identical prompts and images were also assessed by GPT4o. Results showed that 12b models (gemma3:12b and llama3.2-vision) matched GPT-4o’s performance with 98.75% accuracy across graded criteria, albeit with significantly higher inference times. Medium-sized models like llava and minicpm-v showed moderate performance, achieving over 70% accuracy, but occasionally struggled with hallucinations and image detail. Smaller models (e.g., qwen2.5vl:3b, gemma3:4b, and llava-phi3) experienced a marked drop in accuracy and consistency. Overall, the study highlights a tradeoff between model size, inference time, and reasoning quality, with top-performing models demonstrating promise for future use in realworld defense applications—pending further testing on more ambiguous imagery.
