The authors conducted systematic experiments to validate the model's comprehensive performance across multimodal tasks.
The evaluation covers five major directions: in image captioning, Qwen-VL achieves an 85.8 CIDEr score on Flickr30K, significantly surpassing large-scale models such as Flamingo-80B; in general visual question answering (VQA), it attains 79.5%, 58.6%, and 59.3% accuracy on VQAv2, OKVQA, and GQA respectively, demonstrating strong reasoning capabilities; in text-oriented VQA tasks (e.g., TextVQA, DocVQA), the model excels at text recognition and chart understanding (e.g., 63.8% accuracy on TextVQA); in referring expression comprehension tasks (e.g., the RefCOCO series), grounding accuracy exceeds 89%, proving its fine-grained visual localization ability; furthermore, in few-shot learning scenarios, the model's performance approaches that of models with 10x more parameters, highlighting its efficient transferability.
Ultimately, the instruction-tuned Qwen-VL-Chat leads comprehensively in real-world dialogue evaluations (e.g., TouchStone, MME), with particularly pronounced advantages in Chinese multi-turn dialogue and multi-image understanding tasks, confirming its strong generalization capability in practical applications.