Publications

SPECI: Fine-grained Evaluation of Specificity for Image Captioning

AACL-IJCNLP

Publication date: November 6, 2026

Son Luu, Hiep Nguyen, Trung Bui, Viet Lai, Le-Minh Nguyen

Image Captioning lies at the center of synthetic data generation in computer vision. The quality of the caption defines the behavior of the resulting model trained on image-caption pairs generated by Image Captioning. Measuring the specificity of image captioning plays an important role in supplying high-quality captions for the downstream tasks. To address the challenge of automatically assessing the quality of the caption, we propose Speci, an automatic metric that decomposes dense captions into atomic facts and evaluates specificity across visual criteria such as granularity, spatial detail, color, and linguistic focus using WordNet-based hyponym levels. Empirical results on the benchmark datasets demonstrate significant correlation with human judgments, and its application in image captioning and image reconstruction tasks highlights the practical utility of Speci.


Research Area:  Adobe Research iconNatural Language Processing