Large Vision-Language Models (LVLMs) are gaining traction for their remarkable ability to process and integrate visual and textual data. Despite their popularity, the capacity of LVLMs to generate precise, fine-grained textual descriptions has not been fully explored. This study addresses this gap by focusing on \textit{distinctiveness} and \textit{fidelity}, assessing how models like Open-Flamingo, IDEFICS, and MiniGPT-4 can distinguish between similar objects and accurately describe visual features. We proposed the Textual Retrieval-Augmented Classification (TRAC) framework, which, by leveraging its generative capabilities, allows us to delve deeper into analyzing fine-grained visual description generation. This research provides valuable insights into the generation quality of LVLMs, enhancing the understanding of multimodal language models. Notably, MiniGPT-4 stands out for its better ability to generate fine-grained descriptions, outperforming the other two models in this aspect. The code is provided at \url{https://anonymous.4open.science/r/Explore_FGVDs-E277}.

该研究使用大规模视觉语言模型(LVLMs)来评估它们在识别相似对象和准确描述视觉特征方面的独特性和忠实度，并提出了文本检索增强分类(TRAC)框架以深入分析细粒度的视觉描述生成。研究结果表明，在生成细粒度描述方面，MiniGPT-4比其他两个模型表现更好。

大型视觉语言模型生成的描述的独特性和准确性探究