Compositional Zero-Shot Learning (CZSL) aims to recognize novel concepts formed by known states and objects during training. Existing methods either learn the combined state-object representation, challenging the generalization of unseen compositions, or design two classifiers to identify state and object separately from image features, ignoring the intrinsic relationship between them. To jointly eliminate the above issues and construct a more robust CZSL system, we propose a novel framework termed Decomposed Fusion with Soft Prompt (DFSP)1, by involving vision-language models (VLMs) for unseen composition recognition. Specifically, DFSP constructs a vector combination of learnable soft prompts with state and object to establish the joint representation of them. In addition, a cross-modal decomposed fusion module is designed between the language and image branches, which decomposes state and object among language features instead of image features. Notably, being fused with the decomposed features, the image features can be more expressive for learning the relationship with states and objects, respectively, to improve the response of unseen compositions in the pair space, hence narrowing the domain gap between seen and unseen sets. Experimental results on three challenging benchmarks demonstrate that our approach significantly outperforms other state-of-the-art methods by large margins.

提出了一种名为 DFSP 的新型框架，它结合了视觉-语言模型(VLM)用于无人先前经验认知的建立，通过可学习的软提示与状态和对象的矢量组合来建立它们之间的共同表示，并在语言和图像分支之间设计了一种跨模式分解融合模块，从而更好地学习它们之间的关系，提高了成对空间中未知构成的反应，从而缩小了已知集和未知集之间的域间隙。实验结果表明，该方法在三个具有挑战性的基准测试数据集上对于已有的最先进方法有显着的改善。

分解软提示引导融合增强组合式零样本学习