We extend the effective SKIP-GRAM model of Mikolov et al. (2013) by taking visual information into account. Like S KIP - GRAM , our multimodal models (MMSKIP-GRAM) build vector-based word representations by learning to predict linguistic contexts in text corpora. However, for a restricted set of words, the models are also exposed to visual representations of the objects they denote (extracted from natural images), and must predict linguistic and visual features jointly. The MMS KIP - GRAM models achieve excellent performance on a variety of semantic benchmarks. Moreover, since they propagate visual information to all words, we also use them to improve image labeling and retrieval in the challenging zero-shot setup, where the test concepts are not seen in training. Finally, the MMS KIP - GRAM models discover intriguing vision-related properties of abstract words, paving the way to realistic implementations of embodied theories of meaning.

本研究通过将视觉信息纳入 SKIP-GRAM 模型，创新性地提出了一种多模式的词向量表达方式，并取得了良好的语义基准表现。同时，该模型还能够将视觉信息传递到所有词中，用于改进零样本图像标注和检索，并探索了抽象词汇的有趣视觉属性，为意义的具体化实现奠定了基础。

结合语言和视觉的多模式跳字模型