This paper presents a novel method, termed Bridge to Answer, to infer correct answers for questions about a given video by leveraging adequate graph interactions of heterogeneous crossmodal graphs. To realize this, we learn question conditioned visual graphs by exploiting the relation between video and question to enable each visual node using question-to-visual interactions to encompass both visual and linguistic cues. In addition, we propose bridged visual-to-visual interactions to incorporate two complementary visual information on appearance and motion by placing the question graph as an intermediate bridge. This bridged architecture allows reliable message passing through compositional semantics of the question to generate an appropriate answer. As a result, our method can learn the question conditioned visual representations attributed to appearance and motion that show powerful capability for video question answering. Extensive experiments prove that the proposed method provides effective and superior performance than state-of-the-art methods on several benchmarks.

本文提出了一种名为 Bridge to Answer 的新方法，通过利用异构交叉模式图的充分图交互来推断有关给定视频的问题的正确答案，通过学习问题调节的视觉图，对视觉节点使用问题 - 视觉交互来包含视觉和语言线索，并通过将问题图作为中间桥梁来将两个互补的视觉信息放在一起,使可靠的信息传递，以生成适当的答案，从而证明了该方法在视频问答方面提供了有效的上乘表现。

桥接到答案: 面向视频问答的结构感知图交互网络