Training audio-to-image generative models requires an abundance of diverse audio-visual pairs that are semantically aligned. Such data is almost always curated from in-the-wild videos, given the cross-modal semantic correspondence that is inherent to them. In this work, we hypothesize that insisting on the absolute need for ground truth audio-visual correspondence, is not only unnecessary, but also leads to severe restrictions in scale, quality, and diversity of the data, ultimately impairing its use in the modern generative models. That is, we propose a scalable image sonification framework where instances from a variety of high-quality yet disjoint uni-modal origins can be artificially paired through a retrieval process that is empowered by reasoning capabilities of modern vision-language models. To demonstrate the efficacy of this approach, we use our sonified images to train an audio-to-image generative model that performs competitively against state-of-the-art. Finally, through a series of ablation studies, we exhibit several intriguing auditory capabilities like semantic mixing and interpolation, loudness calibration and acoustic space modeling through reverberation that our model has implicitly developed to guide the image generation process.

本研究解决了音频到图像生成模型训练所需的音视频配对数据稀缺问题。我们提出了一种可扩展的图像声化框架，通过现代视觉语言模型的推理能力，将不同模态的数据进行人工配对。研究结果显示，该方法训练的模型在性能上与最先进的技术相当，并展示了多种有趣的听觉能力，如语义混合和声场建模等。

通过视觉组装声音进行音频到图像生成