We present a deep convolutional GAN which leverages techniques from MP3/Vorbis audio compression to produce long, high-quality audio samples with long-range coherence. The model uses a Modified Discrete Cosine Transform (MDCT) data representation, which includes all phase information. Phase generation is hence integral part of the model. We leverage the auditory masking and psychoacoustic perception limit of the human ear to widen the true distribution and stabilize the training process. The model architecture is a deep 2D convolutional network, where each subsequent generator model block increases the resolution along the time axis and adds a higher octave along the frequency axis. The deeper layers are connected with all parts of the output and have the context of the full track. This enables generation of samples which exhibit long-range coherence. We use MP3net to create 95s stereo tracks with a 22kHz sample rate after training for 250h on a single Cloud TPUv2. An additional benefit of the CNN-based model architecture is that generation of new songs is almost instantaneous.

本文提出了一种基于卷积神经网络的生成对抗网络，应用了音频压缩和MDCT数据表示等技术生成长时间和高质量的音频样本，并利用人耳的听觉掩蔽效应和心理声学感知限制来拓宽真实分布并稳定训练过程。经过250小时的训练，使用单个Cloud TPUv2可以创造出95秒的立体声音轨，且模型具有快速生成新歌曲的优势。

MP3net: 用简单的卷积 GAN 从原始音频中生成连贯分钟级音乐