The news of July 21stByteDance The team of Seed released yesterday, Seed Audio 1.0, which uses a unified framework to model human, audio and environmental voices and to organize roles, lines, emotions, background and voice when they enter。

time control: the model will break the creative intent into a structured time line, currently supporting 100 ms space accuracy。
Sound and duration: Supports text or reference audio input, which can generate about 2 minutes of audio at a single time and can continue to be extended to maintain the role ' s voice and expression。
Multilingual: covering more than 20 languages, including Chinese, English, Japanese, Korean and Spanish, the same sound adjusts rhythm, accent and emotions to the target language。
The official assessment of the scenes, such as video, animation, podcast conversations, live tape, drama and speech synthesis, stated that the audio availability of most scenes exceeded 90%. Seed Audio 1.0 is on line at the Volcanic Ark Experience Centre; follow-up programmes support input and generation methods such as video reference, long audio and sub-orbit。