ended4월 29일· 1 sources
VibeVoice: Scaling Speech Synthesis to Full-Length Podcasts
90분 대화도 한 번에, Microsoft VibeVoice가 바꿀 Speech AI의 미래
Why it matters
VibeVoice solves the fundamental architectural bottleneck of traditional speech AI by enabling the processing of 90-minute audio files in a single pass through a revolutionary 7.5 Hz tokenizer. This shift from short-segment processing to long-context understanding marks a significant leap for high-quality podcast generation and complex multi-speaker transcription.
1
Sources
+0
24h
—
Growth
138d
Active
VibeVoiceMicrosoft ResearchSpeech AIAudio TokenizationICLR 2026Long-form Audio