VibeVoice: Microsoft's Open-Source Frontier Voice AI
Microsoft's VibeVoice: 90-minute single-pass voice synthesis, frontier open-source models that beat ElevenLabs on MOS benchmarks
TOPIC_INDEX
7 published entries in this topic.
Microsoft's VibeVoice: 90-minute single-pass voice synthesis, frontier open-source models that beat ElevenLabs on MOS benchmarks
Poolside's Laguna XS.2 and M.1 MoE models feature the Muon optimizer—a novel training method that cuts steps by 15% vs AdamW. XS.2 is Apache 2.0 licensed, local-ready.
A landmark paper from 40+ researchers introduces the first comprehensive taxonomy for AI world models, mapping the evolution from passive predictors to autonomous simulators.
Tencent releases first open-source SOTA 3D world model, generating actual geometry from text/images that imports directly into Unity, Unreal, and Blender.
NVIDIA's Lyra 2.0 generates persistent 3D worlds from single images by solving spatial forgetting and temporal drifting—the two fundamental failure modes of long-horizon video generation.
LPSR is a training-free method that improves 8B model MATH-500 scores from 28.8% to 44.0%, beating a 70B model with 8.75x fewer parameters.
MiniMax M2.7 is the first AI model to autonomously optimize its own behavioral scaffolding, achieving 30% improvement over 100+ iterations without weight changes.