FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking Portrait Paper β’ 2412.01064 β’ Published Dec 2, 2024 β’ 25
Running 139 π Whisper Large V3 Turbo WebGPU ML-powered speech recognition directly in your browser
MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators Paper β’ 2404.05014 β’ Published Apr 7, 2024 β’ 33
MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? Paper β’ 2403.14624 β’ Published Mar 21, 2024 β’ 51
Mora: Enabling Generalist Video Generation via A Multi-Agent Framework Paper β’ 2403.13248 β’ Published Mar 20, 2024 β’ 78
VastGaussian: Vast 3D Gaussians for Large Scene Reconstruction Paper β’ 2402.17427 β’ Published Feb 27, 2024 β’ 9