Upload a portrait image (9:16 vertical, full-body/half-body facing front) and a music file (wav/mp3, 10s-30s).
Select a dance genre and customize generation parameters if desired, then click Generate.
ZeroGPU Optimization: To stay within Hugging Face ZeroGPU time limits (120s max), keep Inference Steps around 12-15 and Max Audio Duration around 10-15s.
Input
1030
0120
Dance Genre
1050
115
Result
Examples
Portrait Image (will be auto-cropped to 9:16 vertical)