Instructions to use lightx2v/Minimax-h3-Turbo with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use lightx2v/Minimax-h3-Turbo with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline from diffusers.utils import load_image, export_to_video # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("lightx2v/Minimax-h3-Turbo", dtype=torch.bfloat16, device_map="cuda") pipe.to("cuda") prompt = "A man with short gray hair plays a red electric guitar." image = load_image( "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/guitar-man.png" ) output = pipe(image=image, prompt=prompt).frames[0] export_to_video(output, "output.mp4") - Notebooks
- Google Colab
- Kaggle
FL2V Turbo 4-step v1.2 (768p) released — improved audio quality
We’re excited to release the MiniMax-H3 FL2V Turbo 4-step v1.2 768p LoRA!
Compared with v1.1, this release focuses on improving the audio generation quality. It produces cleaner and more stable audio with fewer artifacts while retaining the same fast 4-step generation setup.
Model Weights
Recommended Inference Settings
- Steps: 4
- Video shift: 6
- Audio shift: 3
- Sampler: Euler
- Resolution: Up to 768p
Comparison with v1.1
The following comparisons use the same input, prompt, seed, and inference settings. Please enable audio when playing the videos.
v1.1
v1.2
nice job!
Video shift: 6
Audio shift: 3
how do i change this on the default comfy template?
Video shift: 6
Audio shift: 3how do i change this on the default comfy template?
Node is called ModelSamplingMiniMaxH3. Connect it after the Load Diffusion Model node.
When looking at the samples it looks like the quality took a big hit in v1.2.
When looking at the samples it looks like the quality took a big hit in v1.2.
agreed. His hair is way worse.
why do we push for 4 step? Can we keep it at 8 step please?
When looking at the samples it looks like the quality took a big hit in v1.2.
agreed. His hair is way worse.
why do we push for 4 step? Can we keep it at 8 step please?
I totally agree with you... every speedhacks degrade the models quality, mostly by alot, why exagerate with just 4 steps? 8 is enough, damn even a turbo lora for 10 steps would be better.
For people asking why 4 step vs 8 step.
Simple, upscaling. Use normal or 8-step for lowres, then use 4 step for upscale. It saves many hours in the long run. Upscaling doesn't require the model to generate coherence from scratch so using the 4 step lora for say 2 steps at 0.25 denoise works great at hires fix 768 when generating from a low-res at 256.
I use spectrum sampler at low-res, upscale to 768 with 4-step lora for 3 steps, then 1024 again with 4-step lora for 2 step. Allows me to get crisp 1080p footage in somewhat reasonable time.
For people asking why 4 step vs 8 step.
Simple, upscaling. Use normal or 8-step for lowres, then use 4 step for upscale. It saves many hours in the long run. Upscaling doesn't require the model to generate coherence from scratch so using the 4 step lora for say 2 steps at 0.25 denoise works great at hires fix 768 when generating from a low-res at 256.
I use spectrum sampler at low-res, upscale to 768 with 4-step lora for 3 steps, then 1024 again with 4-step lora for 2 step. Allows me to get crisp 1080p footage in somewhat reasonable time.
Can you send me ur workflow please?
For people asking why 4 step vs 8 step.
Simple, upscaling. Use normal or 8-step for lowres, then use 4 step for upscale. It saves many hours in the long run. Upscaling doesn't require the model to generate coherence from scratch so using the 4 step lora for say 2 steps at 0.25 denoise works great at hires fix 768 when generating from a low-res at 256.
I use spectrum sampler at low-res, upscale to 768 with 4-step lora for 3 steps, then 1024 again with 4-step lora for 2 step. Allows me to get crisp 1080p footage in somewhat reasonable time.
oooooooooh ! ok that makes more sense
