FL2V Turbo 4-step v1.2 (768p) released — improved audio quality

#52
by lightx2v - opened
Owner

We’re excited to release the MiniMax-H3 FL2V Turbo 4-step v1.2 768p LoRA!

Compared with v1.1, this release focuses on improving the audio generation quality. It produces cleaner and more stable audio with fewer artifacts while retaining the same fast 4-step generation setup.

Model Weights

Recommended Inference Settings

  • Steps: 4
  • Video shift: 6
  • Audio shift: 3
  • Sampler: Euler
  • Resolution: Up to 768p

Comparison with v1.1

The following comparisons use the same input, prompt, seed, and inference settings. Please enable audio when playing the videos.

v1.1

v1.2

nice job!

Video shift: 6
Audio shift: 3

how do i change this on the default comfy template?

Video shift: 6
Audio shift: 3

how do i change this on the default comfy template?

Node is called ModelSamplingMiniMaxH3. Connect it after the Load Diffusion Model node.

Thanks for the Lora! She is perfect. This Combo works perfect for me @0.4mp resolution.
lx

When looking at the samples it looks like the quality took a big hit in v1.2.

Thanks for the Lora! She is perfect. This Combo works perfect for me @0.4mp resolution.
lx

What is the use of this (beta sampaling scheduler)?

When looking at the samples it looks like the quality took a big hit in v1.2.

agreed. His hair is way worse.

why do we push for 4 step? Can we keep it at 8 step please?

When looking at the samples it looks like the quality took a big hit in v1.2.

agreed. His hair is way worse.

why do we push for 4 step? Can we keep it at 8 step please?

I totally agree with you... every speedhacks degrade the models quality, mostly by alot, why exagerate with just 4 steps? 8 is enough, damn even a turbo lora for 10 steps would be better.

For people asking why 4 step vs 8 step.

Simple, upscaling. Use normal or 8-step for lowres, then use 4 step for upscale. It saves many hours in the long run. Upscaling doesn't require the model to generate coherence from scratch so using the 4 step lora for say 2 steps at 0.25 denoise works great at hires fix 768 when generating from a low-res at 256.

I use spectrum sampler at low-res, upscale to 768 with 4-step lora for 3 steps, then 1024 again with 4-step lora for 2 step. Allows me to get crisp 1080p footage in somewhat reasonable time.

For people asking why 4 step vs 8 step.

Simple, upscaling. Use normal or 8-step for lowres, then use 4 step for upscale. It saves many hours in the long run. Upscaling doesn't require the model to generate coherence from scratch so using the 4 step lora for say 2 steps at 0.25 denoise works great at hires fix 768 when generating from a low-res at 256.

I use spectrum sampler at low-res, upscale to 768 with 4-step lora for 3 steps, then 1024 again with 4-step lora for 2 step. Allows me to get crisp 1080p footage in somewhat reasonable time.

Can you send me ur workflow please?

For people asking why 4 step vs 8 step.

Simple, upscaling. Use normal or 8-step for lowres, then use 4 step for upscale. It saves many hours in the long run. Upscaling doesn't require the model to generate coherence from scratch so using the 4 step lora for say 2 steps at 0.25 denoise works great at hires fix 768 when generating from a low-res at 256.

I use spectrum sampler at low-res, upscale to 768 with 4-step lora for 3 steps, then 1024 again with 4-step lora for 2 step. Allows me to get crisp 1080p footage in somewhat reasonable time.

oooooooooh ! ok that makes more sense

Sign up or log in to comment