How to use from the
Use from the
Diffusers library
pip install -U diffusers transformers accelerate
import torch
from diffusers import DiffusionPipeline
from diffusers.utils import export_to_video

# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("MiniMaxAI/MiniMax-H3", dtype=torch.bfloat16, device_map="cuda")
pipe.load_lora_weights("ZhengmingYu/DMAD")

prompt = "A man with short gray hair plays a red electric guitar."

output = pipe(prompt=prompt).frames[0]
export_to_video(output, "output.mp4")

DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation

4-step MiniMax-H3 students for joint audio-video generation

Project Page Paper Code Demo Video

Zhengming Yu1,2, Junkun Yuan2, Haotian Yang2, Gordon Guocheng Qian2, Yizhi Wang2, Angtian Wang2, Yiding Yang2, Bo Liu2, Xin Li1, Wenping Wang1, Chongyang Ma2
1Texas A&M University, 2ByteDance

Videos generated by the 4-step DMAD student of MiniMax-H3 (video only; every clip also has generated audio)

This repository holds the DMAD students of the paper: 4-step students of MiniMax-H3 (33B, text-to-audio-video) and of Wan2.1-T2V (1.3B and 14B), 4- and 1-step students of SDXL, and 1-step students of the EDM ImageNet-64 teacher. The inference and training code is in the code repository (train/h3, train/wan, train/image).

MiniMax-H3 (text-to-audio-video)

Rank-128 LoRAs on the H3 transformer that turn the 50-step teacher into a 4-step generator of 1344x768 video with native stereo audio.

File Checkpoint Size
minimax_h3/dmad_minimax_h3_4step_lora_critic.safetensors the checkpoint of the paper: EMA of the student at iteration 800 of the main run 1.4 GB
minimax_h3/dmad_minimax_h3_4step_full_critic.safetensors the student of a run whose critic backbone is fully trained (the paper's run keeps it frozen under a LoRA): iteration 1600, live weights; it scores higher on AVGen-Bench 1.4 GB
minimax_h3/dmad_minimax_h3_4step_lora_critic_comfyui.safetensors lora_critic in ComfyUI's MiniMax-H3 key layout (exact conversion) 2.0 GB
minimax_h3/dmad_minimax_h3_4step_full_critic_comfyui.safetensors full_critic in ComfyUI's MiniMax-H3 key layout (exact conversion) 2.0 GB

LoRA layout of the first two: Diffusers keys (<module>.lora.down.weight = A [128, in], <module>.lora.up.weight = B [out, 128]) over attn.to_q/to_k/to_v/to_out.0, ff.net.0.proj, ff.net.2 of all 50 transformer blocks and the 2 token-refiner blocks (312 modules). alpha = rank = 128. The safetensors metadata repeats this.

The inference code lives in the code repository: inference.py with the sampler the paper used (re-noise step rule) and a Diffusers-pipeline example; its README covers the environment. Sampling settings: 4 steps, time shift 12 (video) and 2 (audio), no classifier-free guidance, 124 frames at 24 fps. --low-vram runs it on a 16 GB GPU.

git clone https://github.com/Yzmblog/DMAD.git && cd DMAD   # code + environment setup (see its README)
hf download MiniMaxAI/MiniMax-H3 --local-dir models/MiniMax-H3 --exclude "FL2VA/*" --exclude "Ref2VA/*" --exclude "transformer_ref/*"
hf download ZhengmingYu/DMAD --include "minimax_h3/*_critic.safetensors" --local-dir ckpt
python inference.py --model-dir models/MiniMax-H3 --lora ckpt/minimax_h3/dmad_minimax_h3_4step_lora_critic.safetensors \
    --prompt-file prompts/dmad_sweater.txt --seed 42 --output-dir outputs/dmad_sweater

ComfyUI

minimax_h3/dmad_minimax_h3_4step_{lora_critic,full_critic}_comfyui.safetensors are the same two LoRAs converted exactly to ComfyUI's MiniMax-H3 key layout (rank 128; q/k/v fused into attn.qkv_proj adapters of rank 384, alpha = rank, scale 1.0 โ€” no rank reduction; 2.0 GB each). Load with LoraLoaderModelOnly at strength 1.0, cfg 1.0, ModelSamplingMiniMaxH3 with shift 12 / audio shift 2, and sample with ComfyUI's lcm sampler and simple scheduler (the re-noise multistep rule and sigma grid the students were trained with; ODE samplers such as euler are not their operating point), or with the equivalent DMAD Sampler + DMAD Sigmas nodes from comfyui/ComfyUI-DMAD.

wget -P /path/to/ComfyUI/models/loras https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/dmad_minimax_h3_4step_full_critic_comfyui.safetensors
wget -P /path/to/ComfyUI/models/loras https://huggingface.co/ZhengmingYu/DMAD/resolve/main/minimax_h3/dmad_minimax_h3_4step_lora_critic_comfyui.safetensors

Ready-to-run workflows (24 GB GPU)

Two complete text-to-audio-video workflows in minimax_h3/workflows that make 15 s of 1344x768 video with stereo audio on a 24 GB GPU. Load the .json, or drag the example .mp4 (it embeds the workflow) into ComfyUI; missing models are offered for download from links stored in the workflow. Both sample with stock nodes (lcm + simple, full_critic LoRA); the ComfyUI-DMAD nodes are not needed.

4 steps 8 steps
Workflow dmad_h3_4step_15s_podcast.json dmad_h3_8step_15s_wok.json
Example (workflow embedded) dmad_h3_4step_15s_podcast.mp4, seed 2 dmad_h3_8step_15s_wok.mp4, seed 4
Output 1344x768, 362 frames = 15 s at 24 fps, stereo audio 1344x768, 362 frames = 15 s at 24 fps, stereo audio
Peak GPU memory 24.6 GiB 24.6 GiB
Time (H200, 24 GiB cap) 448 s, of which sampling 393 s 845 s, of which sampling 785 s

The 4-step workflow in ComfyUI

folder file (from Comfy-Org/MiniMax-H3 unless noted)
models/diffusion_models/ minimax_h3_fl2va_pruned_int8_convrot.safetensors
models/text_encoders/ qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
models/vae/ minimax_h3_video_vae_fp16.safetensors, minimax_h3_audio_vae_fp32.safetensors
models/loras/ dmad_minimax_h3_4step_full_critic_comfyui.safetensors (this repository)

Video lengths are 5 + 17k frames (124 = 5 s, 243 = 10 s, 362 = 15 s). On a 24 GB GPU, 15 s needs the H3 Memory Optimization node of H3-Optimizations (install with ComfyUI-Manager), which both workflows include. Without it, 5 s (length 124) fits in 24 GB with stock nodes only, and 15 s needs more than 32 GB (it fits in 40 GB); the video keeps its composition and action but is not bit-identical to the one made with the memory node. Times are end to end (model loading, text encoding, sampling, decoding) on an H200 with the PyTorch allocator capped at 24 GiB; a consumer GPU is slower, and sampling time scales linearly with steps.

16 GB and 22 GB GPUs. Both workflows also run unchanged with the GPU memory capped at 14 GiB (room for a 16 GB card's desktop and CUDA context): ComfyUI's dynamic VRAM keeps less of the model resident and streams more of it from system memory, and the videos are bit-identical to the 24 GiB runs. About 46 GiB of system memory is in use, so 64 GB of RAM is recommended. Outside ComfyUI, inference.py --low-vram in the code repository makes 5 s and 15 s videos in under 14 GiB of GPU memory, bit-identical to its larger-GPU runs.

Wan2.1 (text-to-video)

Full generators (EMA, spectral norm folded in) in the .pth format of train/wan: 4 steps, 480p, 81 frames, no classifier-free guidance.

File Model Size
wan2.1/dmad_wan2pt1_1pt3B.pth Wan2.1-T2V-1.3B student, iteration 17k 2.8 GB
wan2.1/dmad_wan2pt1_14B.pth Wan2.1-T2V-14B student, iteration 19.5k 29 GB
hf download ZhengmingYu/DMAD --include "wan2.1/*" --local-dir ckpt
bash experiments/dmad/sample.sh 1.3B ckpt/wan2.1/dmad_wan2pt1_1pt3B.pth my_prompts.json outputs/samples   # in train/wan

SDXL (text-to-image)

UNet state dicts in fp16 (spectral norm folded in) that load into the standard SDXL pipeline; see train/image for the sampling code.

File Model Size
sdxl/dmad_sdxl_4step_unet_fp16.bin 4-step student (backward simulation), iteration 17k 5.1 GB
sdxl/dmad_sdxl_1step_unet_fp16.bin 1-step student (ODE init), iteration 49k 5.1 GB
sdxl/dmad_sdxl_1step_frozencritic_unet_fp16.bin 1-step student (ODE init, frozen critic backbone), iteration 17.5k 5.1 GB
import torch
from diffusers import DiffusionPipeline, LCMScheduler, UNet2DConditionModel
from huggingface_hub import hf_hub_download

base_model_id = "stabilityai/stable-diffusion-xl-base-1.0"
unet = UNet2DConditionModel.from_config(base_model_id, subfolder="unet").to("cuda", torch.float16)
unet.load_state_dict(torch.load(hf_hub_download("ZhengmingYu/DMAD", "sdxl/dmad_sdxl_4step_unet_fp16.bin"), map_location="cuda"))
pipe = DiffusionPipeline.from_pretrained(base_model_id, unet=unet, torch_dtype=torch.float16, variant="fp16").to("cuda")
pipe.scheduler = LCMScheduler.from_config(pipe.scheduler.config)
image = pipe(prompt="a photo of a cat", num_inference_steps=4, guidance_scale=0, timesteps=[999, 749, 499, 249]).images[0]
# 1-step models: num_inference_steps=1, timesteps=[399]

ImageNet-64 (class-conditional)

1-step EMA generators of the EDM ImageNet-64 teacher, one folder per critic setting of the paper, in the checkpoint_model_<iteration>/pytorch_model_ema.bin layout that train/image's evaluation reads directly.

Folder Setting Size
imagenet64/dmad_imagenet_gaproute/ teacher-UNet critic + gap routing, iteration 568k 1.2 GB
imagenet64/dmad_imagenet_pgcritic/ pretrained-feature critic, iteration 108k 1.2 GB
imagenet64/dmad_imagenet_frozencritic/ frozen teacher-UNet critic + gap routing, iteration 221k 1.2 GB
hf download ZhengmingYu/DMAD --include "imagenet64/dmad_imagenet_gaproute/*" --local-dir ckpt
python main/edm/test_folder_edm.py --folder ckpt/imagenet64/dmad_imagenet_gaproute --run_once ...   # in train/image

Citation

@misc{yu2026dmad,
  title         = {DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation},
  author        = {Zhengming Yu and Junkun Yuan and Haotian Yang and Gordon Guocheng Qian and Yizhi Wang and
                   Angtian Wang and Yiding Yang and Bo Liu and Xin Li and Wenping Wang and Chongyang Ma},
  year          = {2026},
  eprint        = {2610.02188},
  archivePrefix = {arXiv}
}

License

Each family of weights is a derivative of its base model and is distributed under that model's license:

Images and videos produced with these weights are AI-generated.

Downloads last month
-
Inference Providers NEW

This task can take several minutes

Model tree for ZhengmingYu/DMAD

Adapter
(118)
this model

Spaces using ZhengmingYu/DMAD 3

Paper for ZhengmingYu/DMAD