FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree-NVFP4

This UniServe deployment checkpoint applies calibrated NVIDIA Model Optimizer NVFP4 W4A4 to the fused gate/up and down projections in all 50 H3 denoiser gated MLPs and to the qkv, attention-output, fused gate/up, and down projections in all 36 Video VAE decoder Transformer blocks; denoiser attention, the BF16 text encoder, token refiner, modulation, normalization, latent and conditioning boundaries, convolutions, residual/affine paths, and audio decoder remain BF16/FP32. Calibration used 1,000 prompts with complete four-forward denoising trajectories and deployment-distribution VAE decoding, ModelOpt commit 6a4b3f147e14a6fec690fedbced8df402344085d, E2M1 values, K16 FP8 E4M3 block scales, FP32 tensor scales, and calibration-derived activation tensor scales.

This repository is a UniServe deployment checkpoint in NVIDIA Model Optimizer's unified Hugging Face layout, derived from FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree. It generates synchronized video and audio with four transformer forwards and retains the source checkpoint's VSA-H3 execution contract.

Run with UniServe

The quantized transformer and vae components store packed E2M1 values in .weight, K16 FP8 E4M3 block scales in .weight_scale, the FP32 weight scale in .weight_scale_2 and the static activation scale in .input_scale, and each declares a quantization_config (quant_method: modelopt, quant_algo: NVFP4) in its config.json; layers listed under ignore stay dense. UniServe detects this configuration automatically. The checkpoint does not require NVIDIA Model Optimizer at deployment time and rejects runtime precision presets or component overrides that would conflict with its calibrated weights and scales. The tensors are the calibrated values and scales of this repository's earlier packed layout, renamed and described bit for bit, so the validation below applies unchanged.

uniserve serve skx618/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree-NVFP4 \
  --served-model-name MiniMax-H3 \
  --dtype bfloat16

The checkpoint was validated on four NVIDIA B200 GPUs. The GPU count must divide H3's 56 attention heads.

Validation status

Packed values and scales passed 39/39 projection conformance checks against ModelOpt fake quantization. Candidate A reduced measured end-to-end latency by 16.3โ€“28.4% and peak aggregate GPU memory by 12.8โ€“17.7% versus the BF16/FP16 reference over the fixed benchmark points. Same-seed metrics measure trajectory drift rather than absolute perceptual quality; the frozen blinded human review has not been completed, so this checkpoint is not claimed to be perceptually equivalent to the reference.

A separate executable-checkpoint review generated five 5-second, 1344ร—768, 24-fps synchronized audio/video samples from this checkpoint and five same-prompt, same-seed quality-reference samples successfully. Across those five pairs, mean RGB PSNR was 13.81 dB, SSIM 0.4645, AlexNet LPIPS 0.4729, and audio log-mel cosine 0.9260. These prompts came from the calibration-source dataset, so the measurements describe trajectory drift and generation viability, not independent held-out quality acceptance.

License

This derivative inherits the MiniMax H3 Community License and the source checkpoint's usage conditions.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree-NVFP4