MiniMax-H3, FP8 and INT8

Pre-quantized PyTorch checkpoints of MiniMaxAI/MiniMax-H3, for diffusers plus torchao.

MiniMax H3 is an omni-modal generative system that produces video with native stereo audio, up to 15 seconds at 24 FPS with 32 kHz stereo audio. The files here quantize the fl2va_pruned H3-Base first-and-last-frame variant, which takes zero, one or two input images plus text, and the ref2va_pruned omni-reference variant, which takes a prompt plus up to twelve image, video and audio references. Both schemes quantize the same 200 main-block matmuls, 95.8% of the parameters, and leave every 1-D gain, every bias and the whole modulation path bit-identical.

Examples

Same prompt, same seed, same settings, one clip per checkpoint. 960x544, 124 frames, 24 FPS, 8 steps, guidance 1.0, seed 11, on a single card.

a red panda stepping along a mossy log in a misty forest, cinematic

INT8 FP8

The GIFs are downsampled and silent. For the full 960x544 clips with their native 32 kHz stereo audio track, play assets/h3_int8.mp4 and assets/h3_fp8.mp4. H3 generates the audio jointly with the video, so the audio is part of the model output rather than something added afterwards.

Stills from other prompts, three per format:

INT8

FP8

Files

H3 ships two denoiser partitions and a load opens exactly one. transformer/ serves the keyframe workflow fl2va, which also covers text-only generation; transformer_ref/ serves the omni-reference workflow ref2va. They are separate weights, so each partition has its own checkpoints here and the two sets are not interchangeable.

Keyframe and text-only (fl2va, from transformer/):

File Size
MiniMax-H3-INT8.pt 18.86 GiB
MiniMax-H3-INT8-ConvRot.pt 18.86 GiB
MiniMax-H3-FP8.pt 18.87 GiB

Omni-reference (ref2va, from transformer_ref/):

File Size
MiniMax-H3-Ref2VA-INT8-ConvRot.pt 18.86 GiB
MiniMax-H3-Ref2VA-FP8.pt 18.87 GiB

Every checkpoint lives here, INT8 included. One repo on purpose: these are one artifact built a few ways, and a loader picks between them by filename. Against the 40.23 GB bf16 pruned source any of them takes the checkpoint to 20.25 GB and the end-to-end render peak from 57.11 GB to 36.97 GB. Prefer INT8 with torch.compile and FP8 without it.

Pick the pair that matches the workflow you are running. The two partitions have the same class, the same config and the same 635-key state dict, so a keyframe checkpoint seeded into the reference workflow loads cleanly, passes every metadata check and generates from the wrong weights rather than failing. The Ref2VA in the filename is the only thing that distinguishes them.

The -ConvRot INT8 files store their weights in a Hadamard-rotated basis and need a loader that rotates the activations to match; they carry a format tag that makes an older loader refuse them rather than read them as plain INT8. Where both an INT8 and an INT8-ConvRot file exist, the ConvRot one is the newer build.

Reference (Ref2VA) build recipe

Identical to the keyframe rungs apart from the source file: 313 Linears quantized, 55 skipped, the adaLN modulation path left at full precision, and on the INT8 arm a ConvRot Hadamard rotation at group size 256 over all 313. The source is minimax_h3_ref2va_pruned_bf16.safetensors from Comfy-Org/MiniMax-H3, which is the reference partition with the same curve-form modulation pruning as the keyframe one, and both files record it in metadata.base_checkpoint.

Verified before publishing: each file loads through the pre-quantized path against the transformer_ref config with no missing or unexpected keys, and renders a 640x384, 124-frame reference clip. Measured against the reference-partition bfloat16 denoiser at the same prompt, seed and shape (20 steps, two prompts, one image reference), the same composition comes back with fine detail redistributed: SSIM 0.92 and 0.82 for INT8 and 0.88 and 0.78 for FP8, against a determinism ceiling of 1.00 for the bfloat16 arm rendered twice. Those numbers are a divergence measure, not a quality score.

Both are torchao pre-quantized transformer state dicts, per-output-channel absmax scales, INT8 symmetric and FP8 e4m3. The skeleton is built on meta and the quantized subclass tensors are assigned rather than copied, so dense bf16 never touches the GPU. Each file carries base_model_id = MiniMaxAI/MiniMax-H3 and a base_checkpoint naming the exact source it was cut from (.../minimax_h3_fl2va_pruned_bf16.safetensors or .../minimax_h3_ref2va_pruned_bf16.safetensors) in its metadata, which the loader checks before accepting it.

The keyframe source is minimax_h3_fl2va_pruned_bf16.safetensors from Comfy-Org/MiniMax-H3, which already has the modulation pruned to [96768, 8] per block plus a shared adaln_t_table [1025, 8]. Note that stable-diffusion.cpp cannot load these, it has no int8 linear path that reads external per-channel scales, so for sd.cpp use unsloth/MiniMax-H3-GGUF instead.

Licence

MiniMax H3 Community License Agreement, from MiniMax-H3. Full text in LICENSE. Read it before use: it defines an Applicable Territory and excludes some jurisdictions from it. MiniMax also publish a Q&A about the licence.

These files are Model Derivatives, not a plain copy: the transformer is quantized and its modulation is pruned, both of which change the numerics. Section III of the licence wants that stated, so NOTICE lists every change along with the attribution. Not an official MiniMax product, and not endorsed by MiniMax.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for unsloth/MiniMax-H3-FP8

Quantized
(40)
this model