Instructions to use sakamakismile/Ling-3.0-flash-W4A4-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use sakamakismile/Ling-3.0-flash-W4A4-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="sakamakismile/Ling-3.0-flash-W4A4-NVFP4", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("sakamakismile/Ling-3.0-flash-W4A4-NVFP4", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use sakamakismile/Ling-3.0-flash-W4A4-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "sakamakismile/Ling-3.0-flash-W4A4-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sakamakismile/Ling-3.0-flash-W4A4-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/sakamakismile/Ling-3.0-flash-W4A4-NVFP4
- SGLang
How to use sakamakismile/Ling-3.0-flash-W4A4-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "sakamakismile/Ling-3.0-flash-W4A4-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sakamakismile/Ling-3.0-flash-W4A4-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "sakamakismile/Ling-3.0-flash-W4A4-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sakamakismile/Ling-3.0-flash-W4A4-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use sakamakismile/Ling-3.0-flash-W4A4-NVFP4 with Docker Model Runner:
docker model run hf.co/sakamakismile/Ling-3.0-flash-W4A4-NVFP4
Ling-3.0-flash — NVFP4 W4A4 (calibrated, MTP verified working)
The first W4A4 NVFP4 quantization of Ling-3.0-flash (124B total / 5.1B active, hybrid KDA+MLA MoE) — weights and activations in NVFP4, calibrated with 128 samples via llm-compressor.
| Size | 81.4 GB (from 255 GB BF16) |
| Format | compressed-tensors nvfp4-pack-quantized, W4A4 (E2M1 + FP8-E4M3 group-16 scales, static input global scales) |
| Measured | 110 tok/s single / 493 tok/s @ 8-way generation, ~7,000 tok/s prefill — 8×RTX PRO 2000 Blackwell (16 GB), TP=8, 32K context |
| Long context | 131,072 tokens verified (eager mode; prefill 5,900–8,000 tok/s) |
| MTP | Works: 93% acceptance (k=1). Off by default — see below |
| Sanity | 6/6 (math / logic / code / JP idiom / instruction-following / long-form JP); tool-calling 4/4 |
Recommended: fp8 KV cache — 131K context at full speed
With the bundled vllm_patch/triton_decode_attention.py
(1-line shared-memory fix for sm_120), --kv-cache-dtype fp8 works and everything fits at once:
| context | KV budget | generation | prefill | |
|---|---|---|---|---|
| fp8 KV + patch (recommended) | 131,072 | 458,752 tok | 112.6 / 482–493 tok/s (1 / 8-way) | 1.9k single / 6.5k @ 4-way |
| bf16 KV, 32K | 32,768 | ~187K tok | 110 / 493 tok/s | up to 7.6k single (4096 chunks) |
bf16 KV, 131K (--enforce-eager) |
131,072 | ~187K tok | 16.5 / 31.7 tok/s | 5.9k–8.0k |
458K tokens of KV = three concurrent 131K streams, or eight at 57K each.
Quality battery under fp8 KV: 6/6 (math / logic / code / JP idiom / instruction-following /
long-form JP) — no measured degradation vs bf16 KV.
Single-stream prefill in the recommended config is chunking-bound (--max-num-batched-tokens 2048, the largest stable value at 131K on 16 GB); parallel prefill reaches ~6,500 tok/s.
Why context is cheap here but headroom is not. Only 7 of 42 layers are MLA (the other 35 are KDA linear attention and carry no KV), so the KV cost is 8.2 KB/token — roughly 9× cheaper per token than an all-MLA model like DeepSeek-V3. The constraint is the card: weights + non-torch overhead take 12.66 GB of a 16 GB GPU, leaving ~1.5 GB for KV and KDA state — a total budget of about 187K tokens to split between length and concurrency (131K × 1 stream, or 23K × 8 streams, etc.). On 80 GB cards this limit essentially vanishes.
How this differs from olka-fi/Ling-3.0-flash-NVFP4 (weight-only W4A16): this build quantizes activations too (W4A4) with real calibration data, and ships a working MTP path.
Quantization boundary
Same proven boundary as the W4A16 build: routed experts only (layers 2–41,
mlp.experts.*.{gate,up,down}_proj, 120.8B of 127.4B params) → NVFP4 W4A4.
Kept in BF16: shared experts, KDA & gated-MLA attention, router gate + expert_bias,
dense layers 0–1, the MTP layer 42, embeddings, lm_head, norms.
Serving (vLLM fork required)
Upstream vLLM does not support bailing_hybrid v3. Use the prebuilt image
olkafi/vllm-bailing-v3 (inclusionAI/vllm @ ling_3_0 + SwiGLU-clamp patch).
This exact command reproduces every number in this card (131K context, fp8 KV, 458,752-token KV budget, ~112 tok/s single-stream). Both bundled patch files are mounted over the image's own copies — no rebuild needed, they are pure Python:
REPO=/path/to/this-repo # e.g. $(huggingface-cli download sakamakismile/Ling-3.0-flash-W4A4-NVFP4)
docker run --gpus all -d --name ling3 --ipc=host -p 8000:8000 \
-e NCCL_P2P_DISABLE=1 -e NCCL_CUMEM_ENABLE=0 \
-v $REPO:/models/ling3:ro \
-v $REPO/vllm_patch/unquantized.py:/opt/vllm-ling3/vllm/model_executor/layers/fused_moe/oracle/unquantized.py:ro \
-v $REPO/vllm_patch/triton_decode_attention.py:/opt/vllm-ling3/vllm/v1/attention/ops/triton_decode_attention.py:ro \
olkafi/vllm-bailing-v3:latest /models/ling3 \
--served-model-name ling3 --trust-remote-code --host 0.0.0.0 --port 8000 \
--tensor-parallel-size 8 --disable-custom-all-reduce \
--kernel-config '{"moe_backend":"marlin"}' --kv-cache-dtype fp8 \
--compilation-config '{"max_cudagraph_capture_size":16}' \
--gpu-memory-utilization 0.95 --max-model-len 131072 --max-num-seqs 8 \
--max-num-batched-tokens 2048 \
--enable-prefix-caching --mamba-cache-mode align \
--enable-auto-tool-choice --tool-call-parser ling3 --reasoning-parser ling3
Notes for reproducers:
--max-num-batched-tokens 2048: 4096 OOMs at 131K on 16 GB cards (chunked-prefill activations); 2048 is stable.- On cards with more than 16 GB, drop the two patch mounts if you don't need them
(
flashinfer_cutlassmay still be broken on sm_120 — keepmoe_backend=marlinthere), and raise--max-num-batched-tokens/--max-num-seqsfreely. - NCCL flags are for machines without GPU P2P; harmless otherwise.
Sampling (from the base card): temperature=0.6, top_p=0.95, top_k=20, thinking enabled.
⚠️ Known landmines (sm_120 / consumer Blackwell)
moe_backend=marlinis mandatory. The auto-selectedflashinfer_cutlassbackend silently corrupts output (endless!!!!) on sm_120 with EP, and plain TP=8 hitsNotImplementedError(per-rank intermediate 768/8=96 needs unsupported padding).cutlassrejects EP. Marlin works and is the source of the numbers above.- KV headroom is tight — on 16 GB cards use
--gpu-memory-utilization 0.95,--max-num-seqs 8and--compilation-config '{"max_cudagraph_capture_size":16}'. 32K fits comfortably that way; 131K needs--enforce-eager(see table above). --kv-cache-dtype fp8needs the bundled kernel patch on sm_120. The TRITON_MLA fp8 decode kernel asks for 102,400 bytes of shared memory atnum_stages=2; consumer Blackwell caps at 101,376 — 1 KB short.vllm_patch/triton_decode_attention.pydrops fp8-KV MLA tonum_stages=1, which fits with no measured speed loss (112.6 tok/s single-stream, same as bf16 KV).- The bundled
vllm_patch/unquantized.pymakes the unquantized MTP-layer MoE fall back to triton whenmoe_backend=marlinis forced globally — without it, serving with--speculative-configfails at startup.
MTP / speculative decoding: verified working, off by default
The W4A16 release reported 0% acceptance and shipped with MTP disabled. With this build + the bundled patch, MTP works:
| config | single-stream | acceptance |
|---|---|---|
| no MTP (recommended) | 113.2 tok/s | — |
| MTP k=1 | 88.3 tok/s | 93.0% |
| MTP k=2 | 110.0 tok/s | 68%/tok |
The draft (BF16 MTP layer on the triton path) currently costs more than the accepted
tokens buy back on this hardware, so MTP is a proof-of-life, not a speed win — leave
--speculative-config off for throughput. To try it:
--speculative-config '{"method":"mtp","num_speculative_tokens":1}' (add
--compilation-config '{"max_cudagraph_capture_size":32}' --max-num-seqs 32 for memory).
Bake provenance
llm-compressor 0.12 (lna-lab/llmc-bake:v0.24.0), 128 calibration samples
(neuralmagic/calibration, chat-templated) @ 2048 tokens, basic pipeline with
CPU-resident weights and single-GPU onload calibration (fla/KDA kernels are
triton/cuda-only, so pure-CPU forward is impossible). Recipe quirks that will bite
reproducers: transformers 5.x needs three shims in the custom modeling file
(is_torch_fx_available, ROPE_INIT_FUNCTIONS['default'], _tied_weights_keys);
rope_scaling must be forced back to None after config normalization;
llm-compressor drops model_type from the saved config (breaking vLLM's MLA
detection — restore "model_type": "bailing_hybrid"); and post-calibration weight
observation must be moved to CPU (set_onload_device) or 61,440 expert observers
accumulate on one GPU and OOM.
Acknowledgements
Built on inclusionAI's Ling-3.0-flash (MIT). Quantization boundary and serving groundwork follow olka-fi's pioneering W4A16 release and vendor-fork patches.
MIT license, inherited from the base model.
- Downloads last month
- 157
Model tree for sakamakismile/Ling-3.0-flash-W4A4-NVFP4
Base model
inclusionAI/Ling-3.0-flash