Aztec-Coder-4B-GGUF

Aztec-Coder-4B is a 4-billion-parameter agentic coding model from San Diego State University's James Silberrad Brown Center for AI Research (JSBCAI), fine-tuned from Qwen3.5-4B to investigate bugs, edit files, run commands, and verify its own fixes in real software repositories — a capability class the authors note has typically required 27B+ models — while fitting on consumer hardware (~8GB VRAM in BF16, ~5GB as an NVFP4 quantized variant, which retains roughly half the generalization capability at 38.0%/12-32 on the same evaluation slice). It was trained in three stages: seed demonstrations from ~1,875 test-verified coding trajectories generated by GLM-5.3, reinforcement learning via 145 batches of on-policy GRPO over 237 curated software-engineering problems using test-pass/fail as the only reward signal, and repeated generalization checks against unseen bugs, alongside general instruction-following data drawn from NVIDIA's Nemotron-Post-Training-Dataset-v2. On 121 held-out, never-trained-on real bugs verified by running each project's hidden test suite, it jumped from 10.1% (base Qwen3.5-4B) to 82.9% solve rate, while also improving general capabilities rather than trading them off — IFEval rose from 84.66 to 87.21 and MMLU-Pro from 64.0% to 70.0% — alongside gains on Live-60 real-world engineering tasks (15.0% to 21.7%), with Terminal-Bench 1.0 held flat at 33.8% and Terminal-Bench 2.1 results still pending. It's served via vLLM with the Qwen3.5 chat template, qwen3 reasoning parser, and qwen3_coder tool-call format, is tuned specifically for sandboxed container agent loops rather than general deployment, inherits its safety behavior unmodified from the base model (the RL phase optimized purely for test-passing with no safety-specific training), and is released under Apache-2.0.

Model Files

File Name Quant Type File Size File Link Description
Aztec-Coder-4B.BF16.gguf BF16 8.42 GB Link Full BF16 weights. Highest quality, largest file size.
Aztec-Coder-4B.Q3_K_L.gguf Q3_K_L 2.42 GB Link Lower quality but usable, good for low RAM availability.
Aztec-Coder-4B.Q3_K_M.gguf Q3_K_M 2.26 GB Link Low quality.
Aztec-Coder-4B.Q4_K_M.gguf Q4_K_M 2.71 GB Link Good quality, default size for most use cases, recommended.
Aztec-Coder-4B.Q4_K_S.gguf Q4_K_S 2.56 GB Link Slightly lower quality with more space savings, recommended.
Aztec-Coder-4B.Q5_K_M.gguf Q5_K_M 3.07 GB Link High quality, recommended.
Aztec-Coder-4B.Q5_K_S.gguf Q5_K_S 2.99 GB Link High quality, recommended.
Aztec-Coder-4B.Q6_K.gguf Q6_K 3.46 GB Link Very high quality, near perfect, recommended.

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Downloads last month
215
GGUF
Model size
4B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

6-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/Aztec-Coder-4B-GGUF

Finetuned
Qwen/Qwen3.5-4B
Quantized
(1)
this model

Dataset used to train prithivMLmods/Aztec-Coder-4B-GGUF

Collection including prithivMLmods/Aztec-Coder-4B-GGUF