Qwen3.5-9B-Heretic-v2-EQ-v5.1-GGUF

GGUF quantizations of nivvis/Qwen3.5-9B-EQ-v5.1 for llama.cpp, Ollama, and LM Studio.

Example outputs: EQ 9B v5.1 vs Vanilla Qwen3.5-9B — same prompt, same settings.

Available quantizations

Quant Size Notes
F16 ~17.9 GB Full precision, lossless conversion
Q4_K_M ~5.6 GB Best 4-bit balance, recommended for most users

Qwen3.5-9B-EQ-v5.1

A DPO fine-tune of trohrbaugh/Qwen3.5-9B-heretic-v2 for emotional intelligence and empathetic response quality.

This is still intended as a general use model (agentic, coding, general chat). Tuning was light and precise — no capability regression.

What this model does

  • Validates without sycophancy — empathizes with frustration without rubber-stamping bad behavior
  • Sets boundaries warmly — names uncomfortable truths without lecturing
  • Sounds human — conversational tone, not therapist-speak. Better tone vs vanilla Qwen 3.5, e.g. "It sounds like"

Benchmarks

All benchmarks: thinking=on, t=1.0, top_p=0.95 unless noted.

EQ-Bench 3

Rubric score = avg of 6 scored criteria × 5 (0-100), Opus 4.6 judged. Leaderboard scores from eqbench.com. Only qualitative criteria count (empathy, pragmatic EI, insight, social dexterity, emotional reasoning, message tailoring).

# Model Score
1 gpt-5.4 80.9
2 claude-sonnet-4-6 79.9
3 claude-opus-4-6 78.6
4 gemma-4-31B-it 72.8
5 Qwen3.5-397B 70.1
6 Qwen3.5-35b-EQ-v5.0 68.2
7 Qwen3.5-35b-EQ-v5.1 66.8
8 Qwen3.5-9b-EQ-v5.1 65.4
9 Qwen3-235B 61.4
10 gpt-4.5-preview 59.7
11 Qwen3.5-9b (vanilla) 59.1
12 o4-mini 58.1
13 DeepSeek-V3-0324 57.9
14 Qwen3.5-35b (vanilla) 55.4
15 gpt-oss-120b 51.4

HumanEval+

Model HumanEval base HumanEval+
EQ 9B v5.1 92.7% 86.0%
Vanilla Qwen3.5-9B 93.9% 87.2%

GSM8K

Metric EQ 9B v5.1 Vanilla Qwen3.5-9B
Accuracy 84.9% 82.8%

IFBench

Metric EQ 9B v5.1 Vanilla Qwen3.5-9B
Loose (leaderboard) 67.3% 65.0%
Strict 56.8% 56.8%

How to use

llama-server (OpenAI-compatible API)

llama-server \
  -m Qwen3.5-9B-Heretic-v2-EQ-v5.1-Q4_K_M.gguf \
  --host 0.0.0.0 --port 30000 \
  -ngl 99 --jinja

Ollama

ollama run hf.co/nivvis/Qwen3.5-9B-Heretic-v2-EQ-v5.1-GGUF:Q4_K_M

Thinking mode

This model supports thinking mode. To disable (for faster, direct responses):

{"chat_template_kwargs": {"enable_thinking": false}}

Sampling recommendations

Use the same settings as Qwen3.5-9B:

Mode temp top_p top_k presence_penalty
Thinking (general) 1.0 0.95 20 1.5
Thinking (coding) 0.6 0.95 20 0.0
Non-thinking (general) 0.7 0.8 20 1.5
Non-thinking (reasoning) 1.0 0.95 20 1.5

Other formats

Lineage

Qwen/Qwen3.5-9B
  → trohrbaugh/Qwen3.5-9B-heretic-v2 (decensored)
    → nivvis/Qwen3.5-9B-EQ-v5.1 (DPO for EQ)
      → this repo (GGUF quantizations)

License

Apache 2.0, following the base Qwen3.5 license.

Downloads last month
179
GGUF
Model size
9B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including nivvis/Qwen3.5-9B-EQ-v5.1-GGUF