Independent builder working on local AI, reasoning models, and practical agents. I test models, experiment with small weight edits, and share results, GGUFs, and evaluation tools.
Abliterated thinking models have a habit: they decide what to say, then keep arguing with themselves until the token budget runs out. Qwen3.8-27B Abliterated ThinkFix is an uncensored Qwen3.8-27B with one output row edited so the model closes its thinking when it's ready to answer.
On 40 sensitive prompts the edit never saw, 8k output cap: 33 replies finished cleanly instead of 22, 5 hit the limit instead of 17, 61 minutes for the set instead of 75.
Agent use was checked four ways and the edit changes nothing there. MTP draft head kept. Q4_K_M to Q8_0, each tested after the edit. The card lists what the edit was fit on and what it does not do.
I released a one-row weight edit for Qwen3.8 27B and Flash-Next to address a frustrating failure: using the entire output budget thinking, then returning no final answer.
The edit changes only the output-layer row that scores </think> and is packaged into ordinary GGUF files.
On 200 MATH-500 problems at a 4,096-token output limit:
• 27B: 31 empty answers → 0 • Flash-Next: 29 → 0
For comparison, llama.cpp's --reasoning-budget 2048 also eliminated blanks, with similar accuracy. The aim here is to put the control in the model file.
The report includes correctness scores, regressions, compatibility checks, evaluation limitations, and links to both models.