Blackfrost-AI/Qwen3.8-27B-ABLITERATED-GGUF Image-Text-to-Text • 27B • Updated 1 day ago • 91.6k • 136
sakamakismile/DeepSeek-V4-Flash-0731-Abliterated-NVFP4 Text Generation • 304B • Updated 16 days ago • 1.47k • 12
Hikari07jp/Ternary-Bonsai-27B-Abliterated-LowDeg-GGUF Text Generation • 27B • Updated 27 days ago • 17.1k • 25
Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms Paper • 2607.07769 • Published Jul 8 • 11
view post Post 7758 Frontier models use distillation as a step of their post-training pipelines. In 2026 it has three jobs: compress a big model into a small one, merge RL experts into a single model, and let a model teach itself.I wrote up which frontier models use each one and how: https://huggingface.co/blog/sergiopaniego/distillation-2026It pairs with Class 2 of the Training an Agent series Ben and I are doing, where we teach these techniques hands-on with TRL! See translation 3 replies · 👍 14 14 🔥 7 7 ❤️ 3 3 + Reply