JBrightmanAI/Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSORED 9B • Updated 2 days ago • 20 • 1
Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms Paper • 2607.07769 • Published 12 days ago • 8
view post Post 7631 Frontier models use distillation as a step of their post-training pipelines. In 2026 it has three jobs: compress a big model into a small one, merge RL experts into a single model, and let a model teach itself.I wrote up which frontier models use each one and how: https://huggingface.co/blog/sergiopaniego/distillation-2026It pairs with Class 2 of the Training an Agent series Ben and I are doing, where we teach these techniques hands-on with TRL! See translation 3 replies · 👍 14 14 🔥 7 7 ❤️ 3 3 + Reply
yuxinlu1/gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2 Text Generation • 12B • Updated 20 days ago • 127k • 69
yuxinlu1/gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-GGUF Text Generation • 12B • Updated about 1 month ago • 535k • 1.23k
yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF Text Generation • 12B • Updated about 1 month ago • 438k • 2.74k