Post-Training Language Models for Gold-Medal Performance in Coding Competitions Paper • 2609.02849 • Published 2 days ago • 8
Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Model Paper • 2607.22083 • Published Jul 27 • 10
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability Paper • 2608.30320 • Published 4 days ago • 50
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published 4 days ago • 134
Self-Improving Pretraining: using post-trained models to pretrain better models Paper • 2601.21343 • Published Jan 29 • 21
A Programming Paradigm for Spatiotemporal Composability Paper • 2608.25512 • Published 9 days ago • 14
D^3-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation Paper • 2608.24987 • Published 10 days ago • 25
Agent-G^2: Gaussian Guidance for Agentic Reinforcement Learning Paper • 2608.23318 • Published 11 days ago • 30
Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher Paper • 2608.26872 • Published 8 days ago • 80
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 11 days ago • 205
Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation Paper • 2608.15062 • Published 9 days ago • 10