Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 3 days ago • 182
OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs Paper • 2608.21360 • Published 6 days ago • 28
Towards Quantifying Benchmark Optimization in ASR Models Paper • 2608.19936 • Published 7 days ago • 11
SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science? Paper • 2608.19799 • Published 7 days ago • 63
EnvHarness: Awakening Static Worlds for Agent Learning Paper • 2608.19880 • Published 7 days ago • 263
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence Paper • 2608.11341 • Published 16 days ago • 64
Pass the Baton: Trajectory-Relayed On-Policy Distillation Paper • 2607.26057 • Published about 1 month ago • 33
Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation Paper • 2607.27372 • Published 29 days ago • 19
Flux-OPD: On-Policy Distillation with Evolving Contexts Paper • 2607.28022 • Published 28 days ago • 44
MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training Paper • 2606.30406 • Published Jun 29 • 25
BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms Paper • 2607.26497 • Published 28 days ago • 51
PhiZero: A World Model Built Around Physical Language Paper • 2607.28624 • Published 28 days ago • 170
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering Paper • 2607.28568 • Published 28 days ago • 185
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published 28 days ago • 309
MemPrivacy: Privacy-Preserving Personalized Memory Management for Edge-Cloud Agents Paper • 2605.09530 • Published May 10 • 150