SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure Paper • 2608.11079 • Published 1 day ago • 5
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? Paper • 2608.10366 • Published 1 day ago • 1
The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows Paper • 2608.06714 • Published 6 days ago • 9
Characterizing the Quality Profile of AI-Generated C++ in Production Paper • 2608.06640 • Published 7 days ago • 10
Modular TTT: Rethinking Test-Time Training as Composable Modules Paper • 2608.07110 • Published 6 days ago • 8
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Paper • 2608.05747 • Published 7 days ago • 46
HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Paper • 2608.06301 • Published 7 days ago • 34
DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation Paper • 2608.06374 • Published 7 days ago • 23
ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation Paper • 2608.04436 • Published 8 days ago • 58
HelloWorld: Enabling Socially Interactive Characters in Video World Models Paper • 2608.05070 • Published 8 days ago • 37
OPD-V: Visual On-Policy Self-Distillation with Modality Balance Paper • 2608.05131 • Published 7 days ago • 12
PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents Paper • 2608.04003 • Published 8 days ago • 33
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion Paper • 2608.03974 • Published 9 days ago • 91
Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing Paper • 2608.02711 • Published 9 days ago • 88