Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation Paper • 2609.11115 • Published 11 days ago • 171
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments Paper • 2609.04148 • Published 18 days ago • 241
SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation Paper • 2608.21500 • Published Aug 21 • 41
VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End? Paper • 2608.15265 • Published Aug 15 • 60
HRBench: Benchmarking and Understanding Thinking-Mode Switch Strategies in Hybrid-Reasoning LLMs Paper • 2605.28398 • Published May 27 • 14