Running 43 DeepResearch Bench 🔍 43 Explore rankings of Deep Research agents in a searchable leaderboard
MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing Paper • 2607.27616 • Published 8 days ago • 38
Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation Paper • 2607.27816 • Published 8 days ago • 33
BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms Paper • 2607.26497 • Published 8 days ago • 49
Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory Paper • 2607.24368 • Published 11 days ago • 33
Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory Paper • 2607.24368 • Published 11 days ago • 33
LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues Paper • 2605.12493 • Published May 12 • 5
Running 43 DeepResearch Bench 🔍 43 Explore rankings of Deep Research agents in a searchable leaderboard
Stream-T1: Test-Time Scaling for Streaming Video Generation Paper • 2605.04461 • Published May 6 • 109
Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation Paper • 2605.03849 • Published May 5 • 129
Running 43 DeepResearch Bench 🔍 43 Explore rankings of Deep Research agents in a searchable leaderboard