EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction Paper • 2609.02783 • Published 8 days ago • 118
CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild Paper • 2608.23181 • Published 17 days ago • 34
Second Thought: Reasoning in Parallel as LLM Agents Act and Observe Paper • 2608.13667 • Published 27 days ago • 17
How Do Practitioners Build SE Agents? Insights from a Mixed-Methods Study Paper • 2607.10856 • Published Jul 12 • 7
Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents? Paper • 2607.01211 • Published Jul 1 • 14
Executing as You Generate: Hiding Execution Latency in LLM Code Generation Paper • 2604.00491 • Published Apr 1 • 6
Rethinking the Value of Agent-Generated Tests for LLM-Based Software Engineering Agents Paper • 2602.07900 • Published Feb 8 • 4
CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding Paper • 2602.01785 • Published Feb 2 • 97