OmniCapBench: A Deep-Structured Evaluation Framework for Fine-Grained Audio-Visual Captioning Paper • 2610.12458 • Published 3 days ago • 15
QuantCode Model: Specializing Language Models for Executable Algorithmic Trading Code Paper • 2609.39420 • Published 11 days ago • 19
ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks Paper • 2609.18805 • Published 25 days ago • 67
Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training Paper • 2608.26730 • Published Aug 27 • 154