ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition Paper • 2607.25565 • Published 3 days ago • 58
ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition Paper • 2607.25565 • Published 3 days ago • 58
From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models Paper • 2607.06553 • Published 22 days ago • 20
SPACENUM: Revisiting Spatial Numerical Understanding in VLMs Paper • 2605.23898 • Published May 22 • 7
Reward-Decomposed Reinforcement Learning for Immersive Video Role-Playing Paper • 2605.04733 • Published Jun 3
On-Policy Distillation with Best-of-N Teacher Rollout Selection Paper • 2605.09725 • Published May 13
CocoaBench: Evaluating Unified Digital Agents in the Wild Paper • 2604.11201 • Published Apr 13 • 37
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model Paper • 2603.22281 • Published Mar 23 • 20
Vision Language Models Cannot Reason About Physical Transformation Paper • 2603.07109 • Published Mar 7 • 2
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks Paper • 2602.12670 • Published Feb 13 • 65
Less is Enough: Synthesizing Diverse Data in Feature Space of LLMs Paper • 2602.10388 • Published Feb 11 • 246
Team RAS in 11th ABAW Competition: Multimodal Ambivalence Recognition Approach Paper • 2607.14702 • Published 15 days ago
Team LEYA in 10th ABAW Competition: Multimodal Ambivalence/Hesitancy Recognition Approach Paper • 2603.12848 • Published Mar 13