-
DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
Paper • 2511.19399 • Published • 63 -
ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
Paper • 2511.07685 • Published • 10 -
OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment
Paper • 2510.07743 • Published • 14
Collections
Discover the best community collections!
Collections including paper arxiv:2511.19399
-
DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
Paper • 2511.19399 • Published • 63 -
rl-research/DR-Tulu-8B
Text Generation • 8B • Updated • 3.1k • 73 -
rl-research/DR-Tulu-SFT-8B
Text Generation • 8B • Updated • 333 • 5 -
rl-research/dr-tulu-sft-data
Viewer • Updated • 13.1k • 249 • 28
-
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
Paper • 2509.07980 • Published • 105 -
Robot Learning from a Physical World Model
Paper • 2511.07416 • Published • 32 -
MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning
Paper • 2511.06805 • Published • 13 -
GigaEvo: An Open Source Optimization Framework Powered By LLMs And Evolution Algorithms
Paper • 2511.17592 • Published • 121
-
Diffusion Augmented Agents: A Framework for Efficient Exploration and Transfer Learning
Paper • 2407.20798 • Published • 24 -
Offline Reinforcement Learning for LLM Multi-Step Reasoning
Paper • 2412.16145 • Published • 38 -
REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models
Paper • 2501.03262 • Published • 104 -
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Paper • 2502.18449 • Published • 75
-
Fathom-DeepResearch: Unlocking Long Horizon Information Retrieval and Synthesis for SLMs
Paper • 2509.24107 • Published • 80 -
Beyond Turn Limits: Training Deep Search Agents with Dynamic Context Window
Paper • 2510.08276 • Published • 10 -
DeepWideSearch: Benchmarking Depth and Width in Agentic Information Seeking
Paper • 2510.20168 • Published • 28 -
Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics
Paper • 2510.17797 • Published • 11
-
microsoft/bitnet-b1.58-2B-4T
Text Generation • 0.8B • Updated • 15.5k • 1.44k -
M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models
Paper • 2504.10449 • Published • 15 -
nvidia/Llama-3.1-Nemotron-8B-UltraLong-2M-Instruct
Text Generation • 8B • Updated • 122 • 17 -
ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
Paper • 2504.11536 • Published • 63
-
DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
Paper • 2511.19399 • Published • 63 -
ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
Paper • 2511.07685 • Published • 10 -
OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment
Paper • 2510.07743 • Published • 14
-
DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
Paper • 2511.19399 • Published • 63 -
rl-research/DR-Tulu-8B
Text Generation • 8B • Updated • 3.1k • 73 -
rl-research/DR-Tulu-SFT-8B
Text Generation • 8B • Updated • 333 • 5 -
rl-research/dr-tulu-sft-data
Viewer • Updated • 13.1k • 249 • 28
-
Fathom-DeepResearch: Unlocking Long Horizon Information Retrieval and Synthesis for SLMs
Paper • 2509.24107 • Published • 80 -
Beyond Turn Limits: Training Deep Search Agents with Dynamic Context Window
Paper • 2510.08276 • Published • 10 -
DeepWideSearch: Benchmarking Depth and Width in Agentic Information Seeking
Paper • 2510.20168 • Published • 28 -
Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics
Paper • 2510.17797 • Published • 11
-
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
Paper • 2509.07980 • Published • 105 -
Robot Learning from a Physical World Model
Paper • 2511.07416 • Published • 32 -
MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning
Paper • 2511.06805 • Published • 13 -
GigaEvo: An Open Source Optimization Framework Powered By LLMs And Evolution Algorithms
Paper • 2511.17592 • Published • 121
-
microsoft/bitnet-b1.58-2B-4T
Text Generation • 0.8B • Updated • 15.5k • 1.44k -
M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models
Paper • 2504.10449 • Published • 15 -
nvidia/Llama-3.1-Nemotron-8B-UltraLong-2M-Instruct
Text Generation • 8B • Updated • 122 • 17 -
ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
Paper • 2504.11536 • Published • 63
-
Diffusion Augmented Agents: A Framework for Efficient Exploration and Transfer Learning
Paper • 2407.20798 • Published • 24 -
Offline Reinforcement Learning for LLM Multi-Step Reasoning
Paper • 2412.16145 • Published • 38 -
REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models
Paper • 2501.03262 • Published • 104 -
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Paper • 2502.18449 • Published • 75