Skip to content

Latest commit

 

History

125 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

Awesome Deep Research Agent

We maintain a curated collection of papers exploring the path towards Deep Research (DR) Agents, focusing on formulating core concepts and mapping the research landscape.

⌛️ We’re continuously compiling and updating cutting‑edge insights. Feel free to suggest any related work you find valuable!

Build a digital assistant on your screen. Generated by DALL-E-3.

🔥 WELCOME CONTRIBUTE!

🔥 This project is actively maintained, and we welcome your contributions. If you have any suggestions, such as missing papers or information, please feel free to open an issue or submit a pull request.

📰 News

  • [2025.09.03] 📄 Updated version of our paper "Deep Research Agents: A Systematic Examination And Roadmap" is now available on ArXiv with expanded analysis and refined research directions!
  • [2025.06.18] 🚀 Our comprehensive survey "Deep Research Agents: A Systematic Examination And Roadmap" is officially released on ArXiv - providing systematic insights into the current state and future of deepresearch agents!

🏗️ Our Works Towards DR Agents

✨✨✨ Deep Research Agents: A Systematic Examination And Roadmap

Structural overview of a DR agent An overview of a DR agent

📚 Awesome Papers

Table of Contents

  1. Search Engine Integration
  2. Tool Use
  3. Architecture & Workflow
  4. Tuning Methods
  5. Industrial Applications
  6. Benchmarks for DR Agents

Search Engine Integration

📊 Search Engine · API vs Browser Comparison
Legend ✔️ Primary focus 🟫 Secondary/minor focus — Not present
DR Agent API Browser GAIA HLE QA Base Model
Avatar 🟫 Stark Claude-3-Opus, GPT-4
CoSearch-Agent ✔️ GPT-3.5-turbo
MMAC-Copilot ✔️ ✔️ GPT-3.5, GPT-4
Storm 🟫 FreshWiki GPT-3.5-turbo
OpenResearcher ✔️ Private QA DeepSeek-V2-Chat
The AI Scientist ✔️ MLE-Bench GPT-4o, o1-mini, o1-preview
Gemini DR ✔️ ✔️ ✔️ GPQA Gemini-2.0-Flash
Agent Laboratory ✔️ MLE-Bench GPT-4o, o1-preview
Search-o1 ✔️ GPQA·NQ·TriviaQA QwQ-32B-preview
WebWalker WebWalkerQA GPT-4o, Qwen-2.5
Agentic Reasoning ✔️ GPQA DeepSeek-R1, Qwen2.5
AutoAgent ✔️ ✔️ Claude-Sonnet-3.5
Grok DeepSearch ✔️ ✔️ GPQA Grok 3
OpenAI DR ✔️ ✔️ ✔️ ✔️ GPT-o3
Perplexity DR ✔️ 🟫 ✔️ SimoleQA Flexible
Towards an AI Co-Scientist ✔️ GPQA Gemini 2.0
Nouswise
AgentRxiv ✔️ GPQA·MedQA GPT-4o-mini
Agent-R1 ✔️ HotpotQA Qwen2.5-1.5B-Inst
AutoGLM Rumination ✔️ GPQA GLM-Z1-Air
Copilot Researcher ✔️ o3-mini
H2O.ai DR ✔️ ✔️ ✔️ h2ogpt-oasst1-512-12b
Manus ✔️ ✔️ Claude3.5, GPT-4o
OpenManus ✔️ ✔️ Claude3.5, GPT-4o
OWL ✔️ ✔️ ✔️ DeepSeek-R1, Gemini-2.5-Pro, GPT-4o
R1-Searcher 🟫 2WikiMultiHopQA, HotpotQA Llama3.1-8B-Inst, Qwen2.5-7B
ReSearch 🟫 2WikiMultiHopQA, HotpotQA Qwen2.5-7B, Qwen2.5-7B-Inst
Search-R1 🟫 2WikiMultiHopQA, HotpotQA, NQ, TriviaQA Llama3.2-3B, Qwen2.5-3B/7B
DeepResearcher ✔️ HotpotQA, NQ, TriviaQA Qwen2.5-7B-Inst
Genspark Super Agent ✔️ ✔️ ✔️ Mixture of 9 LLMs
WebThinker ✔️ ✔️ ✔️ GPQA, WebWalkerQA QwQ-32B
SWIRL ✔️ HotQA, BeerQA Gemma 2-27B
SimpleDeepSearcher ✔️ ✔️ 2WikiMultiHopQA Qwen-2.5-7B/32B-In, DeepSeek-D-Qwen-2.5-32B, QwQ-32B
Suna AI ✔️ ✔️ GPT-4o, Claude
AgenticSeek ✔️ GPT-4o, DeepSeek-R1, Claude
Alita ✔️ ✔️ ✔️ PathVQA GPT-4o, Claude-Sonnet-4
DeerFlow ✔️ Doubao-1.5-Pro-32k, DeepSeek-R1, GPT-4o, Qwen
PANGU DEEPDIVER ✔️ C-SimpleQA, HotpotQA, ProxyQA Pangu-7B-Reasoner
WebDancer ✔️ ✔️ GAIA, WebWalkerQA Qwen-2.5, QwQ-32B, DeepSeek-R1, GPT-4o
O-agents ✔️ ✔️ GPT-4o, GPT-4.1, Claude-3.7-Sonnet, DeepSeek-R1, Gemini-2.5
Kimi-Researcher ✔️ ✔️ ✔️ SimpleQA Kimi k1.5/k2
WebSailor ✔️ ✔️ SimpleQA Qwen-2.5
Agent-KB ✔️ ✔️ SWE-bench GPT-4o, GPT-4.1, Claude-3.7-Sonnet, o3-mini, Qwen-3, DeepSeek-R1
WebShaper ✔️ ✔️ WebWalkerQA Qwen-2.5, QwQ-32B
Deep Researcher with Test-Time Diffusion ✔️ ✔️ ✔️ Gemini-2.5-Pro
ChatGPT-Agent
AWorld ✔️ ✔️ ✔️ HotpotQA Gemini-2.5-Pro, GPT-4o
Cognitive Kernel-Pro ✔️ ✔️ ✔️ AgentWebQA, WebWalkerQA, Multi-hop URLQA, DocBench, TableBench Claude-3.7-Sonnet, CK-Pro-8B
WebWatcher ✔️ ✔️ Browsercom-VL, LiveVQA, MMSearch Qwen-2.5-VL-32B
WideSearch ✔️ WideSearch DeepSeek-R1, Doubao-Seed-1.6, Claude Sonnet 4, Gemini-2.5-Pro
MiroRL ✔️ ✔️ Qwen3-14B

Tool Use 

📊 Tool Use Capabilities Comparison
Legend ✔️ Involved 🟫 Non Disclosure — Not present
DR Agent Code Interp. Data Analytics Multimodal Release
CoSearchAgent ✔️ Feb-2024
Storm ✔️ Jul-2024
The AI Scientist ✔️ Aug-2024
Agent Laboratory ✔️ Jan-2025
Agentic Reasoning ✔️ Feb-2025
AutoAgent ✔️ ✔️ Feb-2025
Genspark DR ✔️ ✔️ ✔️ Feb-2025
Grok DeepSearch ✔️ ✔️ ✔️ Feb-2025
OpenAI DR ✔️ ✔️ ✔️ Feb-2025
Perplexity DR ✔️ ✔️ ✔️ Feb-2025
Towards an AI co-scientist ✔️ ✔️ Feb-2025
Agent-R1 ✔️ Mar-2025
AutoGLM Romination ✔️ ✔️ Mar-2025
Copilot Researcher ✔️ ✔️ 🟫 Mar-2025
Manus ✔️ ✔️ ✔️ Mar-2025
OpenManus ✔️ ✔️ Mar-2025
OWL ✔️ ✔️ ✔️ Mar-2025
H2O.ai DR ✔️ ✔️ ✔️ Mar-2025
Genspark Super Agent ✔️ ✔️ ✔️ Apr-2025
WebThinker ✔️ Apr-2025
Suna Ai ✔️ ✔️ Apr-2025
Tool-Star ✔️ ✔️ May-2025
AgenticSeek ✔️ ✔️ May-2025
Alita ✔️ 🟫 🟫 May-2025
DeerFlow ✔️ ✔️ May-2025
O-agents ✔️ ✔️ ✔️ Jun-2025
Kimi-Researcher ✔️ ✔️ Jun-2025
Agent-KB ✔️ ✔️ ✔️ Jul-2025
AWorld ✔️ ✔️ ✔️ Jul-2025
Cognitive Kernel-Pro ✔️ ✔️ ✔️ Aug-2025
WebWatcher ✔️ ✔️ ✔️ Aug-2025
MiroRL ✔️ ✔️ Aug-2025

Architecture & Workflow

Architecture & Workflow

Static Workflow

Dynamic Single‑Agent Workflow

Dynamic Multi‑Agent Workflow

Tuning Methods

📊 Tuning Methods Comparison
Legend ✔️ Implemented 🟫 Details Unknown — Not present
DR Agent SFT RL Base Model Data Reward Design
Gemini DR 🟫 🟫 Gemini-2.0-Flash 🟫
WebWalker GPT-4o, Qwen-2.5 (7–72B) WebWalkerQA
Grok DeepSearch 🟫 Grok 3 🟫
OpenAI DR 🟫 GPT-o3 🟫
Agentic Reasoning ✔️ DeepSeek-R1, Qwen2.5 GPQA Rule-Outcome
AutoAgent ✔️ Claude-Sonnet-3.5
Towards an AI co-scientist Gemini 2.0
Agent-R1 PPO · Reinforce++ · GRPO Qwen2.5-1.5B-Inst HotpotQA Rule-Outcome
AutoGLM Rumination 🟫 🟫 GLM-Z1-Air 🟫
H2O.ai DR ✔️ 🟫 h2ogpt-oasst1-512-12b 🟫
Copilot Researcher 🟫 🟫 o3-mini
ReSearch GRPO Qwen2.5-7B-Inst · Qwen2.5-32B-Inst 2WikiMultiHopQA Rule-Outcome
R1-Searcher ✔️ Reinforce++ · GRPO Qwen2.5-7B-Inst / LLaMA-3.1-8B-Inst 2WikiMultiHopQA · HotpotQA Rule-Outcome
Search-R1 ✔️ PPO · GRPO Qwen2.5-3B/7B · LLaMA3.2-3B-Inst NQ · HotpotQA Rule-Outcome
Nouswise 🟫 🟫 Nouswise 🟫
DeepResearcher GRPO Qwen2.5-7B-Inst NQ · HotpotQA Rule-Outcome
Genspark Super Agent 🟫 Mixture of Agents 🟫
WebThinker ✔️ Iterative Online DPO QwQ-32B Expert Dataset Rule-Outcome
SWIRL Offline-RL Gemma 2-27B HotPotQA
SimpleDeepSearcher ✔️ PPO Qwen-2.5-7B-In · Qwen-2.5-32B-In · Deepseek-Distilled-Qwen-32B · QwQ-32B NQ · HotpotQA · 2WikiMultiHopQA · Musique · SimpleQA · MultiHop-RAG Process-based reward
PANGU DEEPDIVER ✔️ GRPO Pangu-7B-Reasoner WebPuzzle Rule-Outcome
Tool-Star ✔️ GRPO Qwen-2.5 NuminaMath · HotpotQA · 2WikiMultiHopQA Rule-Outcome
WebDancer ✔️ DAPO Qwen-2.5-7B/32B · QwQ-32B · DeepSeek-R1 · GPT-4o CRAWLQA · E2HQA Rule-Outcome
O-agents GPT-4o · GPT-4.1 · Claude-3.7-Sonnet · DeepSeek-R1 · Gemini-2.5
Kimi-Researcher REINFORCE Kimi k1.5/k2 Rule-Outcome
WebSailor ✔️ DUPO Qwen-2.5-3B/7B/32B/72B SailorFog-QA Rule-Outcome
Agent-KB GPT-4o · GPT-4.1 · Claude-3.7-Sonnet · o3-mini · Qwen-3 32B · DeepSeek-R1
WebShaper ✔️ GRPO Qwen-2.5-3B/7B/32B/72B · QwQ-32B WebShaper Rule-Outcome
Cognitive Kernel-Pro ✔️ Claude-3.7-Sonnet · CK-Pro-8B OpenWebVoyager · Multi-hop URLQA · AgentWebQA · WebWalkerQA · DocBench · TableBench
WebWatcher GRPO Qwen-2.5-VL-32B BrowseComp-VL · Long-tail VQA · Hard VQA Rule-Outcome
MiroRL ✔️ GRPO Qwen3-14B MiroRL-GenQA Rule-Outcome
Atom-Searcher ✔️ GRPO Qwen2.5-7B-Inst 2WikiMultiHopQA · HotpotQA Atomic Thought Reward (ATR)

Benchmarks for DR Agents

📊 QA Benchmarks (Hotpot / 2Wiki / NQ / TQ / GPQA)
DR Agent Base Model Hotpot 2Wiki NQ TQ GPQA Release
Search-o1 QwQ-32B-preview 57.3 71.4 49.7 74.1 57.9 Jan-2025
Agentic Reasoning DeepSeek-R1, Qwen2.5 67.0 Feb-2025
Grok DeepSearch Grok3 84.6 Feb-2025
AgentRxiv GPT-4o-mini 41.0 Mar-2025
R1-Searcher Qwen2.5-7B-Base 71.9 63.8 Mar-2025
ReSearch Qwen2.5-32B-Inst 67.7 50.0 Mar-2025
Search-R1 Qwen2.5-7B-Inst 34.5 36.9 40.9 55.2 Mar-2025
DeepResearcher Qwen2.5-7B-Inst 64.3 66.6 61.9 85.0 Apr-2025
WebThinker QwQ-32B 68.7 Apr-2025
SimpleDeepSearch QwQ-32B 73.5 Apr-2025
SWIRL Gemma 2-27B 72.0 Apr-2025
Tool-Star Qwen2.5-3B 51.9 40.0 May-2025
📊 GAIA (Test and Val) Benchmarks
DR Agent Base Model GAIA L-1 L-2 L-3 Ave. Release Split
MMAC-Copilot GPT-3.5, GPT-4 45.16 20.75 6.12 25.91 Mar-2024 Test
H2O.ai DR Claude-3.7-Sonnet 89.25 79.87 61.22 79.73 Mar-2025 Test
Alita Claude-Sonnet-4, GPT-4o 92.47 71.70 55.10 75.42 May-2025 Test
Agent-KB GPT-4.1, Claude-3.7 84.91 74.42 57.69 75.15 Jul-2025 Test
O-agents Claude-3.7 83.02 74.42 53.85 73.93 Jun-2025 Test
WebDancer QwQ-32B 61.5 50.0 25.0 51.5 May-2025 Test
WebShaper Qwen-2.5-72B 69.2 63.4 16.6 60.1 Jul-2025 Test
Deep Researcher w/ Test-Time Diffusion Gemini-2.5-Pro 69.1 Jul-2025 Test
Cognitive Kernel-Pro Claude-3.7-Sonnet 83.02 68.60 53.85 70.91 Aug-2025 Test
AutoAgent Claude-Sonnet-3.5 71.7 53.5 26.9 55.2 Feb-2025 Dev
OpenAI DR GPT-o3-customized 78.7 73.2 58.0 67.4 Feb-2025 Dev
Manus Claude 3.5, GPT-4o 86.5 70.1 57.7 71.4 Mar-2025 Dev
OWL Claude-3.7-Sonnet 84.9 68.6 42.3 69.7 Mar-2025 Dev
H2O.ai DR h2ogpt-oasst1-512-12b 67.92 67.44 42.31 63.64 Mar-2025 Dev
Genspark Super Agent Claude 3 Opus 87.8 72.7 58.8 73.1 Apr-2025 Dev
WebThinker QwQ-32B 53.8 44.2 16.7 44.7 Apr-2025 Dev
SimpleDeepSearch QwQ-32B 50.5 45.8 13.8 43.9 Apr-2025 Dev
Alita Claude-Sonnet-4, GPT-4o 75.15 87.27 May-2025 Dev

📄 Citation

If you find this work helpful, please cite our paper:

@article{huang2025deep,
  title={Deep Research Agents: A Systematic Examination And Roadmap},
  author={Huang, Yuxuan and Chen, Yihang and Zhang, Haozheng and Li, Kang and Fang, Meng and Yang, Linyi and Li, Xiaoguang and Shang, Lifeng and Xu, Songcen and Hao, Jianye and others},
  journal={arXiv preprint arXiv:2506.18096},
  year={2025}
}

About

No description, website, or topics provided.

Resources

Stars

638 stars

Watchers

17 watching

Forks

Releases

Packages

Contributors