All three deep research agents now plan, search, read and write a cited report, and all three offer a limited free tier. The real differences are in speed, control, integrations and how you pay once you outgrow the free allowance.
A deep research agent is not a fast chatbot. It takes a question, builds a research plan, reads dozens to hundreds of sources, reasons over them and returns a structured, cited report minutes later.
The three products started at different times and have since changed engines:
Tool | Launched | Engine at launch | Engine today (Sept 2026) |
|---|---|---|---|
December 2024 | Gemini 1.5 Pro | Gemini 3.x family (API agent: Gemini 3.1 Pro) | |
February 2, 2025 | o3 variant trained with RL on browsing and Python | GPT-5.2 since February 10, 2026; “latest models” by default, legacy models selectable | |
February 14, 2025 | Customized DeepSeek R1 | Claude Opus 4.6 (Max first, rolling out to Pro) inside Perplexity’s search and sandbox stack |
The engines keep changing, so the numbers below will age. The structure of each offer is more stable than the figures.
How each tool works today
All three now show you a plan or progress and let you steer. The old split of “Gemini plans, OpenAI grinds, Perplexity searches” no longer holds as cleanly.
ChatGPT Deep Research: the most controllable
You describe the outcome, optionally pick websites or files, and ChatGPT may ask clarifying questions. It then proposes a research plan you can edit before it starts. While it runs you can watch progress, interrupt, refine the focus and change which sources it may use (OpenAI Help).
Since the February 10, 2026 update it runs on GPT-5.2 and can restrict or prioritize specific sites. It reads from connected apps (for example Google Drive or SharePoint) using read-only actions. OpenAI still quotes 5 to 30 minutes per task, depending on complexity.
Gemini Deep Research: built around your Google data
Gemini drafts a multi-step plan, lets you edit it, then searches, reads, self-critiques and writes. Its main advantage is reach into Gmail, Drive and Chat, so research over your own material needs no copy-pasting.
The consumer model has moved from 1.5 Pro to 2.0 Flash Thinking, 2.5 Pro and now the Gemini 3 family. The developer version, released on the Gemini API in April 2026, comes in two tiers: Deep Research and Deep Research Max, both on Gemini 3.1 Pro. It adds code execution, file search and MCP connections.
Perplexity Deep Research: search-native
Perplexity started as a citation-first answer engine, and its agent runs many searches, reads hundreds of sources and writes a report with inline citations. Reports export to PDF, documents or Perplexity Pages.
In February 2026 Perplexity upgraded the agent: first to Claude Opus 4.5, then a week later to Opus 4.6, available to Max subscribers first and rolling out gradually to Pro. Perplexity pairs Anthropic’s model with its own search engine and sandbox. It also publishes its own Sonar models, so it is both a harness around frontier models and a model builder.
What it costs
Every tool has a small free allowance and a roughly $20 paid tier, and each also sells a ~$200 tier for heavy use. Deep research is no longer a $200-only product anywhere.
Tool | Free | ~$20/month tier | Top tier | API |
|---|---|---|---|---|
ChatGPT | Limited (Free and Go plans) | Included in Plus; also Business, Enterprise | Pro: “maximum” access | o3-deep-research at $10 / $40 per million input / output tokens; o4-mini-deep-research cheaper |
Gemini | Limited; about 5 reports a month in Google’s older tables, compute-based limits since May 2026 | Higher limits with Google AI Pro ($19.99) | Highest limits with Google AI Ultra | Google estimates ~$1–3 per task (Deep Research) and ~$3–7 (Max) |
Perplexity | Limited; exact daily allowance not firmly published | Pro ($20): Deep Research included, Advanced mode rolling out | Max ($200): Advanced Deep Research first, highest limits | Sonar API (separate product) |
ChatGPT limits have changed three times. At launch in February 2025, Deep Research was Pro-only, with 100 queries a month. In April 2025 OpenAI set 5 a month for Free, 25 for Plus, Team, Enterprise and Edu, and 250 for Pro, with part of each allowance served by a lighter o4-mini version. In 2026 OpenAI stopped publishing one quota table: usage now varies by plan, an in-product counter shows what remains, and fixed allowances reset every 30 days from first use (OpenAI Help, Wikipedia).
Gemini opened a free tier in March 2025, when Google let non-paying users try Deep Research a few times a month (Search Engine Journal). Note that Deep Think, Gemini’s extended-reasoning mode, is a separate feature from Deep Research and is the one tied to AI Ultra.
On the API, compare API with API. Google’s documentation estimates ~80 searches and ~250k input tokens for a standard task, and up to ~160 searches and ~900k input tokens for Max (Google AI). OpenAI prices its deep research models per token instead. One independent estimate puts an o4-mini-deep-research task at about $0.41, cheaper than Gemini’s standard agent because it runs fewer searches (TokenCost). Treat all of these as estimates that shift with every model update.
Limits and prices here change every few months. Check the current plan page before you budget anything.

Which is most accurate? The benchmark fight
No independent benchmark compares the current versions of all three. The loudest numbers come from vendors, and several were measured on engines that have since been replaced.
Perplexity’s DRACO
Perplexity built DRACO (Deep Research Accuracy, Completeness, and Objectivity) from 100 anonymized real user tasks across 10 domains, each graded against about 40 expert criteria. The paper has a Harvard co-author and the benchmark is open-sourced (arXiv).
System (as tested, early 2026) | Normalized score |
|---|---|
Perplexity Deep Research (Opus 4.6) | 70.5% |
Perplexity Deep Research (Opus 4.5) | 67.2% |
Claude Opus 4.6 (with tools) | 59.8% |
Gemini Deep Research | 59.0% |
OpenAI Deep Research (o3) | 52.1% |
OpenAI Deep Research (o4-mini) | about 42% |
Perplexity’s launch blog reported slightly different figures (67.15% vs 58.97% for Gemini and 52.06% for OpenAI o3), top pass rates of 89.4% in Law and 82.4% in Academic, and an average latency of 459.6 seconds against 592 to 1,808 seconds for the others (Perplexity).
Three caveats matter. Perplexity designed the test and ran it. OpenAI was tested on its o3 and o4-mini versions, not the GPT-5.2 engine it switched to days later. And the strongest non-Perplexity system was plain Claude Opus 4.6, which suggests the harness around a model matters as much as the brand name.
OpenAI’s launch numbers
At launch OpenAI reported 26.6% on Humanity’s Last Exam, against 9.4% for DeepSeek-R1 and 6.2% for Gemini Thinking at the time, plus a 67.36% average on GAIA. These are 2025 results on different tests and cannot be lined up against DRACO.
Independent and informal tests
The independent DeepResearch Bench (mid-2025) found Gemini 2.5 Pro Deep Research and OpenAI Deep Research roughly level and clearly ahead of Perplexity and Grok (arXiv). Informal user tests tell the same moving story: early 2025 reviews such as UX Tigers favored OpenAI, while testers who repeated their comparisons after Google moved to Gemini 2.5 Pro often preferred Gemini.
The honest summary: the lead has changed hands, results depend on when and how people tested, and every system has been upgraded since the most-cited numbers were published.
Output, citations and control
All three cite their sources, all three let you steer, and all three export. They differ in emphasis.
ChatGPT | Gemini | Perplexity | |
|---|---|---|---|
Citations | Citations or source links, a sources-used section and an activity history | Cited reports; Google advises checking citations | Inline citations on every claim, its core identity |
Control before the run | Clarifying questions, editable plan, pick sites and files | Editable research plan | Clarifying questions in Advanced mode |
Control during the run | Watch progress, interrupt, change focus and allowed sources | Progress view; API supports collaborative planning | Follow-up questions mid-run, progress and key findings |
Export | Markdown, Word, PDF | Google Docs in the app; API returns report text | PDF, documents, Perplexity Pages |
Typical run time | 5–30 minutes (OpenAI) | Minutes; API tasks up to 60 minutes | A few minutes; about 8 minutes average in Perplexity’s own DRACO test |
On speed, Perplexity is usually the fastest, but “two to four minutes” is optimistic. Its own benchmark measured about 460 seconds on average, against roughly 10 to 30 minutes for the other agents on the same tasks.
None of these agents is immune to errors: they can state wrong facts or misjudge a source. OpenAI’s documentation says so directly, and Google warns about prompt injection from uploaded files and malicious web pages.
Which tool for which job
Pick by task, not by brand.
Quick, well-cited market or topic scans: Perplexity. It is the fastest, citations are its strength and PDF export is built in. The free allowance is small, so batch work needs Pro.
High-stakes analysis you want to steer: ChatGPT. Clarifying questions, an editable plan, site restrictions, mid-run interruption and Python-based analysis of your files. Save your monthly allowance for the questions that deserve it.
Research over your own material: Gemini, if your sources live in Gmail, Drive and Chat. ChatGPT can also read connected apps such as Google Drive, but Gemini’s Google integration is the most native.
Building research into your own product: compare the Gemini Deep Research API (code execution, MCP, per-task estimates of ~$1–7) with OpenAI’s o3 and o4-mini deep research models (per-token pricing). Run the same tasks on both before choosing.
Questions that grow into reports: Perplexity, where a quick cited answer can turn into a full report with follow-ups.
Tip: metaprompting. Explain your problem to a regular chat model, ask it to write the full research prompt, review it, then give that prompt to the deep research agent. A better brief produces a better report in any of the three.

Limits and risks nobody has solved
Hallucination is not fixed. These are browsing agents on an adversarial web. Citations exist so you can check them, not so you can skip checking.
The comparison data has gaps. There is no independent head-to-head benchmark of the current engines (GPT-5.2, Gemini 3.x, Perplexity on Opus 4.6). Free-tier allowances for Gemini and Perplexity are not published as firm numbers. Nobody has published solid 2026 adoption or satisfaction data.
The ground moves monthly. In 2026 alone, Perplexity upgraded its agent twice in February, OpenAI switched engines and dropped its fixed quota table, and Google launched a developer version with two tiers and moved consumer limits to a compute-based system. Re-test before you commit budget.
People also ask
Which deep research tool is most accurate? No independent benchmark of the current versions settles it. Perplexity’s own DRACO benchmark ranks Perplexity first, but it was designed by Perplexity and tested OpenAI’s older o3 engine. Independent tests in 2025 put OpenAI and Gemini roughly level.
Is ChatGPT Deep Research better than Gemini Deep Research? It depends on the job. ChatGPT gives the most control during a run; Gemini is the most natural fit if your sources live in Google Workspace. Both offer a limited free tier and are included in their ~$20 plans.
How does Perplexity compare with ChatGPT? Perplexity is usually faster and citation-first, with easy PDF export. ChatGPT is slower but offers more steering and deeper file analysis. Perplexity’s paid tiers run on Anthropic’s Claude Opus models inside its own search stack.
Which gives the best citations? Perplexity puts an inline citation on every claim. ChatGPT adds a sources-used section and an activity history so you can audit the process. Gemini cites too. Treat every citation as a lead to verify.
So which one should you pick?
If you live in Google’s ecosystem, start with Gemini. If you want the most control, use ChatGPT. If you want fast, cited, exportable reports, use Perplexity. The best test costs nothing: run one real question from your work through all three free tiers, compare the reports, and keep the one whose citations survive your checking.

