DRACO grades "complex deep research tasks spanning 10 domains," and its authors call out what it actually values: accuracy, completeness, and objectivity across end-to-end work. That lens punishes hand-wavy output and rewards careful sourcing and cross-domain reasoning. It's not a parlor game; it's enterprise reality distilled.
Here's the scoreboard that shook up 2025: Perplexity Deep Research leads at 70.5% (Opus 4.6), with Opus 4.5 at 67.2%. Gemini's deep-research preview lands around 59.0%, while OpenAI's o3 deep-research build trails at 52.1%. The biggest gaps show up where buyers actually make money—finance (+21.6 points), shopping/product comparison (+10.9), technology (+9.8), and academic research (+9.3).
"Speed is a tactic, accuracy is a strategy."
There's a cost for that lead: Perplexity gulps a hefty average of 778,711 input tokens, yet still posts superior accuracy. OpenAI o3, by contrast, shows the slowest stride at 1,808.1 seconds and mid-tier output around 24,944 tokens.
Translate that into ROI. On workflows like expert reports and market comparisons, a 20–30% lift in task completion rate means fewer escalations, fewer rewrites, and a shorter path to publication-grade deliverables. The org feels it first in analyst bandwidth, then in cycle time to decisions. You're buying accuracy to buy back time.
A mini scorecard you can steal
- Accuracy and completeness: weighted 40%. Judge against domain truth, not vibes.
- Objectivity and citation quality: 25%. If you can't trace it, you can't ship it.
- Token efficiency: 15%. Pay compute where it saves headcount, not everywhere.
- Latency: 10%. Batches forgive; live teams don't. Segment by use case.
- Cross-domain stability: 10%. Finance today, tech tomorrow—same rails or bust.
One more note buyers keep repeating: orchestration beats one-off cleverness. A "super" agent that dazzles solo but can't hand off to retrieval, enrichment, and verification at scale gets lonely in production. Joe's Site has seen teams ship faster when they standardize traces, prompts, and QA gates across every agent—even when models differ under the hood.