Buyers don't need another glossy demo. They need proof that an agent can chase a thread through messy data, make choices without babysitting, and return something your CFO won't tear apart. Benchmarks step into that gap—when they actually measure the right things.
Enter DRACO, the deep research benchmark quietly shaping 2025 AI budgets. It doesn't ask trick trivia. It simulates what your team does under pressure: synthesize cross-domain facts, reason through conflicts, and justify every step. And yes, speed counts—because meetings don't wait and markets don't pause.
On DRACO, Perplexity Deep Research posts a 67.15% mean score, topping Gemini Deep Research at 58.97% and OpenAI o3 at 52.06 across ten domains—technology, law, general knowledge, more. The gaps aren't noise; they're operating realities you'll feel as fewer escalations, tighter briefs, and decisions that don't wobble.
Benchmarks don't buy software—operators do. Still, when the spread hits double digits in core domains (up to 11.2 points in general knowledge, 9.8 in technology), the math nudges you. Time saved compounds. Risk avoided compounds faster.
Latency tips the scale. Perplexity turns results in 459.6 seconds on average—even with heavier token use—so orchestration flows move without stalling. That's not trivia. That's whether your sales research agent lands insights before the rep walks into the room.
There's a wrinkle: finance remains a soft spot across models. It's where nuance, time-series volatility, and source fragility collide. Smart buyers don't pretend it's solved; they route around it with tighter retrieval, audit hooks, and, bluntly, a human in the loop when dollars move.