October 7, 2026 · Cole Dermott
The cheapest model with data beat the best model without it

Claude Haiku is one of the cheapest models you can run an agent on. Claude Opus is one of the most capable. We gave both the same work tasks, the kind where the answer has to come from real data. Haiku with access to paid data finished 60% of them. Opus without it finished 50%. Per finished task, they cost about the same.
That held across the board. We tested nine models from five labs on 82 tasks, and every model did better with data than the best model did without it. The weakest model with data, Grok 4.7, finished 54%. That edge over Opus is small enough to be noise, but the direction never flipped.
Most of the conversation about agents is about which model reasons best. For a lot of the work companies actually want agents to do, we think that matters less than people assume. The work runs into a wall that has nothing to do with reasoning: the answer lives behind a login, an API key or a paywall, and the agent can't get to it.
We run Locus Pro, which sells exactly this kind of access, so you should read this with that in mind. That's also why we wrote down the method and froze it before running anything, why we report the results that make us look bad, and why every task, grader and trace is public so you can rerun it.

What we tested
The tasks were the kind of thing people ask agents to do at work. Find the CEO of Pleo and a work email that will actually deliver. Get the exact Google Maps review count for Dishoom Covent Garden. Find the cheapest nonstop from San Francisco to New York on November 5. People who never saw our catalog wrote them, and we checked every answer against a source we trusted, at the time each run was graded.
Each model ran every task three times in three setups. The first had no tools, which tells you what a model already knows. The second was a standard agent with web search, a page fetcher and Python. The third was that same agent with one Locus Pro key, which reaches thousands of paid APIs through one MCP server. We also wired seven well-known vendors directly into the same agent, to see whether a key like ours adds anything over doing it yourself.
The models were Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5, GPT-6.1 Sol and GPT-6 Luna, Gemini 3.1 Pro and 3.8 Flash, Grok 4.7, and DeepSeek V4 Pro. That's about 7,700 graded runs.
1. Data mattered more than the model
Without paid data, the nine models landed between 34% and 50% on the data tasks, a 16-point spread from worst to best. Adding data was worth 25 points on average, and every model gained at least 15.
Put another way, the gap between the best and worst model was smaller than what data added to any single one. Haiku with data cost about 19 cents per finished task, against about 18 for Opus without it. GPT-6 Luna and Gemini 3.8 Flash with data also beat every model without it. If your agent keeps failing at this kind of task, we'd try giving it data before paying for a bigger model.
2. Without data, agents guess
Contact research showed the gap most clearly. Agents without paid data found the right person 82% of the time. Then they needed an email, and without a contact database they had to guess one or copy whatever turned up on a public page. 45% of those addresses weren't deliverable. With paid data, 91% of the emails for the right person were deliverable, and success on these tasks went from 41% to 82%.
Flights were similar. A standard agent can find flights, but it rarely finds a fare it can stand behind, and success went from 10% to 43% with data. That 43% is a floor. Our grader required each fare to be within 10% of the cheapest fare any agent reported, and a couple of web-search agents reported fares that were approximate or out of date. With the grading corrections described under Limits, it's 51%. Counting every answer that met the route, date and price limits, it's 72%. Numbers behind logins and bot walls, like Google Ads search volume, Maps review counts, Amazon prices and X follower counts, went from 29% to 61%.
When agents couldn't get the data, they also wasted a lot of effort. On the data tasks, 32% of standard-agent runs used all 30 turns without answering, against 15% with data.

3. Access doesn't fix reasoning
Data only helped where data was the problem. Multi-step research jobs, where the hard part is putting several findings together, stayed around 65% with or without it. Pages that block scrapers stayed in the mid 90s, because the models already handled them. On 15 control questions that plain search can answer, success went from 92% to 91%. Adding a catalog of 16,000 endpoints didn't distract the agents, but it didn't make them smarter either.
4. Agents won't pay for data unless you tell them to
In our first tests, we connected the paid tools to six popular agent harnesses and asked them the same questions. Five of the six ignored the new tools and kept using their built-in web search.
What fixed this was a short set of instructions, packaged as a skill, telling the agent when paid data is worth it. With the skill, Claude Code went from 42% to 52% on the data tasks compared with the tools alone, and OpenClaw went from 52% to 72%. Across five harnesses as they ship, adding data with the skill raised success by 27 to 38 points. Claude Code went from 25% to 52%, Codex from 46% to 83%, Hermes from 46% to 73%, OpenClaw from 35% to 72% and the OpenAI Agents SDK from 60% to 90%.

5. Once agents can pay, spending gets lumpy
Data costs money, and we don't want to hide that. On the data tasks, agents spent about 8 cents per run on data on average, and cost per finished task went up for every model. Sonnet 5.5 went from about 13 cents to 25 cents per finished task.
Some of that cost comes back in a place you might not look. Agents with data made fewer tool calls (8 per run against 11) and ran out of turns less than half as often. Each run's model bill was a little higher, about 8 cents against 6, because the tool list and the data coming back make every turn longer. But far more of those runs finished, so the model bill per finished task fell from about 15 cents to 11, and it fell for eight of the nine models. The data itself cost more than that saved. If you only watch your model bill, though, adding data looks cheaper, not more expensive.

The way they spent is the more useful finding. We gave agents five open-ended tasks with a stated budget between $1 and $3. About half the runs with paid access spent nothing. A smaller group spent heavily: 16 of 129 used more than half their budget, and one used $2.95 of $3. A single research call cost $1.62. With the vendors wired directly, only 1 of 135 runs went past half.
No agent went over budget, and none actually hit a hard cap. The ones near the line stopped on their own. We'd still set one. Average spend per task tells you very little when the spread is this wide, and a per-run budget is cheap insurance against the expensive tail.

Where Locus Pro fits
The obvious question is whether any of this needs Locus, or just the vendors behind it. So we wired seven vendors directly: Apollo, Hunter, Prospeo, Firecrawl, Exa, Tavily and E2B, using their own MCP servers where they have one.
On the 34 data tasks those vendors can serve, success was the same: 76% with Locus and 75% with the vendors. The vendors are where the data comes from. Locus doesn't make that data better.

The difference is what it takes to run. Wiring the vendors directly means seven accounts and seven keys, and their tool definitions add 79,648 tokens to every request before the agent reads the task. With Locus it's one key and 17,104 tokens, one balance, and per-run budgets. For most models, wiring the vendors also cost a little less per finished task in our estimate, although that estimate leaves out the monthly plans most of those vendors require.

What it found in our own product
Running this benchmark exposed a problem in Locus Pro, and we fixed it before the main run. Tool search was ranking obscure scrapers above our main integrations. Fixing the ranking took the right tool from first place in 4 of 18 held-out questions to 9 of 18.
Limits
- We make Locus Pro and ran this ourselves. The method was frozen before the run, every deviation is logged in the repo, and you can rerun everything with your own keys.
- Live data changes, so we captured each answer within ten minutes of the run that used it.
- After the run we found three grading mistakes, mostly by looking at the tasks where data seemed to hurt. Our flight check let stale or approximate fares from other agents set the price bar, and sometimes read the user's budget ("under your $350 limit") as the fare. One contact task accepted only a company's new email domain, though its old domain still delivers. One research task asked the judge to check roles against a careers-page snapshot that we never showed it. With those fixed, the gain is 27 points instead of 25. We kept the pre-registered 25 as the headline, and the repo has both.
- Every run had a 30-turn limit. It hurt agents without data more, and we report how often each setup hit it.
- We also ran Gemini CLI in the harness comparison. Its runs without paid data could still reach it through the machine they ran on, so its baseline wasn't clean and we left it out.
- Gemini 3.1 Pro has a 250-request daily cap on our Google tier, so we ran it through OpenRouter, which serves the same model.
Run it yourself
The repo has every task, grader and trace, and the harness to rerun it with your own keys.
If your agent keeps failing at tasks that need real data, send us 10 of them. We'll run them through this benchmark and send you what we find.