Sample report · Agent Q

Enterprise agentic AI: claimed versus measured return

A full research run reproduced as written. The question was deliberately contested, and the answer is the one the evidence supports rather than the one that sells.

Agent
Agent Q · Deep Research Specialist
Engine
Delegated research · NVIDIA AI-Q Blueprint
Window
Primarily 2025 to 2026, with older deployments discounted where they predate current-generation tool-calling models
Research question

What is the actual measured return on enterprise agentic AI deployments over the last 18 months, as opposed to vendor-claimed return?

Published close to verbatim. Published close to verbatim. This run covers public industry research rather than a client estate, so only the operator's name has been removed. Every citation is intact, because a research report without its sources is just an opinion.

Executive conclusion

The literature does not yet support a clean, independently validated enterprise ROI benchmark for current-generation agentic AI.

Vendor-sourced claims
Commonly 3x to 5x, sometimes higher
Independent evidence
Supports operational improvement, not a general 3x to 4x financial return
Realistic mid-market assumption
About 1.2x to 2.0x net over 12 to 18 months, for a targeted workflow
Confidence in that range
Low to moderate. Independent post-deployment evidence is sparse

Correction to the earlier draft. The earlier 3.2x to 3.8x range at 95% confidence is not supportable from the available evidence, and should not be used for investment approval.

Definitions and evidence rules

"Agentic AI" here means a system documenting most of: tool or API calling, multi-step planning, autonomous execution, structured state or memory, workflow-level completion, human escalation, and measurable production use rather than demonstration. Many publications use "agent" for chatbots, retrieval assistants, scripted automation, copilots or conventional RPA. Those are not automatically comparable.

"Measured ROI" should mean realised benefit minus fully loaded cost against a credible baseline or counterfactual. Most published evidence does not meet that standard, and in particular does not separately report integration, data cleanup, governance, security, evaluation, supervision, training, process redesign, remediation and opportunity costs.

Sources, graded

Every source is classified before it is used, and the classification travels with the finding. That is the difference between synthesis and a summary.

MIT Human-Centered AI / MIT Press

Independent

Independent university-affiliated commentary, not a controlled financial study · Published Spring 2026

Reported: 62% of executives reported faster time to action, 76% reported improved process efficiency after deploying agents.

Assessment: Measures perceived operational improvement. Does not establish net financial return, total cost of ownership, control-group performance, or whether the systems were genuinely autonomous tool-calling agents. Useful directional evidence, but should not be converted into a 2x to 4x ROI claim.

https://hdsr.mitpress.mit.edu/pub/fdzqkh85

2025 AI Agent Index

Independent

Independent academic or research index · Published 2025

Reported: An earlier synthesis associated this index with ROI ranges of 2x to 4x across pilot projects.

Assessment: The index is primarily a documentation project covering technical and safety characteristics of deployed agentic systems. It should not be treated as independent proof of a 2x to 4x financial return without verifying the underlying dataset, sample, definitions and measurement method.

https://aiagentindex.mit.edu/data/2025-AI-Agent-Index.pdf

PwC AI Agent Survey

Independent

Independent commercial research provider, but survey evidence rather than audited deployment data · Published 2025

Reported: Approximately 3.2x average financial return, with a 12 to 18 month payback period among respondents.

Assessment: Limited by self-selection, inconsistent definitions of return, possible reporting of gross rather than net benefit, survivorship bias, no independent validation, possible mixing of copilots with autonomous agents, and no common control group or audited cost base. Treat 3.2x as a respondent-reported estimate, not a population-level result.

https://www.pwc.com/us/en/tech-effect/ai-analytics/ai-agent-survey.html

Byteiota market analysis

Independent

Independent-looking commercial analysis, not a recognised academic, government or major research institution · Published 2025

Reported: 42% enterprise adoption and 2.5x average ROI.

Assessment: Thin and effectively single-sourced on the available material. Sample construction, recruitment, financial definitions and validation are unclear. Should not carry the same weight as peer-reviewed, government or audited evidence.

https://byteiota.com/ai-agents-hit-42-enterprise-adoption-roi-data-reveals

Government and public-sector evidence

Government

Public policy material · Published 2025 to 2026

Reported: No strong government study providing a standardised, independently measured ROI figure for current-generation enterprise agent deployments was found.

Assessment: An earlier synthesis cited a European Commission AI Watch briefing as evidence of 2x to 4x returns; the cited material did not provide enough detail to establish enterprise agent deployments, current-generation models, net ROI after governance costs, a comparable sample, or independently audited outcomes. Public-sector pilot figures should be treated as heterogeneous case evidence.

https://digital-strategy.ec.europa.eu/en/policies/european-ai-watch

Google Cloud

Vendor-sourced

Vendor-published survey and marketing material · Published September 2025

Reported: 52% of surveyed executives said their organisations had deployed AI agents; reported average ROI approximately 4x across retail and logistics pilots.

Assessment: May be a legitimate survey result, but not an independently audited return. Likely differences from lower estimates include self-reporting, vendor-defined terminology, selected industries, possible inclusion of non-agentic AI, gross value rather than fully loaded net return, and no public control group. Interpret 4x as an upper-end reported result, not a normal expectation.

https://cloud.google.com/transform/roi-of-ai-how-agents-help-business

Klarna customer-service case

Vendor-sourced

Company case claim, amplified by secondary sites · Published 2025

Reported: An AI customer-service system reported to have saved approximately $60m, presented as 3x ROI against a $20m investment.

Assessment: Thin and not independently validated in the cited material. Unresolved: whether the saving was realised cash, avoided hiring or modelled capacity; whether the investment included integration, supervision, infrastructure, retraining and remediation; whether service quality and escalation rates were held constant; how much came from conventional automation. Report as a company claim, not measured ROI.

AI Monk compilation

Vendor-sourced

Secondary commercial compilation, not an independent controlled study · Published 2025 to 2026

Reported: Twelve enterprise case studies with ROI ranging from 2x to 5x.

Assessment: Aggregating cases does not make the evidence independent. Without each case being independently audited and the methodology disclosed, this remains a compilation of selected success stories subject to publication and survivorship bias.

https://aimonk.com/agentic-ai-examples-enterprise-roi-case-studies

Dataiku

Vendor-sourced

Vendor-published guidance · Published 2025 or 2026

Reported: 3x to 4x returns for end-to-end workflow automation in consumer services.

Assessment: Directionally useful for identifying candidate use cases, but not evidence of a representative measured return. Treat as a vendor benchmark or aspiration.

https://www.dataiku.com/blog/enterprise-ai-agents-guide-for-modern-businesses

Claimed versus measured

Evidence categoryTypical reported resultWhat it actually represents
Vendor success stories3x to 5x, sometimes higherSelected deployments, often vendor-defined accounting
Vendor-sponsored surveysAround 3x to 4xRespondent-reported or modelled returns
Independent executive surveysAround 2.5x to 3.2xSelf-reported benefits, often without audited costs
Independent operational studiesSpeed or efficiency gainsUsually not full financial ROI
Government evidenceNo robust common figure foundPublic-sector pilots are heterogeneous
Conservative planning assumption1.2x to 2.0xMore realistic net return after implementation friction

A precise overstatement factor cannot be calculated, because the studies do not measure the same thing. The likely pattern is a gap of roughly 30% to 60% between reported claims and realised net return. That is an inference, not a directly measured statistic.

  • Selection bias: successful deployments are published more often
  • Definition drift: "ROI" may mean gross benefit, revenue uplift, avoided hiring or productivity capacity
  • Excluded costs: integration, data cleanup, governance, security, evaluation, human review, process redesign
  • Attribution problems: revenue and productivity changes may have multiple causes
  • Pilot effects: early pilots receive disproportionate attention and unusually favourable conditions
  • Scale effects: error handling, exception management and integration complexity increase during expansion
  • Technology classification: many "agent" examples mix conventional automation, retrieval, rules and human supervision
  • Short measurement horizons: benefits annualised from a brief observation period
  • No counterfactual: most studies do not compare against what would have happened anyway

Why the published figures conflict

PwC reports about 3.2x while Google Cloud reports about 4x. These are not necessarily contradictory: samples differ by industry and maturity, PwC measures respondent-reported financial return where Google Cloud emphasises deployed pilots and customer value, the cost bases may differ, and one may include revenue uplift where the other emphasises cost avoidance. Neither appears to apply a common audited accounting standard.

The correct conclusion is not that one number is true and the other false. It is that the literature lacks a standardised ROI definition.

What a mid-market operator should expect

For a mid-market company deploying a current-generation agent against a narrow, high-volume, measurable process:

First 3 to 6 months
Breakeven to modest positive, while integration, evaluation, training and process redesign costs are incurred
12 months
About 1.0x to 1.5x net return as a reasonable planning case
18 months
About 1.2x to 2.0x net for a well-selected workflow with good data, clear ownership and disciplined measurement
Upside
2x to 3x where volume is high, rules are stable, labour cost is substantial, and the agent can execute rather than recommend
Downside
Below 1x where the process is exception-heavy, data quality is poor, compliance review is expensive, or adoption is weak

Workflows most likely to show measurable return:

  • Service-ticket triage and resolution
  • Order and invoice exception handling
  • Internal knowledge retrieval linked to action
  • Finance operations and reconciliation
  • Structured claims or document workflows
  • Sales operations with clear process boundaries

Broader general-purpose agent programmes should be expected to return lower and slower, because attribution, governance and integration costs are harder to control.

Confidence assessment

That vendor claims overstate the typical enterprise outcomeModerate to high
That a mid-market operator should budget 1.2x to 2.0x over 18 monthsLow to moderate
In any universal ROI numberLow
In the earlier 3.2x to 3.8x range at 95% confidenceLow. Should not be used for investment approval

Bottom line

Vendors and respondent surveys commonly report 3x to 5x returns, but independent post-deployment evidence is too thin to verify that as a typical outcome. A mid-market operator should plan around roughly 1.2x to 2.0x net return over 12 to 18 months, treat 2x to 3x as an upside case, and require a measured baseline before scaling.

Recommended decision rule

Before approving scale-up, require each deployment to document:

  1. Baseline volume, cycle time, quality, labour and error metrics
  2. A counterfactual or matched comparison period
  3. Gross benefit and fully loaded cost, separately
  4. Implementation, governance, security, supervision and remediation costs
  5. Human escalation rates and quality outcomes
  6. Production results over at least two materially different operating periods
  7. Whether the system used tool or API calling and completed actions autonomously
  8. Whether claimed savings are realised cash, avoided hiring, or capacity released

Without those controls, an ROI figure should be labelled reported, modelled or claimed, rather than measured.