Outcomes

What AI delivers, and whether it lasts

Compare reported ROI, agent and non-agent results, deployment durability, and the implementation choices behind them — grounded in source-linked enterprise AI evidence.

Cases with outcomes

1,536

39.3% of the catalog
Metrics extracted

3,021

structured quantified results
Percent-comparable

1,365

normalized for benchmarking
Outcome categories

11

with reported evidence

1,536 of 3,905 public AI cases report measurable results. Time & speed leads the evidence: a median 50% reduction across 313 reported metrics.

Durability

Do reported outcomes last?

Of the 166 deployments first documented in 2023 that can be judged, 97% still have live public evidence today. This measures evidence durability, not whether a system remains in production.

Organization expanded

66.9%

Original source live

30.1%

Lost public footprint

3%

Explore deployment persistence

Implementation

Build, buy, or compose?

How teams delivered the outcomes in 2,748 deployments where the implementation approach is documented.

BuildBuyComposeMixed

2,748 of 3,826 cases classified (72%)

Compare build vs buy

Agents vs classic AI

Do agents change the ROI math?

Median reported result per outcome category, agentic vs non-agentic deployments · A larger magnitude is a larger reported change.

Agentic deployments report larger median results in productivity & throughput (+60% vs +40%), cost savings (−47% vs −40%), customer experience, but smaller revenue & growth and quality & accuracy and automation & deflection results, and time & speed is a wash.

Agentic deploymentsClassic AI

Time & speedeven

median reduction · n 118 vs 195

Cost savingsagents +7 pp

median reduction · n 42 vs 137

Productivity & throughputagents +20 pp

median improvement · n 35 vs 73

Revenue & growthclassic +14 pp

median improvement · n 24 vs 45

Quality & accuracyclassic +12 pp

median improvement · n 21 vs 42

Customer experienceagents +5 pp

median improvement · n 27 vs 26

Automation & deflectionclassic +18 pp

median reduction · n 11 vs 13

277 agentic and 585 classic deployments report a comparable quantified result in this window.

ⓘ How this is measured

Each row compares the median normalized result reported by agentic deployments (IsAgentCase) against all other AI deployments, within one outcome category and one direction of change — the same direction-split medians as the outcome benchmarks above, so the modules reconcile. A category is shown only when both cohorts carry at least 8 comparable metrics in this window.

These are reported outcomes from vendor-linked evidence, not a controlled comparison — teams choose what to publish, and agent projects may simply target different work. The 2024+ toggle holds publication vintage constant because agent cases skew newer. Risk/safety (medians pin near 100%) and sustainability (too few agent metrics) are excluded.

See which outcomes are documented most often within each industry.

Compare industries

Quantified outcomes

Median results by outcome category

Each category's reported reduction or improvement, with the p25–p75 spread — normalized so you can benchmark magnitudes side by side.

Time & speed

891 cases · 1,303 metrics

Reported reduction

50%median · 313 metrics

p25 30%median 50%p75 77.8%

Reported improvement

60%median · 143 metrics

p25 30%median 60%p75 200%

Cost savings

270 cases · 311 metrics

Reported reduction

40%median · 179 metrics

p25 25%median 40%p75 70.7%

Reported improvement

53.5%median · 16 metrics

p25 28.8%median 53.5%p75 425%

Productivity & throughput

154 cases · 170 metrics

Reported reduction

50%median · 9 metrics

p25 46%median 50%p75 51.4%

Reported improvement

40%median · 108 metrics

p25 22.9%median 40%p75 70%

Quality & accuracy

115 cases · 128 metrics

Reported reduction

35%median · 26 metrics

p25 20%median 35%p75 49.8%

Reported improvement

41%median · 63 metrics

p25 20.5%median 41%p75 94%

Revenue & growth

90 cases · 112 metrics

Reported reduction

31%median · 2 metrics

p25 29%median 31%p75 33%

Reported improvement

40%median · 69 metrics

p25 15%median 40%p75 135%

Automation & deflection

79 cases · 88 metrics

Reported reduction

50%median · 24 metrics

p25 40%median 50%p75 71.2%

Reported improvement

49%median · 19 metrics

p25 28.5%median 49%p75 70%

Adoption & scale

73 cases · 94 metrics

Reported reduction

56%median · 2 metrics

p25 39%median 56%p75 73%

Reported improvement

250%median · 8 metrics

p25 134.2%median 250%p75 1,100%

Customer experience

58 cases · 68 metrics

Reported reduction

30%median · 8 metrics

p25 27.5%median 30%p75 57.5%

Reported improvement

25%median · 53 metrics

p25 20%median 25%p75 90%

Strategic outcomes

Non-quantified strategic results

3,900 cases report a strategic outcome that isn't a single number — a new business model, market expansion, a new capability, and more.

New product / capability

2,328 cases

Speed & agility

2,159 cases

Customer experience & trust

2,117 cases

Risk & compliance

1,382 cases

Scale & capacity

1,321 cases

Better decisions & insight

1,102 cases

Cost efficiency

1,038 cases

Employee experience

494 cases

Market & geographic expansion

290 cases

Innovation & culture

271 cases

Sustainability & ESG

227 cases

Ecosystem & partnerships

223 cases

Competitive differentiation

140 cases

New business model

100 cases

Other strategic outcome

258 cases

ⓘ How these benchmarks are measured

Quantified results are extracted from each case's published impact statements, assigned to a fixed outcome taxonomy, and normalized for comparability: percentages are used as reported (implausible figures are excluded), multipliers like "3x faster" are converted to percent equivalents, and time units are converted to hours. Medians and quartiles are computed over percent-comparable metrics only; vendor-published evidence skews toward successes, so treat these as reported outcomes, not guaranteed results.

Last computed 4 minutes ago.